Videos TNwJ1LMiENk
Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel
Scene timeline
51 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 210
- whisperx 210
- chunks
- 36
- from 210 cues
- keyframes
- 37
- kept of 51 captured
- frames with text
- 37
- 572 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.0 MB
- word timings on 210 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 13:42 | 1m 56s |
stt |
done | — | 2026-08-10 13:44 | 26s |
chunk |
done | — | 2026-08-10 13:44 | 0s |
text_embed |
done | — | 2026-08-10 19:50 | 0s |
keyframe |
done | — | 2026-08-10 13:45 | 1m 41s |
ocr |
done | — | 2026-08-10 13:47 | 14s |
frame_embed |
done | — | 2026-08-10 19:50 | 7s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Snorkel0.94
- The Frontier Al Data Lab0.99
- Stop Making Models Bigger.1.00
- Make Them Behave1.00
- Making a 4B Model Outperform 235B on1.00
- yol Use for Financial Analysis0.99
- Kobie Crawford, Developer Advocate1.00
- AlEngineer1.00
- EUROPE1.00
-
- SnorkeJ0.97
- The Frontier Al Data Lab0.98
- Stop Making Models Bigger.1.00
- AIE1.00
- Make Them Behave1.00
- ★1.00
- ★1.00
- Making a 4B Model Outperform 235B on1.00
- Tool Use for Financial Analysis1.00
- Kobie Crawford, Developer Advocate1.00
- Snorkel AIl | Proprietary & Confidential | Not for distribution0.97
- Make Th0.99
- Engineering the future of Al1.00
-
- SnorkeJ0.97
- The Frontier Al Data Lab1.00
- Stop Making Models Bigger.1.00
- *★★0.62
- AIE1.00
- Make Them Behave1.00
- ★1.00
- ★1.00
- ★1.00
- Making a 4B Model Outperform 235B on0.99
- Tool Use for Financial Analysis1.00
- Kobie Crawford, Developer Advocate1.00
- Snorkel AIl | Proprietary & Confidential | Not for distribution0.97
- Make Th0.99
- AlEngineer0.97
- EUROPE1.00
-
- Snorkel0.99
- The Frontier Al Data Lab0.99
- Stop Making Models Bigger.0.99
- Make Them Behave1.00
- Making a 4B Model Outperform 235B on0.99
- e for Financial Analysis0.97
- rd, Developer Advocate0.99
- AlEngineer0.99
- EUROPE1.00
-
- What we'll0.98
- 1 | Research Objective0.93
- *★★0.71
- cover1.00
- 2 | The Approach0.96
- AIE1.00
- 3 The Result0.97
- ★1.00
- ★1.00
- Snorkel Al | Proprietary & Confidential | Not for distribution0.97
- Engineering the future of Al0.99
-
- Background1.00
- Models are used for increasingly complex work in enterprise use cases0.99
- Financial analysis is a prime example:1.00
- ***0.50
- ★1.00
- O0.84
- Tool use1.00
- AIE1.00
- ★1.00
- Multi-step reasoning1.00
- ★1.00
- ★1.00
- ★1.00
- O0.94
- SQL execution across multiple schemas1.00
- Combining retrieved data with numerical calculations to produce correct,1.00
- verifiable answers1.00
- We see a trend where the models get bigger to1.00
- address the application need, but is this necessary?1.00
- Snorkel Al | Proprietary & Confidential | Not for distribution0.97
- 140.94
- Engineering the future of Al0.99
-
- Background1.00
- Models are used for increasingly complex work in enterprise use cases0.99
- Financial analysis is a prime example:0.99
- O0.74
- Tool use1.00
- ★1.00
- AIE1.00
- ★0.97
- Multi-step reasoning1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- O0.83
- SQL execution across multiple schemas1.00
- Combining retrieved data with numerical calculations to produce correct,1.00
- verifiable answers1.00
- We see a trend where the models get bigger to0.99
- address the application need, but is this necessary?1.00
- Snorkel Al | Proprietary & Confidential | Not for distribution0.98
- 140.95
- AlEngineer0.98
- EUROPE1.00
-
- Research Objective: Use RL to Get Smaller Models to0.99
- Achieve Parity with Larger Models on Tool-Use Tasks1.00
- Cost1.00
- Speed1.00
- Security1.00
- Pattern0.99
- /Control1.00
- ★1.00
- AIE1.00
- ★1.00
- Smaller models are1.00
- Less compute = faster1.00
- Easily deployed1.00
- Mirrors the classic1.00
- ★1.00
- dramatically cheaper1.00
- inference1.00
- on-prem, no external1.00
- POC-to-production arc1.00
- ★1.00
- ★1.00
- to run at inference1.00
- inference API1.00
- — prototype with a0.97
- dependency, data1.00
- large capable model,1.00
- remains inside1.00
- productionize with a1.00
- enterprise boundary0.98
- fine-tuned small model1.00
- RL is the right kind/phase of training –0.99
- this is a challenge of behavior, rather than core language capability/knowledge0.99
- Snorkel Al | Proprietary & Confidential | Not for distribution0.98
- 150.95
- Engineering the future of Al1.00
-
- Research Objective: Use RL to Get1.00
- Smaller Models to1.00
- Achieve Parity with Larger Models on1.00
- Tool-Use Tasks0.99
- Security1.00
- /Control1.00
- AlEngine0.98
- EUROPE1.00
-
- Qwen3 235B1.00
- AIE1.00
- ★1.00
- FinQA Tool Use0.98
- AlEngineer0.96
- EUROPE1.00
-
- "What is the YoY growth rate of YouTube ads revenue from1.00
- 2023 to 2024?"0.99
- Qwen3-235B:0.99
- # Skips table discovery, guesses a table name0.99
- AIE1.00
- ★1.00
- ★1.00
- sql_query(company_name="youtube", table_name="revenue_breakdown",0.99
- query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
- ★1.00
- ★1.00
- ★★0.96
- # Error: table revenue_breakdown for company youtube could not be found.0.99
- # Guesses again instead of calling get_table_names0.99
- sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
- query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
- #Error: table us_gaap_RevenueTable for company youtube could not be found.1.00
- # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.98
- # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.99
- Snorkel Al | Proprietary & Confidential | Not for distribution0.96
- |70.84
- Engineering the future of Al1.00
-
- "What is the YoY growth rate of YouTube ads revenue from0.99
- 2023 to 2024?"0.99
- Qwen3-235B:0.99
- # Skips table discovery, guesses a table name0.99
- AIE1.00
- ★1.00
- ★1.00
- sql_query(company_name="youtube", table_name="revenue_breakdown",0.99
- query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
- ★1.00
- ★1.00
- ★★0.96
- # Error: table revenue_breakdown for company youtube could not be found.0.98
- # Guesses again instead of calling get_table_names0.99
- sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
- query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
- #Error: table us_gaap_RevenueTable for company youtube could not be found.0.99
- # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.98
- # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.98
- Snorkel Al | Proprietary & Confidential | Not for distribution0.97
- 170.76
- Engineering the future of Al1.00
-
- "What is the YoY growth rate of YouTube ads revenue from0.99
- 2023 to 2024?"0.99
- AlEngine0.99
- EUROPE0.99
-
- "What is the YoY growth rate of YouTube ads revenue from0.99
- 2023 to 2024?"0.98
- Qwen3-235B:1.00
- # Skips table discovery, guesses a table name0.99
- AIE1.00
- ★1.00
- ★0.96
- sql_query(company_name="youtube", table_name="revenue_breakdown",1.00
- query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
- ★1.00
- ★1.00
- ★★0.92
- # Error: table revenue_breakdown for company youtube could not be found.0.99
- # Guesses again instead of calling get_table_names0.99
- sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
- query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
- #Error: table us_gaap_RevenueTable for company youtube could not be found.1.00
- # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.99
- # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.99
- Snorkel Al | Proprietary & Confidential | Not for distribution0.97
- |70.72
- Braintrust1.00
- WorkOS OpenAI0.95
-
- The Problem with Large Models on Tool-Use Tasks1.00
- Great reasoning, but no discipline in tool use:0.99
- Skip schema inspection steps1.00
- Fail to learn table structures and column names before querying1.00
- AIE1.00
- ★1.00
- Execute poorly formed SQL queries0.99
- ★1.00
- ★1.00
- ★1.00
- →Retrieve bad data0.98
- → Hallucinate answers0.98
- Greater reasoning capability does not guarantee better1.00
- task performance when tool use is undisciplined1.00
- Snorkel Al | Proprietary & Confidential | Not for distribution0.98
- |80.93
- Braintrust1.00
- WorkOS OpenAI0.97
-
- What we'll1.00
- 1 | Research Objective0.94
- *★★0.52
- cover1.00
- 2 The Approach0.98
- AIE1.00
- ★1.00
- 3 The Result0.97
- ★1.00
- ★1.00
- Snorkel Al |Proprietary & Confidential | Not for distribution0.98
- AlEngineer0.96
- EUROPE1.00
Transcript
210 cues· 3,748 words· 20,125 chars
- 0:16 This is the last presentation I have to give this conference, so I'm feeling already a little bit of the euphoria of like, ah, it's all done.
- 0:24 I know we're also close to the end of the whole sequence.
- 0:31 I keep finding these conferences to be some of the highest signal that I get wherever I go.
- 0:37 So generally, do people feel like they're getting what they came for here?
- 0:41 I'm just curious, because we're at Snorkel.
- 0:44 We put in our sponsorship, and we want to know that people are getting what they want.
- 0:47 They know they're going to come back, because we want to sponsor next time.
- 0:49 We want to know if people are happy about it.
- 0:51 Did you guys see what you wanted to see?
- 0:53 Yeah?
- 0:54 Yeah, really good.
- 0:55 Brilliant, brilliant.
- 0:58 So now that it is 345, I'm gonna go ahead and start the official thing.
- 1:03 So my name is Coby Crawford.
- 1:05 I'm a developer advocate at Snorkel.
- 1:08 We call ourselves the Frontier AI Data Lab, and what we're doing right now, our main thing is
- 1:15 Starting from the research-backed work that Snorkel has been doing since its inception, we've been working on a variety of things about data quality.
- 1:23 And at this point now, where we're focused is actually providing data sets where we assure a certain level of quality.
- 1:31 attentive to being very, very, very motivated about making sure the data is high quality.
- 1:35 And part of how we get to high quality is we always make sure to have sort of an expert in the loop as part of the process.
- 1:40 So we have expert contributors that we work with and we bring people in to provide their expertise to make sure that the data that we generate is of top quality.
- 1:48 And then for the top labs that want to use our data to improve their models and get the hill climbing done in the right way, that's what we do at Snorkel.
- 2:00 Because of that, a lot of what goes on is still more research.
- 2:05 And this is a talk that's talking about some of the work that we did, that our research team did.
- 2:10 And one of the keys in this research is that we're looking at
- 2:16 how the best quality data can be best applied and where it is that we need to be looking for where there are opportunities to get that done.
- 2:24 So in this particular case, talking about stop making models bigger, I mean, it's a nice punchy title.
- 2:30 Of course, we don't really mean that models shouldn't be large intrinsically,
- 2:34 The point, broadly speaking, is that sometimes we find great wins to be had with the right data applied to the right problem statement.
- 2:42 And so this is something we're going to talk about, a specific use case that our research team discovered.
- 2:46 And in partnership with the RLLM team, which is a research group, part of UC Berkeley.
- 2:54 And so the UC Berkeley team over there, RLLM, the agentica project, their lab partnered with us on this particular work.
- 3:02 So the goal, as I said, is making a four billion parameter model outperform a 235 billion parameter model on tool use tasks for financial analysis.
- 3:14 So we'll start with the research objective and then we'll iterate through talking about the approach that was used for this particular process and then talk about the results and happy to report that we got what we were looking for, so good things to be had.
- 3:32 Couple of quick level setting backgrounds of what we're talking about here.
- 3:35 First is that as we see enterprise use cases take on some greater complexity.
- 3:42 Obviously, we've got the massive explosion of what people are doing in terms of personal assistance.
- 3:46 And as people are working in the context of enterprise, a lot of times you still need a sort of more constrained
- 3:53 choice about how to implement something and make sure that it's reliable.
- 3:56 When you're looking for things that are going to be done for enterprise production use cases, you kind of also have to make sure there's a lot of safety and security things done.
- 4:05 So these other kind of priorities that fold into what people typically want to do, we're looking at these things and saying, OK, well, these are the enterprise use cases that people have.
- 4:15 And as people try to solve the problems of making the models perform at the level that makes it acceptable for actually being deployed as a production service,
- 4:24 we see very often that people choose like, well, okay, we didn't get the performance that we wanted with this right now, we'll just drop in a larger model, it'll be smarter, it has greater reasoning skills, and we'll just sort of expect that the performance will improve commensurate with the additional load of the size of the model and the greater inference cost that goes along with that.
- 4:43 And in some cases, that might not always be the right thing.
- 4:45 So we see people saying, you know, let's just go get a bigger model, that'll solve the problem.
- 4:49 And sometimes maybe that isn't quite the answer.
- 4:55 In this case, what we're trying to do is to say, can we take a smaller model and then use RL with the right data to yield the kind of performance gains that we're looking for and to deliver the kind of application functionality that we want?
- 5:07 And so that's the target here.
- 5:09 And again, for these various reasons, cost, speed, security, and then the idea that in general, you start with a really big model and make your POC and make it work and everybody's happy that it works.
- 5:21 And it's like, okay, now what do we do to productionize it?
loading