read-only demo

Videos TNwJ1LMiENk

Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel

index_state ready data_status ok

AI Engineer· published 2026-06-10· 0:20:55· en-US· indexed 2026-08-10 19:50

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:56, 1 of 1 keyframes kept
  5. Shot 4, 0:56 to 1:24, 1 of 1 keyframes kept
  6. Shot 5, 1:24 to 1:53, 0 of 1 keyframes kept
  7. Shot 6, 1:53 to 2:21, 1 of 1 keyframes kept
  8. Shot 7, 2:21 to 2:47, 1 of 1 keyframes kept
  9. Shot 8, 2:47 to 3:13, 0 of 1 keyframes kept
  10. Shot 9, 3:13 to 3:29, 1 of 1 keyframes kept
  11. Shot 10, 3:29 to 3:57, 1 of 1 keyframes kept
  12. Shot 11, 3:57 to 4:25, 1 of 1 keyframes kept
  13. Shot 12, 4:25 to 4:53, 0 of 1 keyframes kept
  14. Shot 13, 4:53 to 5:13, 1 of 1 keyframes kept
  15. Shot 14, 5:13 to 5:43, 1 of 1 keyframes kept
  16. Shot 15, 5:43 to 6:11, 0 of 1 keyframes kept
  17. Shot 16, 6:11 to 6:39, 1 of 1 keyframes kept
  18. Shot 17, 6:39 to 7:07, 0 of 1 keyframes kept
  19. Shot 18, 7:07 to 7:34, 1 of 1 keyframes kept
  20. Shot 19, 7:34 to 8:00, 1 of 1 keyframes kept
  21. Shot 20, 8:00 to 8:24, 1 of 1 keyframes kept
  22. Shot 21, 8:24 to 8:31, 1 of 1 keyframes kept
  23. Shot 22, 8:31 to 8:59, 1 of 1 keyframes kept
  24. Shot 23, 8:59 to 9:08, 1 of 1 keyframes kept
  25. Shot 24, 9:08 to 9:41, 1 of 1 keyframes kept
  26. Shot 25, 9:41 to 10:13, 1 of 1 keyframes kept
  27. Shot 26, 10:13 to 10:46, 0 of 1 keyframes kept
  28. Shot 27, 10:46 to 11:04, 1 of 1 keyframes kept
  29. Shot 28, 11:04 to 11:26, 1 of 1 keyframes kept
  30. Shot 29, 11:26 to 12:06, 1 of 1 keyframes kept
  31. Shot 30, 12:06 to 12:37, 1 of 1 keyframes kept
  32. Shot 31, 12:37 to 13:08, 0 of 1 keyframes kept
  33. Shot 32, 13:08 to 13:39, 1 of 1 keyframes kept
  34. Shot 33, 13:39 to 13:51, 1 of 1 keyframes kept
  35. Shot 34, 13:51 to 14:17, 0 of 1 keyframes kept
  36. Shot 35, 14:17 to 14:36, 1 of 1 keyframes kept
  37. Shot 36, 14:36 to 15:08, 0 of 1 keyframes kept
  38. Shot 37, 15:08 to 15:39, 0 of 1 keyframes kept
  39. Shot 38, 15:39 to 16:10, 0 of 1 keyframes kept
  40. Shot 39, 16:10 to 16:17, 1 of 1 keyframes kept
  41. Shot 40, 16:17 to 16:33, 1 of 1 keyframes kept
  42. Shot 41, 16:33 to 16:56, 1 of 1 keyframes kept
  43. Shot 42, 16:56 to 17:31, 0 of 1 keyframes kept
  44. Shot 43, 17:31 to 18:05, 1 of 1 keyframes kept
  45. Shot 44, 18:05 to 18:29, 1 of 1 keyframes kept
  46. Shot 45, 18:29 to 18:53, 0 of 1 keyframes kept
  47. Shot 46, 18:53 to 19:22, 1 of 1 keyframes kept
  48. Shot 47, 19:22 to 19:44, 0 of 1 keyframes kept
  49. Shot 48, 19:44 to 20:07, 1 of 1 keyframes kept
  50. Shot 49, 20:07 to 20:40, 1 of 1 keyframes kept
  51. Shot 50, 20:40 to 20:55, 1 of 1 keyframes kept

51 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
210
whisperx 210
chunks
36
from 210 cues
keyframes
37
kept of 51 captured
frames with text
37
572 lines read
chapters
0
from the source metadata
keyframe bytes
6.0 MB
word timings on 210 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 13:42 1m 56s
stt done 2026-08-10 13:44 26s
chunk done 2026-08-10 13:44 0s
text_embed done 2026-08-10 19:50 0s
keyframe done 2026-08-10 13:45 1m 41s
ocr done 2026-08-10 13:47 14s
frame_embed done 2026-08-10 19:50 7s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:47 #3 done9 line(s)

    shot 3·sharpness 320.5

    1. Snorkel0.94
    2. The Frontier Al Data Lab0.99
    3. Stop Making Models Bigger.1.00
    4. Make Them Behave1.00
    5. Making a 4B Model Outperform 235B on1.00
    6. yol Use for Financial Analysis0.99
    7. Kobie Crawford, Developer Advocate1.00
    8. AlEngineer1.00
    9. EUROPE1.00
  • 1:15 #4 done13 line(s)

    shot 4·sharpness 4223.9

    1. SnorkeJ0.97
    2. The Frontier Al Data Lab0.98
    3. Stop Making Models Bigger.1.00
    4. AIE1.00
    5. Make Them Behave1.00
    6. 1.00
    7. 1.00
    8. Making a 4B Model Outperform 235B on1.00
    9. Tool Use for Financial Analysis1.00
    10. Kobie Crawford, Developer Advocate1.00
    11. Snorkel AIl | Proprietary & Confidential | Not for distribution0.97
    12. Make Th0.99
    13. Engineering the future of Al1.00
  • 1:36 #5 skipped

    shot 5·duplicate of #4

  • 1:59 #6 done16 line(s)

    shot 6·sharpness 4210.9

    1. SnorkeJ0.97
    2. The Frontier Al Data Lab1.00
    3. Stop Making Models Bigger.1.00
    4. *★★0.62
    5. AIE1.00
    6. Make Them Behave1.00
    7. 1.00
    8. 1.00
    9. 1.00
    10. Making a 4B Model Outperform 235B on0.99
    11. Tool Use for Financial Analysis1.00
    12. Kobie Crawford, Developer Advocate1.00
    13. Snorkel AIl | Proprietary & Confidential | Not for distribution0.97
    14. Make Th0.99
    15. AlEngineer0.97
    16. EUROPE1.00
  • 2:27 #7 done9 line(s)

    shot 7·sharpness 316.1

    1. Snorkel0.99
    2. The Frontier Al Data Lab0.99
    3. Stop Making Models Bigger.0.99
    4. Make Them Behave1.00
    5. Making a 4B Model Outperform 235B on0.99
    6. e for Financial Analysis0.97
    7. rd, Developer Advocate0.99
    8. AlEngineer0.99
    9. EUROPE1.00
  • 2:58 #8 skipped

    shot 8·duplicate of #4

  • 3:18 #9 done11 line(s)

    shot 9·sharpness 1554.8

    1. What we'll0.98
    2. 1 | Research Objective0.93
    3. *★★0.71
    4. cover1.00
    5. 2 | The Approach0.96
    6. AIE1.00
    7. 3 The Result0.97
    8. 1.00
    9. 1.00
    10. Snorkel Al | Proprietary & Confidential | Not for distribution0.97
    11. Engineering the future of Al0.99
  • 3:40 #10 done22 line(s)

    shot 10·sharpness 4720.7

    1. Background1.00
    2. Models are used for increasingly complex work in enterprise use cases0.99
    3. Financial analysis is a prime example:1.00
    4. ***0.50
    5. 1.00
    6. O0.84
    7. Tool use1.00
    8. AIE1.00
    9. 1.00
    10. Multi-step reasoning1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. O0.94
    15. SQL execution across multiple schemas1.00
    16. Combining retrieved data with numerical calculations to produce correct,1.00
    17. verifiable answers1.00
    18. We see a trend where the models get bigger to1.00
    19. address the application need, but is this necessary?1.00
    20. Snorkel Al | Proprietary & Confidential | Not for distribution0.97
    21. 140.94
    22. Engineering the future of Al0.99
  • 4:08 #11 done23 line(s)

    shot 11·sharpness 4706.0

    1. Background1.00
    2. Models are used for increasingly complex work in enterprise use cases0.99
    3. Financial analysis is a prime example:0.99
    4. O0.74
    5. Tool use1.00
    6. 1.00
    7. AIE1.00
    8. 0.97
    9. Multi-step reasoning1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. O0.83
    15. SQL execution across multiple schemas1.00
    16. Combining retrieved data with numerical calculations to produce correct,1.00
    17. verifiable answers1.00
    18. We see a trend where the models get bigger to0.99
    19. address the application need, but is this necessary?1.00
    20. Snorkel Al | Proprietary & Confidential | Not for distribution0.98
    21. 140.95
    22. AlEngineer0.98
    23. EUROPE1.00
  • 4:28 #12 skipped

    shot 12·duplicate of #10

  • 4:55 #13 done35 line(s)

    shot 13·sharpness 5082.8

    1. Research Objective: Use RL to Get Smaller Models to0.99
    2. Achieve Parity with Larger Models on Tool-Use Tasks1.00
    3. Cost1.00
    4. Speed1.00
    5. Security1.00
    6. Pattern0.99
    7. /Control1.00
    8. 1.00
    9. AIE1.00
    10. 1.00
    11. Smaller models are1.00
    12. Less compute = faster1.00
    13. Easily deployed1.00
    14. Mirrors the classic1.00
    15. 1.00
    16. dramatically cheaper1.00
    17. inference1.00
    18. on-prem, no external1.00
    19. POC-to-production arc1.00
    20. 1.00
    21. 1.00
    22. to run at inference1.00
    23. inference API1.00
    24. — prototype with a0.97
    25. dependency, data1.00
    26. large capable model,1.00
    27. remains inside1.00
    28. productionize with a1.00
    29. enterprise boundary0.98
    30. fine-tuned small model1.00
    31. RL is the right kind/phase of training –0.99
    32. this is a challenge of behavior, rather than core language capability/knowledge0.99
    33. Snorkel Al | Proprietary & Confidential | Not for distribution0.98
    34. 150.95
    35. Engineering the future of Al1.00
  • 5:37 #14 done8 line(s)

    shot 14·sharpness 231.4

    1. Research Objective: Use RL to Get1.00
    2. Smaller Models to1.00
    3. Achieve Parity with Larger Models on1.00
    4. Tool-Use Tasks0.99
    5. Security1.00
    6. /Control1.00
    7. AlEngine0.98
    8. EUROPE1.00
  • 5:52 #15 skipped

    shot 15·duplicate of #13

  • 6:25 #16 done6 line(s)

    shot 16·sharpness 2073.8

    1. Qwen3 235B1.00
    2. AIE1.00
    3. 1.00
    4. FinQA Tool Use0.98
    5. AlEngineer0.96
    6. EUROPE1.00
  • 7:04 #17 skipped

    shot 17·duplicate of #16

  • 7:25 #18 done22 line(s)

    shot 18·sharpness 4656.0

    1. "What is the YoY growth rate of YouTube ads revenue from1.00
    2. 2023 to 2024?"0.99
    3. Qwen3-235B:0.99
    4. # Skips table discovery, guesses a table name0.99
    5. AIE1.00
    6. 1.00
    7. 1.00
    8. sql_query(company_name="youtube", table_name="revenue_breakdown",0.99
    9. query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
    10. 1.00
    11. 1.00
    12. ★★0.96
    13. # Error: table revenue_breakdown for company youtube could not be found.0.99
    14. # Guesses again instead of calling get_table_names0.99
    15. sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
    16. query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
    17. #Error: table us_gaap_RevenueTable for company youtube could not be found.1.00
    18. # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.98
    19. # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.99
    20. Snorkel Al | Proprietary & Confidential | Not for distribution0.96
    21. |70.84
    22. Engineering the future of Al1.00
  • 7:39 #19 done22 line(s)

    shot 19·sharpness 4647.9

    1. "What is the YoY growth rate of YouTube ads revenue from0.99
    2. 2023 to 2024?"0.99
    3. Qwen3-235B:0.99
    4. # Skips table discovery, guesses a table name0.99
    5. AIE1.00
    6. 1.00
    7. 1.00
    8. sql_query(company_name="youtube", table_name="revenue_breakdown",0.99
    9. query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
    10. 1.00
    11. 1.00
    12. ★★0.96
    13. # Error: table revenue_breakdown for company youtube could not be found.0.98
    14. # Guesses again instead of calling get_table_names0.99
    15. sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
    16. query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
    17. #Error: table us_gaap_RevenueTable for company youtube could not be found.0.99
    18. # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.98
    19. # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.98
    20. Snorkel Al | Proprietary & Confidential | Not for distribution0.97
    21. 170.76
    22. Engineering the future of Al1.00
  • 8:10 #20 done4 line(s)

    shot 20·sharpness 234.3

    1. "What is the YoY growth rate of YouTube ads revenue from0.99
    2. 2023 to 2024?"0.99
    3. AlEngine0.99
    4. EUROPE0.99
  • 8:30 #21 done23 line(s)

    shot 21·sharpness 4727.7

    1. "What is the YoY growth rate of YouTube ads revenue from0.99
    2. 2023 to 2024?"0.98
    3. Qwen3-235B:1.00
    4. # Skips table discovery, guesses a table name0.99
    5. AIE1.00
    6. 1.00
    7. 0.96
    8. sql_query(company_name="youtube", table_name="revenue_breakdown",1.00
    9. query="SELECT youtube_ads_2023, youtube_ads_2024 FROM revenue_breakdown LIMIT0.99
    10. 1.00
    11. 1.00
    12. ★★0.92
    13. # Error: table revenue_breakdown for company youtube could not be found.0.99
    14. # Guesses again instead of calling get_table_names0.99
    15. sql_query(company_name="youtube", table_name="us_gaap_RevenueTable",1.00
    16. query="SELECT * FROM us_gaap_RevenueTable WHERE segment = 'YouTube'")0.98
    17. #Error: table us_gaap_RevenueTable for company youtube could not be found.1.00
    18. # Hallucination: Fabricates a growth rate (Actual is ~14.7%)0.99
    19. # "YouTube ads revenue grew approximately 20-25% year-over-year in 2024..."0.99
    20. Snorkel Al | Proprietary & Confidential | Not for distribution0.97
    21. |70.72
    22. Braintrust1.00
    23. WorkOS OpenAI0.95
  • 8:35 #22 done18 line(s)

    shot 22·sharpness 4711.3

    1. The Problem with Large Models on Tool-Use Tasks1.00
    2. Great reasoning, but no discipline in tool use:0.99
    3. Skip schema inspection steps1.00
    4. Fail to learn table structures and column names before querying1.00
    5. AIE1.00
    6. 1.00
    7. Execute poorly formed SQL queries0.99
    8. 1.00
    9. 1.00
    10. 1.00
    11. →Retrieve bad data0.98
    12. → Hallucinate answers0.98
    13. Greater reasoning capability does not guarantee better1.00
    14. task performance when tool use is undisciplined1.00
    15. Snorkel Al | Proprietary & Confidential | Not for distribution0.98
    16. |80.93
    17. Braintrust1.00
    18. WorkOS OpenAI0.97
  • 9:06 #23 done13 line(s)

    shot 23·sharpness 1546.1

    1. What we'll1.00
    2. 1 | Research Objective0.94
    3. *★★0.52
    4. cover1.00
    5. 2 The Approach0.98
    6. AIE1.00
    7. 1.00
    8. 3 The Result0.97
    9. 1.00
    10. 1.00
    11. Snorkel Al |Proprietary & Confidential | Not for distribution0.98
    12. AlEngineer0.96
    13. EUROPE1.00

Transcript

210 cues· 3,748 words· 20,125 chars

  1. 0:16 This is the last presentation I have to give this conference, so I'm feeling already a little bit of the euphoria of like, ah, it's all done.
  2. 0:24 I know we're also close to the end of the whole sequence.
  3. 0:31 I keep finding these conferences to be some of the highest signal that I get wherever I go.
  4. 0:37 So generally, do people feel like they're getting what they came for here?
  5. 0:41 I'm just curious, because we're at Snorkel.
  6. 0:44 We put in our sponsorship, and we want to know that people are getting what they want.
  7. 0:47 They know they're going to come back, because we want to sponsor next time.
  8. 0:49 We want to know if people are happy about it.
  9. 0:51 Did you guys see what you wanted to see?
  10. 0:53 Yeah?
  11. 0:54 Yeah, really good.
  12. 0:55 Brilliant, brilliant.
  13. 0:58 So now that it is 345, I'm gonna go ahead and start the official thing.
  14. 1:03 So my name is Coby Crawford.
  15. 1:05 I'm a developer advocate at Snorkel.
  16. 1:08 We call ourselves the Frontier AI Data Lab, and what we're doing right now, our main thing is
  17. 1:15 Starting from the research-backed work that Snorkel has been doing since its inception, we've been working on a variety of things about data quality.
  18. 1:23 And at this point now, where we're focused is actually providing data sets where we assure a certain level of quality.
  19. 1:31 attentive to being very, very, very motivated about making sure the data is high quality.
  20. 1:35 And part of how we get to high quality is we always make sure to have sort of an expert in the loop as part of the process.
  21. 1:40 So we have expert contributors that we work with and we bring people in to provide their expertise to make sure that the data that we generate is of top quality.
  22. 1:48 And then for the top labs that want to use our data to improve their models and get the hill climbing done in the right way, that's what we do at Snorkel.
  23. 2:00 Because of that, a lot of what goes on is still more research.
  24. 2:05 And this is a talk that's talking about some of the work that we did, that our research team did.
  25. 2:10 And one of the keys in this research is that we're looking at
  26. 2:16 how the best quality data can be best applied and where it is that we need to be looking for where there are opportunities to get that done.
  27. 2:24 So in this particular case, talking about stop making models bigger, I mean, it's a nice punchy title.
  28. 2:30 Of course, we don't really mean that models shouldn't be large intrinsically,
  29. 2:34 The point, broadly speaking, is that sometimes we find great wins to be had with the right data applied to the right problem statement.
  30. 2:42 And so this is something we're going to talk about, a specific use case that our research team discovered.
  31. 2:46 And in partnership with the RLLM team, which is a research group, part of UC Berkeley.
  32. 2:54 And so the UC Berkeley team over there, RLLM, the agentica project, their lab partnered with us on this particular work.
  33. 3:02 So the goal, as I said, is making a four billion parameter model outperform a 235 billion parameter model on tool use tasks for financial analysis.
  34. 3:14 So we'll start with the research objective and then we'll iterate through talking about the approach that was used for this particular process and then talk about the results and happy to report that we got what we were looking for, so good things to be had.
  35. 3:32 Couple of quick level setting backgrounds of what we're talking about here.
  36. 3:35 First is that as we see enterprise use cases take on some greater complexity.
  37. 3:42 Obviously, we've got the massive explosion of what people are doing in terms of personal assistance.
  38. 3:46 And as people are working in the context of enterprise, a lot of times you still need a sort of more constrained
  39. 3:53 choice about how to implement something and make sure that it's reliable.
  40. 3:56 When you're looking for things that are going to be done for enterprise production use cases, you kind of also have to make sure there's a lot of safety and security things done.
  41. 4:05 So these other kind of priorities that fold into what people typically want to do, we're looking at these things and saying, OK, well, these are the enterprise use cases that people have.
  42. 4:15 And as people try to solve the problems of making the models perform at the level that makes it acceptable for actually being deployed as a production service,
  43. 4:24 we see very often that people choose like, well, okay, we didn't get the performance that we wanted with this right now, we'll just drop in a larger model, it'll be smarter, it has greater reasoning skills, and we'll just sort of expect that the performance will improve commensurate with the additional load of the size of the model and the greater inference cost that goes along with that.
  44. 4:43 And in some cases, that might not always be the right thing.
  45. 4:45 So we see people saying, you know, let's just go get a bigger model, that'll solve the problem.
  46. 4:49 And sometimes maybe that isn't quite the answer.
  47. 4:55 In this case, what we're trying to do is to say, can we take a smaller model and then use RL with the right data to yield the kind of performance gains that we're looking for and to deliver the kind of application functionality that we want?
  48. 5:07 And so that's the target here.
  49. 5:09 And again, for these various reasons, cost, speed, security, and then the idea that in general, you start with a really big model and make your POC and make it work and everybody's happy that it works.
  50. 5:21 And it's like, okay, now what do we do to productionize it?

Open at this second