read-only demo

Videos KMR_RBoCa4M

SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe)

index_state ready data_status ok

AI Engineer· published 2026-07-29· 0:16:29· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:31, 1 of 1 keyframes kept
  5. Shot 4, 0:31 to 1:18, 1 of 1 keyframes kept
  6. Shot 5, 1:18 to 1:33, 1 of 1 keyframes kept
  7. Shot 6, 1:33 to 2:18, 1 of 1 keyframes kept
  8. Shot 7, 2:18 to 3:03, 1 of 1 keyframes kept
  9. Shot 8, 3:03 to 3:30, 1 of 1 keyframes kept
  10. Shot 9, 3:30 to 3:58, 0 of 1 keyframes kept
  11. Shot 10, 3:58 to 4:47, 1 of 1 keyframes kept
  12. Shot 11, 4:47 to 5:15, 0 of 1 keyframes kept
  13. Shot 12, 5:15 to 5:43, 1 of 1 keyframes kept
  14. Shot 13, 5:43 to 6:28, 1 of 1 keyframes kept
  15. Shot 14, 6:28 to 6:41, 1 of 1 keyframes kept
  16. Shot 15, 6:41 to 7:18, 0 of 1 keyframes kept
  17. Shot 16, 7:18 to 7:45, 1 of 1 keyframes kept
  18. Shot 17, 7:45 to 8:13, 0 of 1 keyframes kept
  19. Shot 18, 8:13 to 8:41, 0 of 1 keyframes kept
  20. Shot 19, 8:41 to 9:28, 1 of 1 keyframes kept
  21. Shot 20, 9:28 to 10:07, 0 of 1 keyframes kept
  22. Shot 21, 10:07 to 10:40, 1 of 1 keyframes kept
  23. Shot 22, 10:40 to 11:12, 0 of 1 keyframes kept
  24. Shot 23, 11:12 to 11:53, 1 of 1 keyframes kept
  25. Shot 24, 11:53 to 12:26, 1 of 1 keyframes kept
  26. Shot 25, 12:26 to 12:59, 0 of 1 keyframes kept
  27. Shot 26, 12:59 to 13:31, 0 of 1 keyframes kept
  28. Shot 27, 13:31 to 13:58, 1 of 1 keyframes kept
  29. Shot 28, 13:58 to 14:24, 0 of 1 keyframes kept
  30. Shot 29, 14:24 to 14:56, 1 of 1 keyframes kept
  31. Shot 30, 14:56 to 15:28, 0 of 1 keyframes kept
  32. Shot 31, 15:28 to 16:00, 0 of 1 keyframes kept
  33. Shot 32, 16:00 to 16:02, 1 of 1 keyframes kept
  34. Shot 33, 16:02 to 16:12, 1 of 1 keyframes kept
  35. Shot 34, 16:12 to 16:28, 0 of 1 keyframes kept

35 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
185
whisperx 185
chunks
29
from 185 cues
keyframes
22
kept of 35 captured
frames with text
22
420 lines read
chapters
12
from the source metadata
keyframe bytes
4.0 MB
word timings on 185 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:16 0s
stt done 2026-08-09 13:50 23s
chunk done 2026-08-09 13:50 0s
text_embed done 2026-08-10 19:41 1s
keyframe done 2026-08-09 13:50 2m 13s
ocr done 2026-08-09 13:53 9s
frame_embed done 2026-08-10 19:41 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:14 #3 done2 line(s)

    shot 3·sharpness 300.9

    1. AlEngineer0.99
    2. World's Fair0.97
  • 0:41 #4 done23 line(s)

    shot 4·sharpness 2624.7

    1. AlEngineer0.98
    2. Nu At A Glance0.98
    3. World's Fair0.98
    4. Leading digital bank in Latin America — and a launchpad for production Al.1.00
    5. 115M+1.00
    6. 15M1.00
    7. 135M1.00
    8. PRESENTED BY1.00
    9. Customers in Brazil1.00
    10. Customers in Mexico1.00
    11. Microsoft1.00
    12. customers across Latin America0.99
    13. NPS > 800.99
    14. 5B+1.00
    15. ≈ +4M added in Q1 2026· LatAm's #1 digital bank0.97
    16. Customer love1.00
    17. Quarterly Revenue (Q1'26)0.99
    18. Customer support approach: A hybrid model0.99
    19. Human Xpeers (fanatical customer care) + Al agents resolve issues together → fast, empathetic, correct.0.99
    20. Al handles routine intents end-to-end; humans focus on the hardest, long-tail cases.0.99
    21. NU0.89
    22. Engineering the future of Al0.99
    23. World's Fair0.99
  • 1:28 #5 done12 line(s)

    shot 5·sharpness 2069.9

    1. AlEngineer0.97
    2. World'sFair1.00
    3. THIS TALK IS ABOUT ONE THING ONLY...0.98
    4. If you generate your eval data in sim,0.99
    5. PRESENTED BY1.00
    6. Instead of waiting on prod,1.00
    7. Microsoft1.00
    8. You can ship 20× faster.1.00
    9. nU|0.80
    10. TRACK 3· JULY 2, 20260.95
    11. Al in Finance0.96
    12. World'sFair1.00
  • 1:47 #6 done18 line(s)

    shot 6·sharpness 2550.7

    1. AlEngineer0.98
    2. PROVEN IN PRODUCTION· FIVE DEPLOYMENTS0.97
    3. Results First: Closing gap to human agents0.99
    4. World'sFair1.00
    5. Average Al tNPS across five production agents — rising toward expert-human quality.0.98
    6. AltNPS ↑0.93
    7. Read our KDD paper0.97
    8. Expert human tNPS0.99
    9. +26 pp average tNPS gain0.98
    10. ≈10 pp to human1.00
    11. across five deployments1.00
    12. Launch1.00
    13. Now(V10)1.00
    14. TIME·AGENT VERSIONS→0.96
    15. Averaged across five online A/B tests — card delivery, debt, credit-limit, card management, product explainer. Gupta et al., KDD 2026, Table 3.0.99
    16. TRACK 3· JULY 2, 20260.96
    17. Alin Finance0.99
    18. World'sFair1.00
  • 2:57 #7 done13 line(s)

    shot 7·sharpness 2087.5

    1. AlEngineer0.97
    2. In This Talk0.96
    3. World'sFair1.00
    4. 010.99
    5. Eval data is the bottleneck0.97
    6. 021.00
    7. Why Simulated Eval Data Works As Well As Collected Data1.00
    8. 031.00
    9. Nu: 20x Faster Agent Workflows w/ Simulations0.99
    10. nU|0.86
    11. TRACK 3· JULY 2, 20260.93
    12. Alin Finance0.97
    13. World'sFair1.00
  • 3:27 #8 done14 line(s)

    shot 8·sharpness 2249.0

    1. AlEngineer0.98
    2. Evals are Critical for Building Good Agents1.00
    3. World'sFair1.00
    4. Evals = Metrics + Data1.00
    5. METRICS1.00
    6. DATA1.00
    7. LLM-as-a-judge +1.00
    8. Time consuming &0.96
    9. human alignment1.00
    10. expensive process1.00
    11. nU|0.87
    12. TRACK 3· JULY 2, 20260.94
    13. Alin Finance0.98
    14. World'sFair1.00
  • 3:36 #9 skipped

    shot 9·duplicate of #8

  • 4:41 #10 done41 line(s)

    shot 10·sharpness 1983.6

    1. AlEngineer0.99
    2. Eval Data for Agents is Very Hard1.00
    3. World's Fair0.99
    4. Structured ML data0.98
    5. RAG/Single-turn QA1.00
    6. Multi-turn agents1.00
    7. x10.80
    8. x20.99
    9. x31.00
    10. y1.00
    11. question1.00
    12. answer1.00
    13. I was charged twice this month.0.99
    14. 0.621.00
    15. 1.401.00
    16. 0.081.00
    17. Reset my PIN?0.97
    18. Settings → Security →0.97
    19. Let me take a look.0.99
    20. PIN1.00
    21. 0.311.00
    22. 0.971.00
    23. 0.551.00
    24. getCharges(user, 30d)0.99
    25. 0.880.94
    26. 0.121.00
    27. 0.741.00
    28. Refund window? Within 30 days1.00
    29. refund-agent·issueRefund($9.99)1.00
    30. 0.451.00
    31. 0.631.00
    32. 0.291.00
    33. Min balance?1.00
    34. None on checking1.00
    35. Done— duplicate refunded.0.97
    36. Card limit?0.99
    37. $5,000 default0.99
    38. NUI0.79
    39. TRACK 3· JULY 2, 20260.95
    40. Al in Finance1.00
    41. World'sFair1.00
  • 4:53 #11 skipped

    shot 11·duplicate of #10

  • 5:37 #12 done15 line(s)

    shot 12·sharpness 1954.2

    1. AlEngineer0.97
    2. How Teams Get Eval Data Today1.00
    3. World'sFair0.99
    4. MANUAL AUTHORING0.98
    5. PRODUCTION TRACES0.98
    6. • Hand-plan every state update and tool call0.99
    7. You're testing on real, live users1.00
    8. PRESENTED BY1.00
    9. • Synthetic state has to stay consistent across calls0.99
    10. • Can't parallelize or run experiments at scale0.99
    11. Microsoft1.00
    12. nU0.81
    13. TRACK 3· JULY 2, 20260.95
    14. Alin Finance0.97
    15. World'sFair1.00
  • 5:52 #13 done21 line(s)

    shot 13·sharpness 2074.8

    1. AlEngineer0.99
    2. A Typical Agent Release Cycle1.00
    3. World's Fair0.98
    4. Change agent harness1.00
    5. TODAY1.00
    6. a few hours0.97
    7. PRESENTED BY1.00
    8. Run offline evals on hand-curated data1.00
    9. TODAY1.00
    10. 2-3 days0.95
    11. Microsoft1.00
    12. A/B test in prod & monitor regressions0.99
    13. TODAY1.00
    14. 2-3 weeks0.91
    15. 1week1.00
    16. 2 weeks1.00
    17. 3 weeks1.00
    18. NU0.91
    19. TRACK 3· JULY 2, 20260.96
    20. Al in Finance0.99
    21. World's Fair0.99
  • 6:30 #14 done10 line(s)

    shot 14·sharpness 1690.4

    1. AlEngineer0.97
    2. World'sFair1.00
    3. Simulations short-circuit this timeline:0.99
    4. PRESENTED BY1.00
    5. From 3 weeks to <1 day.1.00
    6. Microsoft1.00
    7. nU|0.81
    8. TRACK 3· JULY 2, 20260.93
    9. Alin Finance0.98
    10. World'sFair1.00
  • 7:03 #15 skipped

    shot 15·duplicate of #13

  • 7:37 #16 done24 line(s)

    shot 16·sharpness 2709.2

    1. AlEngineer0.99
    2. Mechanically, What is a Simulation?1.00
    3. World's Fair0.97
    4. INPUT1.00
    5. 011.00
    6. Wrap your agent0.97
    7. 02 Describe personas & scenarios1.00
    8. Point at the endpoint, no code changes. Snowglobe groks0.99
    9. Who shows up, and what they're trying to get done in plain1.00
    10. the tools it calls for mocking.1.00
    11. language.1.00
    12. Snowglobe runs the sim1.00
    13. OUTPUT1.00
    14. 1,000s1.00
    15. pass / fail0.96
    16. Eval set0.94
    17. of multi-turn conversations0.99
    18. judge labels on every single turn1.00
    19. that drops into your existing pipeline1.00
    20. against your real agent1.00
    21. nU0.69
    22. TRACK 3· JULY 2, 20260.94
    23. Al in Finance0.99
    24. World'sFair1.00
  • 7:54 #17 skipped

    shot 17·duplicate of #16

  • 8:16 #18 skipped

    shot 18·duplicate of #16

  • 9:22 #19 done39 line(s)

    shot 19·sharpness 2320.1

    1. AlEngineer0.98
    2. What Does Simulated Data Look Like?0.99
    3. World'sFair1.00
    4. SIMULATED PERSONA1.00
    5. SIMULATED CONVERSATION1.00
    6. MS1.00
    7. Maria Souza1.00
    8. USER1.00
    9. need a new credit card.0.99
    10. AGENT1.00
    11. Happy to help — let me pull up your account.0.99
    12. Wants to order a new credit card. 34, designer,0.99
    13. L, getCustomer("Maria Souza") [mocked]0.95
    14. first-time credit customer.0.97
    15. AGENT1.00
    16. Ship it to Rua Augusta 1024, São Paulo?0.99
    17. Address1.00
    18. Rua Augusta 1024,1.00
    19. USER1.00
    20. yes.1.00
    21. São Paulo—SP0.98
    22. L, card-agent · issueCard(·.· 4821) [mocked]0.92
    23. Card1.00
    24. 5502……48210.96
    25. Tone1.00
    26. curt1.00
    27. AGENT1.00
    28. Done — your new card arrives in 5 business days.0.99
    29. Verbosity1.00
    30. one-line messages0.98
    31. USER1.00
    32. k.0.97
    33. grounding data1.00
    34. account data1.00
    35. tone & verbosity0.95
    36. NU0.93
    37. TRACK 3· JULY 2, 20260.95
    38. Al in Finance0.95
    39. World'sFair1.00
  • 9:58 #20 skipped

    shot 20·duplicate of #19

  • 10:36 #21 done34 line(s)

    shot 21·sharpness 2690.9

    1. AlEngineer0.98
    2. Self Improving Agents at Nu,1.00
    3. World'sFair1.00
    4. Powered by Snowglobe1.00
    5. On-demand eval data + good metrics1.00
    6. PRESENTED BY0.99
    7. is all you need for agents that improve themselves.0.99
    8. Microsoft1.00
    9. 010.99
    10. 021.00
    11. 031.00
    12. 041.00
    13. 051.00
    14. Ship &1.00
    15. Auto-eval1.00
    16. Generate sim1.00
    17. Optimize1.00
    18. Confirm1.00
    19. observe1.00
    20. Flag failing traces0.99
    21. Reproduce the1.00
    22. Tune the agent1.00
    23. Regression-test in1.00
    24. Watch prod signals0.99
    25. failures1.00
    26. harness1.00
    27. sim1.00
    28. A0.83
    29. RUNS ON ALIGNED EVAL METRICS0.99
    30. NU0.91
    31. TRACK 3· JULY 2, 20260.95
    32. Al in Finance0.96
    33. AlEngineer1.00
    34. World'sFair1.00
  • 11:09 #22 skipped

    shot 22·duplicate of #21

  • 11:25 #23 done25 line(s)

    shot 23·sharpness 3103.6

    1. AlEngineer0.99
    2. Simulation Tracks Real Production Data1.00
    3. World's Fair0.97
    4. ONLINE SIM QUALITY EVAL0.98
    5. OFFLINE SIM QUALITY EVAL0.99
    6. HUMANREVIEW1.00
    7. PRESENTED BY1.00
    8. 200 vs 5%0.94
    9. 0.80/0.740.99
    10. ~80%1.00
    11. Microsoft1.00
    12. Eval metric trends matched in0.99
    13. Spearman & Pearson correlation0.97
    14. Of domain expert labels1.00
    15. an online comparison of 2001.00
    16. on sim vs. real data. Strong1.00
    17. confirmed that sims give usable1.00
    18. sims vs. 5% of live traffic.0.99
    19. enough to rank treatments1.00
    20. greenfield insight.1.00
    21. NU|0.86
    22. TRACK 3· JULY 2, 20260.94
    23. Al in Finance0.98
    24. AlEngineer1.00
    25. World's Fair0.97

Transcript

185 cues· 2,843 words· 15,519 chars

  1. 0:12 Hi, everybody.
  2. 0:13 My name is Shreya.
  3. 0:14 I am the CEO of Snowglobe.
  4. 0:16 And we have with us Aman, who is a principal machine learning engineer at Nu.
  5. 0:21 And this talk is going to be about simulation maxing and how NuBank ships agents 20x faster using simulations.
  6. 0:32 Hey, everyone.
  7. 0:33 I'm Aman.
  8. 0:34 So let me talk about NewBank at a glance.
  9. 0:37 We are the leading digital bank in Latin America.
  10. 0:40 We have 135 million customers in Brazil, Mexico, and Colombia.
  11. 0:44 And we are launching in the US real soon.
  12. 0:47 Our quarterly revenue crossed $5 billion in Q1, 2026.
  13. 0:51 Our NPS customer level is very high.
  14. 0:54 And we are the perfect company for using AI agents for customer support, where human expires are fanatical customer care people and AI agents together solve customer issues in a fast, empathetic, and correct manner.
  15. 1:08 AI handles a lot of routines end to end.
  16. 1:11 Humans focus on the hardest and long-tail cases.
  17. 1:14 And together, we aim to delight our customers.
  18. 1:19 So as Shreya said, this talk is only about one thing, really.
  19. 1:24 If you generate your eval data in SIEM, instead of waiting on production data, you can ship agents 20x faster.
  20. 1:32 And we'll give you evidence for that.
  21. 1:34 So let's start with the results directly.
  22. 1:37 So this is the average of TNPS, which is a measure of customer satisfaction for five of our AI agents in production.
  23. 1:45 And at the beginning, they were not so great.
  24. 1:48 But now, with a few months
  25. 1:51 of work and a few quarters worth of effort, we've been able to massively increase the TNPS and customer love for our AI agents.
  26. 1:59 And many of them are approaching human quality.
  27. 2:01 And this data is a bit stale, as many of them are exceeding human quality.
  28. 2:05 So we are at the stage where we are actually able to show proof that this actually works in production.
  29. 2:11 Here is a QR code for a KDD paper in case you want to check it out.
  30. 2:14 It's going to be presented in Korea in August.
  31. 2:21 Awesome, so we opened with results and it's really about this journey of how do you implement the right systems for evaluation in order to be able to achieve those results.
  32. 2:31 So in this talk, we basically split it up into these three sections.
  33. 2:34 The first is why evals are so important and essential, but why they're also the bottleneck from being able to do a lot of high throughput experimentation to get the results that Aman was showing earlier.
  34. 2:44 And then the second part of this talk is about why simulated data
  35. 2:48 works as well as collected data and helps you circumvent a lot of this bottleneck that we're gonna talk about.
  36. 2:54 And then the third part is really digging deep into the systems and the findings that we had by implementing this framework at scale in NewBank.
  37. 3:05 Awesome, so evals, we know there's a whole talk track dedicated to this conference to evals.
  38. 3:10 Evals are absolutely critical for building good agents, and evals are really only about two things, right?
  39. 3:16 There's metrics and there's data.
  40. 3:19 Metrics, again, I hope you attended many of the amazing talks yesterday on the evals track, but metrics are really
  41. 3:25 while they're challenging, we have a playbook for how to build metrics that are really well aligned with the rubrics that we care about, right?
  42. 3:32 Which is you essentially use LLM as a judge style classifiers, and you align it with human judgment and getting human data, and you can iteratively build on it using auto optimization and auto prompt tuning techniques.
  43. 3:46 The thing that's a bottleneck and that still remains very challenging and unsolved is what is the data that you're actually computing these metrics on?
  44. 3:55 And that process is very time consuming and very expensive, specifically so for agents.
  45. 4:01 So once again,
  46. 4:02 People have been talking about, if you're around machine learning era circa 2018, it would be like ML work is 85% data work, right?
  47. 4:10 So data has always been challenging, but with agents, the level of sophistication that data requires is just so much more expensive.
  48. 4:17 So here's examples of what structured ML data look like,
  49. 4:21 And what even early era of AI data for chatbots or single turn QA, which was much more manageable and tractable, and you could still think of it as these structured rows.
  50. 4:32 But now for multi-turn agents, each data point is a trajectory with a lot of internal tool calls that all need state, et cetera.

Chapters

  1. 0:00 Opening with the results
  2. 0:39 Nubank at 135 million customers
  3. 2:34 Why evals matter most
  4. 2:47 Why simulated data works
  5. 3:49 What makes agent eval data hard
  6. 4:49 How teams get eval data today
  7. 6:44 Simulations in minutes, not weeks
  8. 7:35 Pointing Snowglobe at your agent
  9. 8:48 A grounded simulation: Maria orders a card
  10. 10:31 Ship, observe, simulate, repeat
  11. 11:34 How close simulations are to real
  12. 13:43 Testing models and variants

Open at this second