Videos KMR_RBoCa4M
SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe)
Scene timeline
35 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 185
- whisperx 185
- chunks
- 29
- from 185 cues
- keyframes
- 22
- kept of 35 captured
- frames with text
- 22
- 420 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 4.0 MB
- word timings on 185 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:16 | 0s |
stt |
done | — | 2026-08-09 13:50 | 23s |
chunk |
done | — | 2026-08-09 13:50 | 0s |
text_embed |
done | — | 2026-08-10 19:41 | 1s |
keyframe |
done | — | 2026-08-09 13:50 | 2m 13s |
ocr |
done | — | 2026-08-09 13:53 | 9s |
frame_embed |
done | — | 2026-08-10 19:41 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.97
-
- AlEngineer0.98
- Nu At A Glance0.98
- World's Fair0.98
- Leading digital bank in Latin America — and a launchpad for production Al.1.00
- 115M+1.00
- 15M1.00
- 135M1.00
- PRESENTED BY1.00
- Customers in Brazil1.00
- Customers in Mexico1.00
- Microsoft1.00
- customers across Latin America0.99
- NPS > 800.99
- 5B+1.00
- ≈ +4M added in Q1 2026· LatAm's #1 digital bank0.97
- Customer love1.00
- Quarterly Revenue (Q1'26)0.99
- Customer support approach: A hybrid model0.99
- Human Xpeers (fanatical customer care) + Al agents resolve issues together → fast, empathetic, correct.0.99
- Al handles routine intents end-to-end; humans focus on the hardest, long-tail cases.0.99
- NU0.89
- Engineering the future of Al0.99
- World's Fair0.99
-
- AlEngineer0.97
- World'sFair1.00
- THIS TALK IS ABOUT ONE THING ONLY...0.98
- If you generate your eval data in sim,0.99
- PRESENTED BY1.00
- Instead of waiting on prod,1.00
- Microsoft1.00
- You can ship 20× faster.1.00
- nU|0.80
- TRACK 3· JULY 2, 20260.95
- Al in Finance0.96
- World'sFair1.00
-
- AlEngineer0.98
- PROVEN IN PRODUCTION· FIVE DEPLOYMENTS0.97
- Results First: Closing gap to human agents0.99
- World'sFair1.00
- Average Al tNPS across five production agents — rising toward expert-human quality.0.98
- AltNPS ↑0.93
- Read our KDD paper0.97
- Expert human tNPS0.99
- +26 pp average tNPS gain0.98
- ≈10 pp to human1.00
- across five deployments1.00
- Launch1.00
- Now(V10)1.00
- TIME·AGENT VERSIONS→0.96
- Averaged across five online A/B tests — card delivery, debt, credit-limit, card management, product explainer. Gupta et al., KDD 2026, Table 3.0.99
- TRACK 3· JULY 2, 20260.96
- Alin Finance0.99
- World'sFair1.00
-
- AlEngineer0.97
- In This Talk0.96
- World'sFair1.00
- 010.99
- Eval data is the bottleneck0.97
- 021.00
- Why Simulated Eval Data Works As Well As Collected Data1.00
- 031.00
- Nu: 20x Faster Agent Workflows w/ Simulations0.99
- nU|0.86
- TRACK 3· JULY 2, 20260.93
- Alin Finance0.97
- World'sFair1.00
-
- AlEngineer0.98
- Evals are Critical for Building Good Agents1.00
- World'sFair1.00
- Evals = Metrics + Data1.00
- METRICS1.00
- DATA1.00
- LLM-as-a-judge +1.00
- Time consuming &0.96
- human alignment1.00
- expensive process1.00
- nU|0.87
- TRACK 3· JULY 2, 20260.94
- Alin Finance0.98
- World'sFair1.00
-
- AlEngineer0.99
- Eval Data for Agents is Very Hard1.00
- World's Fair0.99
- Structured ML data0.98
- RAG/Single-turn QA1.00
- Multi-turn agents1.00
- x10.80
- x20.99
- x31.00
- y1.00
- question1.00
- answer1.00
- I was charged twice this month.0.99
- 0.621.00
- 1.401.00
- 0.081.00
- Reset my PIN?0.97
- Settings → Security →0.97
- Let me take a look.0.99
- PIN1.00
- 0.311.00
- 0.971.00
- 0.551.00
- getCharges(user, 30d)0.99
- 0.880.94
- 0.121.00
- 0.741.00
- Refund window? Within 30 days1.00
- refund-agent·issueRefund($9.99)1.00
- 0.451.00
- 0.631.00
- 0.291.00
- Min balance?1.00
- None on checking1.00
- Done— duplicate refunded.0.97
- Card limit?0.99
- $5,000 default0.99
- NUI0.79
- TRACK 3· JULY 2, 20260.95
- Al in Finance1.00
- World'sFair1.00
-
- AlEngineer0.97
- How Teams Get Eval Data Today1.00
- World'sFair0.99
- MANUAL AUTHORING0.98
- PRODUCTION TRACES0.98
- • Hand-plan every state update and tool call0.99
- You're testing on real, live users1.00
- PRESENTED BY1.00
- • Synthetic state has to stay consistent across calls0.99
- • Can't parallelize or run experiments at scale0.99
- Microsoft1.00
- nU0.81
- TRACK 3· JULY 2, 20260.95
- Alin Finance0.97
- World'sFair1.00
-
- AlEngineer0.99
- A Typical Agent Release Cycle1.00
- World's Fair0.98
- Change agent harness1.00
- TODAY1.00
- a few hours0.97
- PRESENTED BY1.00
- Run offline evals on hand-curated data1.00
- TODAY1.00
- 2-3 days0.95
- Microsoft1.00
- A/B test in prod & monitor regressions0.99
- TODAY1.00
- 2-3 weeks0.91
- 1week1.00
- 2 weeks1.00
- 3 weeks1.00
- NU0.91
- TRACK 3· JULY 2, 20260.96
- Al in Finance0.99
- World's Fair0.99
-
- AlEngineer0.97
- World'sFair1.00
- Simulations short-circuit this timeline:0.99
- PRESENTED BY1.00
- From 3 weeks to <1 day.1.00
- Microsoft1.00
- nU|0.81
- TRACK 3· JULY 2, 20260.93
- Alin Finance0.98
- World'sFair1.00
-
- AlEngineer0.99
- Mechanically, What is a Simulation?1.00
- World's Fair0.97
- INPUT1.00
- 011.00
- Wrap your agent0.97
- 02 Describe personas & scenarios1.00
- Point at the endpoint, no code changes. Snowglobe groks0.99
- Who shows up, and what they're trying to get done in plain1.00
- the tools it calls for mocking.1.00
- language.1.00
- Snowglobe runs the sim1.00
- OUTPUT1.00
- 1,000s1.00
- pass / fail0.96
- Eval set0.94
- of multi-turn conversations0.99
- judge labels on every single turn1.00
- that drops into your existing pipeline1.00
- against your real agent1.00
- nU0.69
- TRACK 3· JULY 2, 20260.94
- Al in Finance0.99
- World'sFair1.00
-
- AlEngineer0.98
- What Does Simulated Data Look Like?0.99
- World'sFair1.00
- SIMULATED PERSONA1.00
- SIMULATED CONVERSATION1.00
- MS1.00
- Maria Souza1.00
- USER1.00
- need a new credit card.0.99
- AGENT1.00
- Happy to help — let me pull up your account.0.99
- Wants to order a new credit card. 34, designer,0.99
- L, getCustomer("Maria Souza") [mocked]0.95
- first-time credit customer.0.97
- AGENT1.00
- Ship it to Rua Augusta 1024, São Paulo?0.99
- Address1.00
- Rua Augusta 1024,1.00
- USER1.00
- yes.1.00
- São Paulo—SP0.98
- L, card-agent · issueCard(·.· 4821) [mocked]0.92
- Card1.00
- 5502……48210.96
- Tone1.00
- curt1.00
- AGENT1.00
- Done — your new card arrives in 5 business days.0.99
- Verbosity1.00
- one-line messages0.98
- USER1.00
- k.0.97
- grounding data1.00
- account data1.00
- tone & verbosity0.95
- NU0.93
- TRACK 3· JULY 2, 20260.95
- Al in Finance0.95
- World'sFair1.00
-
- AlEngineer0.98
- Self Improving Agents at Nu,1.00
- World'sFair1.00
- Powered by Snowglobe1.00
- On-demand eval data + good metrics1.00
- PRESENTED BY0.99
- is all you need for agents that improve themselves.0.99
- Microsoft1.00
- 010.99
- 021.00
- 031.00
- 041.00
- 051.00
- Ship &1.00
- Auto-eval1.00
- Generate sim1.00
- Optimize1.00
- Confirm1.00
- observe1.00
- Flag failing traces0.99
- Reproduce the1.00
- Tune the agent1.00
- Regression-test in1.00
- Watch prod signals0.99
- failures1.00
- harness1.00
- sim1.00
- A0.83
- RUNS ON ALIGNED EVAL METRICS0.99
- NU0.91
- TRACK 3· JULY 2, 20260.95
- Al in Finance0.96
- AlEngineer1.00
- World'sFair1.00
-
- AlEngineer0.99
- Simulation Tracks Real Production Data1.00
- World's Fair0.97
- ONLINE SIM QUALITY EVAL0.98
- OFFLINE SIM QUALITY EVAL0.99
- HUMANREVIEW1.00
- PRESENTED BY1.00
- 200 vs 5%0.94
- 0.80/0.740.99
- ~80%1.00
- Microsoft1.00
- Eval metric trends matched in0.99
- Spearman & Pearson correlation0.97
- Of domain expert labels1.00
- an online comparison of 2001.00
- on sim vs. real data. Strong1.00
- confirmed that sims give usable1.00
- sims vs. 5% of live traffic.0.99
- enough to rank treatments1.00
- greenfield insight.1.00
- NU|0.86
- TRACK 3· JULY 2, 20260.94
- Al in Finance0.98
- AlEngineer1.00
- World's Fair0.97
Transcript
185 cues· 2,843 words· 15,519 chars
- 0:12 Hi, everybody.
- 0:13 My name is Shreya.
- 0:14 I am the CEO of Snowglobe.
- 0:16 And we have with us Aman, who is a principal machine learning engineer at Nu.
- 0:21 And this talk is going to be about simulation maxing and how NuBank ships agents 20x faster using simulations.
- 0:32 Hey, everyone.
- 0:33 I'm Aman.
- 0:34 So let me talk about NewBank at a glance.
- 0:37 We are the leading digital bank in Latin America.
- 0:40 We have 135 million customers in Brazil, Mexico, and Colombia.
- 0:44 And we are launching in the US real soon.
- 0:47 Our quarterly revenue crossed $5 billion in Q1, 2026.
- 0:51 Our NPS customer level is very high.
- 0:54 And we are the perfect company for using AI agents for customer support, where human expires are fanatical customer care people and AI agents together solve customer issues in a fast, empathetic, and correct manner.
- 1:08 AI handles a lot of routines end to end.
- 1:11 Humans focus on the hardest and long-tail cases.
- 1:14 And together, we aim to delight our customers.
- 1:19 So as Shreya said, this talk is only about one thing, really.
- 1:24 If you generate your eval data in SIEM, instead of waiting on production data, you can ship agents 20x faster.
- 1:32 And we'll give you evidence for that.
- 1:34 So let's start with the results directly.
- 1:37 So this is the average of TNPS, which is a measure of customer satisfaction for five of our AI agents in production.
- 1:45 And at the beginning, they were not so great.
- 1:48 But now, with a few months
- 1:51 of work and a few quarters worth of effort, we've been able to massively increase the TNPS and customer love for our AI agents.
- 1:59 And many of them are approaching human quality.
- 2:01 And this data is a bit stale, as many of them are exceeding human quality.
- 2:05 So we are at the stage where we are actually able to show proof that this actually works in production.
- 2:11 Here is a QR code for a KDD paper in case you want to check it out.
- 2:14 It's going to be presented in Korea in August.
- 2:21 Awesome, so we opened with results and it's really about this journey of how do you implement the right systems for evaluation in order to be able to achieve those results.
- 2:31 So in this talk, we basically split it up into these three sections.
- 2:34 The first is why evals are so important and essential, but why they're also the bottleneck from being able to do a lot of high throughput experimentation to get the results that Aman was showing earlier.
- 2:44 And then the second part of this talk is about why simulated data
- 2:48 works as well as collected data and helps you circumvent a lot of this bottleneck that we're gonna talk about.
- 2:54 And then the third part is really digging deep into the systems and the findings that we had by implementing this framework at scale in NewBank.
- 3:05 Awesome, so evals, we know there's a whole talk track dedicated to this conference to evals.
- 3:10 Evals are absolutely critical for building good agents, and evals are really only about two things, right?
- 3:16 There's metrics and there's data.
- 3:19 Metrics, again, I hope you attended many of the amazing talks yesterday on the evals track, but metrics are really
- 3:25 while they're challenging, we have a playbook for how to build metrics that are really well aligned with the rubrics that we care about, right?
- 3:32 Which is you essentially use LLM as a judge style classifiers, and you align it with human judgment and getting human data, and you can iteratively build on it using auto optimization and auto prompt tuning techniques.
- 3:46 The thing that's a bottleneck and that still remains very challenging and unsolved is what is the data that you're actually computing these metrics on?
- 3:55 And that process is very time consuming and very expensive, specifically so for agents.
- 4:01 So once again,
- 4:02 People have been talking about, if you're around machine learning era circa 2018, it would be like ML work is 85% data work, right?
- 4:10 So data has always been challenging, but with agents, the level of sophistication that data requires is just so much more expensive.
- 4:17 So here's examples of what structured ML data look like,
- 4:21 And what even early era of AI data for chatbots or single turn QA, which was much more manageable and tractable, and you could still think of it as these structured rows.
- 4:32 But now for multi-turn agents, each data point is a trajectory with a lot of internal tool calls that all need state, et cetera.
loading
Chapters
- 0:00 Opening with the results
- 0:39 Nubank at 135 million customers
- 2:34 Why evals matter most
- 2:47 Why simulated data works
- 3:49 What makes agent eval data hard
- 4:49 How teams get eval data today
- 6:44 Simulations in minutes, not weeks
- 7:35 Pointing Snowglobe at your agent
- 8:48 A grounded simulation: Maria orders a card
- 10:31 Ship, observe, simulate, repeat
- 11:34 How close simulations are to real
- 13:43 Testing models and variants