read-only demo

Videos HEFSExa0xl0

Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs

index_state ready data_status ok

AI Engineer· published 2026-07-08· 0:19:08· en-US· indexed 2026-08-11 04:33

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 0:55, 1 of 1 keyframes kept
  6. Shot 5, 0:55 to 1:23, 1 of 1 keyframes kept
  7. Shot 6, 1:23 to 1:51, 1 of 1 keyframes kept
  8. Shot 7, 1:51 to 2:02, 1 of 1 keyframes kept
  9. Shot 8, 2:02 to 2:18, 1 of 1 keyframes kept
  10. Shot 9, 2:18 to 2:46, 0 of 1 keyframes kept
  11. Shot 10, 2:46 to 3:17, 0 of 1 keyframes kept
  12. Shot 11, 3:17 to 3:49, 1 of 1 keyframes kept
  13. Shot 12, 3:49 to 4:20, 0 of 1 keyframes kept
  14. Shot 13, 4:20 to 4:56, 1 of 1 keyframes kept
  15. Shot 14, 4:56 to 5:33, 0 of 1 keyframes kept
  16. Shot 15, 5:33 to 5:49, 1 of 1 keyframes kept
  17. Shot 16, 5:49 to 6:19, 1 of 1 keyframes kept
  18. Shot 17, 6:19 to 6:45, 1 of 1 keyframes kept
  19. Shot 18, 6:45 to 7:11, 0 of 1 keyframes kept
  20. Shot 19, 7:11 to 7:37, 0 of 1 keyframes kept
  21. Shot 20, 7:37 to 8:20, 1 of 1 keyframes kept
  22. Shot 21, 8:20 to 8:53, 0 of 1 keyframes kept
  23. Shot 22, 8:53 to 9:26, 0 of 1 keyframes kept
  24. Shot 23, 9:26 to 9:54, 1 of 1 keyframes kept
  25. Shot 24, 9:54 to 10:23, 0 of 1 keyframes kept
  26. Shot 25, 10:23 to 10:51, 1 of 1 keyframes kept
  27. Shot 26, 10:51 to 11:19, 0 of 1 keyframes kept
  28. Shot 27, 11:19 to 11:48, 1 of 1 keyframes kept
  29. Shot 28, 11:48 to 12:16, 1 of 1 keyframes kept
  30. Shot 29, 12:16 to 12:47, 0 of 1 keyframes kept
  31. Shot 30, 12:47 to 13:27, 1 of 1 keyframes kept
  32. Shot 31, 13:27 to 14:03, 1 of 1 keyframes kept
  33. Shot 32, 14:03 to 14:38, 0 of 1 keyframes kept
  34. Shot 33, 14:38 to 15:03, 1 of 1 keyframes kept
  35. Shot 34, 15:03 to 15:28, 1 of 1 keyframes kept
  36. Shot 35, 15:28 to 16:06, 0 of 1 keyframes kept
  37. Shot 36, 16:06 to 16:51, 0 of 1 keyframes kept
  38. Shot 37, 16:51 to 17:33, 0 of 1 keyframes kept
  39. Shot 38, 17:33 to 18:07, 1 of 1 keyframes kept
  40. Shot 39, 18:07 to 18:42, 1 of 1 keyframes kept
  41. Shot 40, 18:42 to 18:46, 1 of 1 keyframes kept
  42. Shot 41, 18:46 to 18:53, 1 of 1 keyframes kept
  43. Shot 42, 18:53 to 19:08, 1 of 1 keyframes kept

43 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
154
whisperx 154
chunks
33
from 154 cues
keyframes
28
kept of 43 captured
frames with text
28
407 lines read
chapters
0
from the source metadata
keyframe bytes
4.6 MB
word timings on 154 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 04:30 1m 06s
stt done 2026-08-11 04:31 23s
chunk done 2026-08-11 04:31 0s
text_embed done 2026-08-11 04:31 6s
keyframe done 2026-08-11 04:31 1m 27s
ocr done 2026-08-11 04:33 11s
frame_embed done 2026-08-11 04:33 5s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 827.9

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:22 #3 done11 line(s)

    shot 3·sharpness 460.6

    1. -hing Coding Age0.97
    2. ster Spreadshe0.98
    3. AlEngineer0.95
    4. Nuno Campos0.98
    5. EUROPE1.00
    6. Witan Labs, CTO & Co-Founder0.98
    7. Previously LangChain, LangGraph1.00
    8. AIE London - April 20260.96
    9. EU/ACC0.98
    10. AlEngineer1.00
    11. EUROPE1.00
  • 0:47 #4 done12 line(s)

    shot 4·sharpness 1791.6

    1. 50%→92%1.00
    2. — 4 months, multiple architectures, and many dead ends0.98
    3. — What mattered most: replacing 15 discrete tools with one REPL0.99
    4. 0.53
    5. AIE1.00
    6. 1.00
    7. 1.00
    8. 21.00
    9. AlEngineer0.96
    10. AlEngineer1.00
    11. 20260.98
    12. EUROPE1.00
  • 1:20 #5 done19 line(s)

    shot 5·sharpness 2777.4

    1. The problem1.00
    2. A human sees: revenue table, assumptions,1.00
    3. P&L summary1.00
    4. 1.00
    5. 1.00
    6. An LLM sees: 10,000 cell values,0.99
    7. 0.99
    8. AIE1.00
    9. 1.00
    10. =SUMPRODUCT((B$3:B$50="NEW")*(0.99
    11. 1.00
    12. 1.00
    13. 0.99
    14. 1.00
    15. 1.00
    16. G$3:G$50)),formatting metadata0.99
    17. AlEngineer0.96
    18. AlEngineer1.00
    19. EUROPE1.00
  • 1:37 #6 done14 line(s)

    shot 6·sharpness 2787.1

    1. The problem1.00
    2. A human sees: revenue table, assumptions,1.00
    3. P&L summary1.00
    4. 1.00
    5. An LLM sees: 10,000 cell values,0.97
    6. AIE1.00
    7. 1.00
    8. =SUMPRODUCT((B$3:B$50="NEW")*(0.99
    9. 1.00
    10. 1.00
    11. 1.00
    12. G$3:G$50)),formatting metadata0.99
    13. Engineering the future of Al0.98
    14. AlEngineer1.00
  • 1:58 #7 done14 line(s)

    shot 7·sharpness 2955.5

    1. One dead end0.98
    2. Three specialized agents:1.00
    3. Block discovery — identifies workbook structure0.98
    4. *★0.68
    5. AIE1.00
    6. 1.00
    7. 2. Edit agent — 5-step process: disambiguate, define end state, plan,0.99
    8. 1.00
    9. execute,verify1.00
    10. 3. Question agent— answers questions0.98
    11. Key finding: Rigid architectures don't win1.00
    12. Engineering the future of Al1.00
    13. AlEngineer1.00
    14. 20260.95
  • 2:15 #8 done12 line(s)

    shot 8·sharpness 442.3

    1. end1.00
    2. Throtialized agents:0.94
    3. AlEngineer1.00
    4. Block discovery — identifies workboo0.97
    5. EUROPE1.00
    6. 2. Edit agent – 5-step process: disambig0.97
    7. execute,verify1.00
    8. 3. Question agent- answers questions0.98
    9. EUJACO0.98
    10. Key finding: Rigid architectures don't win1.00
    11. AlEngineer1.00
    12. EUROPE1.00
  • 2:24 #9 skipped

    shot 9·duplicate of #7

  • 2:49 #10 skipped

    shot 10·duplicate of #4

  • 3:42 #11 done22 line(s)

    shot 11·sharpness 3348.4

    1. More dead ends1.00
    2. Representation1.00
    3. Why it failed0.99
    4. TSV views0.99
    5. Lost formatting and structural context0.99
    6. 1.00
    7. AIE0.99
    8. 0.97
    9. SQL views0.99
    10. Flat tables didn't capture visual structure1.00
    11. 1.00
    12. HTML tables0.99
    13. Too verbose, consumed too many tokens1.00
    14. DOT graphs1.00
    15. Useful for analysis, not for interaction1.00
    16. XML/XSLT1.00
    17. LLM struggled with the syntax1.00
    18. None of them worked as a general-purpose representation – but two informed what came next.0.99
    19. 51.00
    20. Engineering the future of Al0.99
    21. AlEngineer0.98
    22. 20261.00
  • 3:53 #12 skipped

    shot 12·duplicate of #11

  • 4:28 #13 done18 line(s)

    shot 13·sharpness 2144.5

    1. November 23rd: The REPL1.00
    2. We replaced all 15 tools with a persistent Node.js REPL and a spreadsheet0.99
    3. API.0.99
    4. *★★0.53
    5. 1.00
    6. AIE1.00
    7. Agent Loop --(JavaScript code)--> Node.js REPL --(JSON-RPC)--> Spreadsheet engine1.00
    8. 1.00
    9. 1.00
    10. ★★0.81
    11. Variables persist1.00
    12. Workbook state0.97
    13. across calls1.00
    14. persists1.00
    15. 61.00
    16. Engineering the future of Al1.00
    17. AlEngineer0.99
    18. 20261.00
  • 5:11 #14 skipped

    shot 14·duplicate of #4

  • 5:47 #15 done21 line(s)

    shot 15·sharpness 2324.2

    1. Before vs. After0.99
    2. Before (10-15 tool calls):1.00
    3. Tool call 1: list_sheets()0.99
    4. -> ["Summary", "Data", "Inputs"]0.98
    5. *★★0.51
    6. Tool call 2: read_range("Summary!A1:D10") -> cell values1.00
    7. AIE1.00
    8. 1.00
    9. Tool call 3: find_celis("Revenue")-> [matches]0.98
    10. 1.00
    11. Tool call 4: read_cell("Summary!C10") -> "$1,234,567"1.00
    12. 1.00
    13. ★★0.94
    14. After (I tool call):0.95
    15. const sheets = await xlsx.listSheets(wb);0.99
    16. const summary = await xlsx.readRangeTsv(wb, ${sheets[0].name}!A1:D10`);0.99
    17. const revenue = await xlsx.findCells(wb, "Revenue", { context: 1 });1.00
    18. console.log(sheets, summary, revenue);1.00
    19. Engineering the future of Al1.00
    20. AlEngineer1.00
    21. 20260.93
  • 6:10 #16 done15 line(s)

    shot 16·sharpness 385.0

    1. levs. REPL0.96
    2. Cod1.00
    3. a the agent writes scripts instead of ma0.98
    4. AlEngineer1.00
    5. naturally. 50-line scripts are common.0.99
    6. EUROPE1.00
    7. REPL = code mode + persistent state - variables s0.98
    8. scripts, reasons between them, builds understandin1.00
    9. BUJACC0.89
    10. console.log(revenue);1.00
    11. revenue =0.99
    12. xlsx.findCells(wb,1.00
    13. Result: accuracy gain on harder tasks - the agent ca0.99
    14. AlEngineer0.99
    15. EUROPE1.00
  • 6:27 #17 done23 line(s)

    shot 17·sharpness 3630.4

    1. Code mode vs. REPL0.98
    2. Code mode — the agent writes scripts instead of making tool calls. Operations compose0.99
    3. naturally. 50-line scripts are common.0.99
    4. 1.00
    5. ***0.74
    6. AIE1.00
    7. 1.00
    8. 0.96
    9. REPL = code mode + persistent state — variables survive across calls. The agent writes shorter0.98
    10. 1.00
    11. scripts, reasons between them, builds understanding incrementally.1.00
    12. 1.00
    13. ★★0.88
    14. // Code mode: one long script, print everything at the end0.99
    15. // REPL: explore, reason, continue0.99
    16. const revenue = await xlsx.findCells(wb, "Revenue", { context: 1 });0.98
    17. console.log(revenue);1.00
    18. // → agent reasons about the output, then writes the next script0.98
    19. Result: accuracy gain on harder tasks — the agent can course-correct mid-exploration.0.99
    20. 81.00
    21. Engineering the future of Al1.00
    22. AlEngineer0.99
    23. 20260.99
  • 7:08 #18 skipped

    shot 18·duplicate of #17

  • 7:24 #19 skipped

    shot 19·duplicate of #4

  • 7:55 #20 done15 line(s)

    shot 20·sharpness 390.6

    1. levs. REPL0.93
    2. Cod1.00
    3. the agent writes scripts instead of ma0.98
    4. AlEngineer1.00
    5. naturally. 50-line scripts are common.0.99
    6. EUROPE1.00
    7. REPL = code mode + persistent state - variables s0.97
    8. scripts, reasons between them, builds understandin1.00
    9. EU/ACC0.94
    10. console.log(revenue);1.00
    11. revenue =0.98
    12. xlsx.findCells(wb,1.00
    13. Result: accuracy gain on harder tasks - the agent ca0.98
    14. AlEngineer1.00
    15. EUROPE0.96
  • 8:24 #21 skipped

    shot 21·duplicate of #4

  • 8:57 #22 skipped

    shot 22·duplicate of #11

  • 9:40 #23 done20 line(s)

    shot 23·sharpness 1832.1

    1. The verification loop1.00
    2. The formula engine and renderer close a loop the agent1.00
    3. uses to check its own work.0.99
    4. Write1.00
    5. This held across three successive frontier model1.00
    6. *★★0.54
    7. releases. Each new model used the same loop more1.00
    8. AIE1.00
    9. effectively.1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. Calculate1.00
    15. Render1.00
    16. 101.00
    17. AlEngineer0.97
    18. AlEngineer1.00
    19. 20261.00
    20. EUROPE1.00

Transcript

154 cues· 2,761 words· 14,713 chars

  1. 0:15 Cool.
  2. 0:16 Hi, everyone.
  3. 0:17 My name is Nunu and I want to talk to you about how we spent the last four months teaching coding agents to master spreadsheets.
  4. 0:27 So essentially our goal was to get coding agents to be as good at spreadsheets as they are at Python, JavaScript, or whatever your favorite language is.
  5. 0:38 We started at around 50% accuracy on a financial analysis benchmark and got to 92%.
  6. 0:46 So I'll chat about what actually moved the needle and what didn't and the dead ends and yeah, let's see.
  7. 0:57 So spreadsheets are a little bit harder for AI than you might think at first.
  8. 1:03 If you think about how you would, if you open Excel, how you'd find your way in an Excel file that you don't know, it's actually a very visual thing and you just instantly see the structure
  9. 1:16 There's a revenue table here, assumptions in there, a chart in there, and it just feels intuitive and you don't even think about it.
  10. 1:24 And an LLM doesn't really see any of this.
  11. 1:29 If you ask it what's the revenue, then it has to figure out which revenue do you mean.
  12. 1:35 net revenue, gross revenue, revenue for this, revenue for that, which quarter, which year, and then is the number found like an actual input?
  13. 1:46 Is it a formula?
  14. 1:48 So it's actually a deceptively hard task.
  15. 1:53 One thing we tried close to the beginning was to split the work into three agents.
  16. 2:00 So the kind of the central one was the edit agent that had like a five step process that you'd define the end state, you'd do a plan, you'd execute, you'd verify, all the things you're supposed to do.
  17. 2:14 And this kind of changed the kind of errors we got.
  18. 2:18 Without it, the agent would just make mistakes while actually building a financial model or something.
  19. 2:24 And with this, it would maybe make those mistakes while planning, which was a lot easier to rectify.
  20. 2:31 But this architecture in the end was too rigid.
  21. 2:34 because discovery ran once up front, and then you couldn't revisit it, and the context wouldn't flow between the different agents, so it just turned out to be one dead end.
  22. 2:47 Then some more dead ends.
  23. 2:50 We, I think, ended up probably trying every conceivable way of representing a spreadsheet to an LLM.
  24. 2:58 None really worked as a standalone representation, but two turned out to be useful as methods inside the REPL that we ended up creating.
  25. 3:10 But they all had something going for them in theory, and that's why we tried it.
  26. 3:13 So SQL has obviously been around for decades, so super popular in LLM training data, so agents are really good at it, supposed to be a great way to deal with structured data, but it turns out that it doesn't quite work for this.
  27. 3:29 XML is how Excel files are represented on disks, so maybe that was a good idea.
  28. 3:33 It wasn't.
  29. 3:36 And many others.
  30. 3:38 In the end, we did get two useful things out of this.
  31. 3:42 One was the concept of having these CSV or TSV views of part of a spreadsheet.
  32. 3:49 This turned out to not be that great as the only way to interact with a spreadsheet, but as one piece of the larger solution, it turns out to be used very, very often.
  33. 4:03 And HTML was also a step in the right direction as it introduced the idea of a layout and formatting and so on.
  34. 4:09 So that ended up resulting in us building a rendering engine to let the agent see what the rendered spreadsheet looked like as an image.
  35. 4:22 And then eventually we hit on what was probably ended up being the biggest breakthrough, which was to replace the many tools that we accumulated over time.
  36. 4:34 I think at that time we had around 15 tools.
  37. 4:37 with a single tool, which was a Node.js REPL.
  38. 4:43 And so all the 15 tools that we had to start just became different JavaScript functions that the agent could combine in this one REPL call.
  39. 4:54 And why JavaScript?
  40. 4:58 We needed a scripting language that's easy for
  41. 5:01 to sandbox, it's easy, LLMs are very familiar with it, and Python would probably work equally well, we just went with JavaScript.
  42. 5:11 But the actual implementation of the code that deals with the spreadsheet is actually in a completely different language in C sharp.
  43. 5:20 And that's kind of the advantage of this architecture.
  44. 5:23 You just use the scripting language for what it's good at, which is letting the agents interact with it and use the right language to then deal with the actual files.
  45. 5:34 And what it looked like before and after.
  46. 5:36 So before it would be you'd have 10 or 15 tool calls usually for an agent to explore a spreadsheet and get to an answer.
  47. 5:46 And this would actually very often end up timing out and taking a long time because it was just doing things sequentially.
  48. 5:53 Even parallel tool calling didn't really help because you couldn't combine the results in any way.
  49. 5:57 And after, the agent would just combine the different things it wanted to do in a single tool call and get all the results at the same time.
  50. 6:09 So some of you will be familiar with the idea of code mode.

Open at this second