read-only demo

Videos CLttOU7n6sI

Respect The Process - Andrew Dumit, Watershed Technology Inc.

index_state ready data_status ok

AI Engineer· published 2026-07-07· 0:16:43· en-US· indexed 2026-08-11 05:03

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:34, 1 of 1 keyframes kept
  2. Shot 1, 0:34 to 1:06, 1 of 1 keyframes kept
  3. Shot 2, 1:06 to 1:38, 1 of 1 keyframes kept
  4. Shot 3, 1:38 to 2:25, 1 of 1 keyframes kept
  5. Shot 4, 2:25 to 2:53, 1 of 1 keyframes kept
  6. Shot 5, 2:53 to 3:21, 1 of 1 keyframes kept
  7. Shot 6, 3:21 to 3:50, 0 of 1 keyframes kept
  8. Shot 7, 3:50 to 4:25, 1 of 1 keyframes kept
  9. Shot 8, 4:25 to 5:00, 0 of 1 keyframes kept
  10. Shot 9, 5:00 to 5:32, 1 of 1 keyframes kept
  11. Shot 10, 5:32 to 6:04, 0 of 1 keyframes kept
  12. Shot 11, 6:04 to 6:36, 0 of 1 keyframes kept
  13. Shot 12, 6:36 to 7:01, 1 of 1 keyframes kept
  14. Shot 13, 7:01 to 7:26, 1 of 1 keyframes kept
  15. Shot 14, 7:26 to 7:51, 0 of 1 keyframes kept
  16. Shot 15, 7:51 to 8:17, 1 of 1 keyframes kept
  17. Shot 16, 8:17 to 8:44, 0 of 1 keyframes kept
  18. Shot 17, 8:44 to 9:14, 1 of 1 keyframes kept
  19. Shot 18, 9:14 to 9:44, 0 of 1 keyframes kept
  20. Shot 19, 9:44 to 10:14, 0 of 1 keyframes kept
  21. Shot 20, 10:14 to 10:45, 1 of 1 keyframes kept
  22. Shot 21, 10:45 to 11:15, 0 of 1 keyframes kept
  23. Shot 22, 11:15 to 11:46, 0 of 1 keyframes kept
  24. Shot 23, 11:46 to 12:17, 1 of 1 keyframes kept
  25. Shot 24, 12:17 to 12:43, 1 of 1 keyframes kept
  26. Shot 25, 12:43 to 13:09, 0 of 1 keyframes kept
  27. Shot 26, 13:09 to 13:35, 0 of 1 keyframes kept
  28. Shot 27, 13:35 to 14:06, 1 of 1 keyframes kept
  29. Shot 28, 14:06 to 14:37, 0 of 1 keyframes kept
  30. Shot 29, 14:37 to 15:08, 0 of 1 keyframes kept
  31. Shot 30, 15:08 to 15:37, 1 of 1 keyframes kept
  32. Shot 31, 15:37 to 16:06, 0 of 1 keyframes kept
  33. Shot 32, 16:06 to 16:35, 0 of 1 keyframes kept
  34. Shot 33, 16:35 to 16:42, 1 of 1 keyframes kept

34 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
140
whisperx 140
chunks
30
from 140 cues
keyframes
18
kept of 34 captured
frames with text
18
351 lines read
chapters
0
from the source metadata
keyframe bytes
4.7 MB
word timings on 140 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:00 1m 06s
stt done 2026-08-11 05:01 19s
chunk done 2026-08-11 05:02 0s
text_embed done 2026-08-11 05:02 1s
keyframe done 2026-08-11 05:02 44s
ocr done 2026-08-11 05:02 8s
frame_embed done 2026-08-11 05:03 3s

Frames, and what the machine read

  • 0:20 #0 done4 line(s)

    shot 0·sharpness 720.3

    1. docs.google.com1.00
    2. Respect the Process1.00
    3. Trustworthy coding agents in a domain full of expert1.00
    4. judgement calls0.98
  • 0:53 #1 done13 line(s)

    shot 1·sharpness 2341.7

    1. docs.google.com1.00
    2. Sustainability is filled with0.98
    3. judgement calls0.99
    4. Questions like:0.99
    5. What are the emissions attributable to1.00
    6. 1 bottle of wine?0.98
    7. Which method should be used to1.00
    8. allocate the emissions to various0.99
    9. co-products?1.00
    10. There are many ways to get to a right answer1.00
    11. the wrong way and many right answers.1.00
    12. verify the process in addition to the0.99
    13. 21.00
  • 1:13 #2 done20 line(s)

    shot 2·sharpness 3256.0

    1. docs.google.com1.00
    2. 0.64
    3. 0.59
    4. Sustainability is filled with0.98
    5. judgement calls0.98
    6. Questions like:1.00
    7. 0.97 - 1.49 kgco2e0.96
    8. What are the emissions attributable to0.97
    9. 1 bottle of wine?0.99
    10. Which method should be used to1.00
    11. allocate the emissions to various1.00
    12. Six experts with the exact0.99
    13. co-products?1.00
    14. same data on the exact same1.00
    15. There are many ways to get to a right answer1.00
    16. bottle of wine0.99
    17. the wrong way and many right answers.1.00
    18. verify the process in addition to the1.00
    19. Scrucca et. al. (2020). "Uncertainty in LCA: An estimation of practitioner-related effects."0.99
    20. 21.00
  • 2:11 #3 done40 line(s)

    shot 3·sharpness 2170.9

    1. docs.google.com1.00
    2. 0.87
    3. Our task: Updating0.98
    4. complex graph data0.99
    5. electricity1.00
    6. Gas &0.98
    7. Softener1.00
    8. Cotton yarn1.00
    9. Sulphur dye0.96
    10. 0.40.98
    11. 0.31.00
    12. 211.00
    13. 0.11.00
    14. Help our users edit complex graphs representing1.00
    15. Zipper1.00
    16. Thread &0.98
    17. labels0.99
    18. Denim1.00
    19. Electricity1.00
    20. supply chains0.98
    21. 0.31.00
    22. 0.11.00
    23. 2.90.90
    24. 0.21.00
    25. Each graph is a DAG representing the flow of1.00
    26. materials and energy1.00
    27. The graph is comprised of 1,000s of nodes with rich1.00
    28. Packaging1.00
    29. Transport1.00
    30. assembly1.00
    31. Jeans1.00
    32. metadata describing the materials and processing at0.99
    33. 0.21.00
    34. 0.31.00
    35. 3.51.00
    36. each step0.95
    37. Dark wash0.96
    38. jeans1.00
    39. 4.01.00
    40. 31.00
  • 2:42 #4 done35 line(s)

    shot 4·sharpness 1486.1

    1. docs.google.com1.00
    2. 0.75
    3. Our agent worked1.00
    4. on one graph...0.98
    5. electricity1.00
    6. Gas &0.92
    7. Softener1.00
    8. Cotton yarn0.97
    9. Sulphur dye0.96
    10. 0.31.00
    11. 2.10.98
    12. 0.11.00
    13. When we first tried to solve1.00
    14. this task a year ago, it worked1.00
    15. Zipper1.00
    16. Thread &0.95
    17. labels1.00
    18. Denim1.00
    19. Electricity1.00
    20. well enough on a single graph0.98
    21. 0.31.00
    22. 2.91.00
    23. 0.20.94
    24. using custom built tools1.00
    25. Packaging0.97
    26. Transport1.00
    27. assembly1.00
    28. Jeans1.00
    29. 0.21.00
    30. 0.31.00
    31. 3.51.00
    32. Dark wash0.98
    33. jeans1.00
    34. 4.01.00
    35. 41.00
  • 3:13 #5 done13 line(s)

    shot 5·sharpness 1894.6

    1. docs.google.com1.00
    2. Our agent worked1.00
    3. on one graph...0.97
    4. When we first tried to solve1.00
    5. this task a year ago, it worked1.00
    6. Suiphur dye0.91
    7. well enough on a single graph1.00
    8. using custom built tools1.00
    9. Zipper0.87
    10. But absolutely broke when we1.00
    11. tried to scale it up to many1.00
    12. graphs1.00
    13. 41.00
  • 3:43 #6 skipped

    shot 6·duplicate of #5

  • 4:20 #7 done16 line(s)

    shot 7·sharpness 2332.1

    1. docs.google.com1.00
    2. So, we naturally swapped in a coding agent0.98
    3. Gives the agent the ability0.99
    4. The agent can explore and1.00
    5. to delight1.00
    6. edit efficiently0.99
    7. 40.73
    8. The agent can find clever ways to1.00
    9. Loops over graphs, scripts to unpack1.00
    10. solve underspecified problems, create1.00
    11. and summarize node content, same1.00
    12. fancy visualizations, and answer1.00
    13. pattern of data exploration as agentic0.99
    14. related questions1.00
    15. data science1.00
    16. 51.00
  • 4:49 #8 skipped

    shot 8·duplicate of #7

  • 5:19 #9 done9 line(s)

    shot 9·sharpness 1456.1

    1. docs.google.com1.00
    2. But, unconstrained code is scary0.98
    3. The agent will find0.99
    4. “creative” ways to solve0.97
    5. every problem1.00
    6. Writing Python when we expect0.99
    7. typescript, editing graph artifacts1.00
    8. directly without lineage1.00
    9. 61.00
  • 5:51 #10 skipped

    shot 10·duplicate of #9

  • 6:26 #11 skipped

    shot 11·duplicate of #9

  • 6:58 #12 done14 line(s)

    shot 12·sharpness 1148.7

    1. docs.google.com1.00
    2. The process ≠ answer story is not new0.98
    3. The Open Proof Corpus, Dekoninck and1.00
    4. Petrov et. al. 20260.97
    5. 100%1.00
    6. Correct Final Answer1.00
    7. Correct Proof1.00
    8. 75%1.00
    9. 50%1.00
    10. 25%1.00
    11. 0%1.00
    12. o30.97
    13. Gemini-Pro1.00
    14. 71.00
  • 7:23 #13 done31 line(s)

    shot 13·sharpness 5406.7

    1. docs.google.com1.00
    2. The process ≠ answer story is not new0.98
    3. The Open Proof Corpus, Dekoninck and1.00
    4. Petrov et. al. 20260.98
    5. "their reasoning often contains subtle logical0.97
    6. 100%1.00
    7. Correct Final Answer0.98
    8. errors masked by fluent language, posing1.00
    9. Correct Proof1.00
    10. significant risks for critical applications."0.99
    11. - Beyond Correctness, Zheng et. al. 20250.98
    12. 75%1.00
    13. 50%1.00
    14. “For every success story like the unit distance counterexample,0.98
    15. 25%1.00
    16. there are likely thousands of pages generated for each of0.99
    17. these problems, which have led nowhere."1.00
    18. - Thomas Bloom on ErdosProblems.com, 20260.99
    19. 0%1.00
    20. o30.93
    21. Gemini-Pro1.00
    22. “72% of reward hacking episodes include0.99
    23. "an LLM agent with access to unit tests0.98
    24. explicit chain-of-thought rationale,1.00
    25. may delete failing tests rather than fix0.99
    26. suggesting models often frame exploits0.99
    27. the underlying bug.1.00
    28. as legitimate problem-solving."1.00
    29. - ImpossibleBench, Zhong et. al. 20250.97
    30. - Reward Hacking Bench, Thaman 20260.99
    31. 71.00
  • 7:29 #14 skipped

    shot 14·duplicate of #13

  • 7:59 #15 done18 line(s)

    shot 15·sharpness 2062.0

    1. docs.google.com1.00
    2. 0.66
    3. 0.85
    4. Constrain the effects, not the expression1.00
    5. reject → retry0.97
    6. User1.00
    7. deterministic1.00
    8. commit as1.00
    9. valid·traceable1.00
    10. request1.00
    11. execution1.00
    12. typed objects1.00
    13. ·replayable0.97
    14. agent writes1.00
    15. typed SDK1.00
    16. free code0.99
    17. + lint1.00
    18. 81.00
  • 8:33 #16 skipped

    shot 16·duplicate of #15

  • 9:05 #17 done34 line(s)

    shot 17·sharpness 3167.5

    1. C0.76
    2. docs.google.com1.00
    3. Our SDK as the only door0.98
    4. import {0.99
    5. defineEditFunction, findNodesByNameExact, assert,0.99
    6. setRate, editNode, type EditProductionGraphState,1.00
    7. } from '@watershed/graphAgentApi';0.99
    8. export default defineEditFunction(0.99
    9. · Typescript SDK with all edit primitives the0.99
    10. 'cut_electricity_and_swap_steel',1.00
    11. async (state: EditProductionGraphState) => {0.99
    12. agent needs to make changes1.00
    13. const [mfg] = findNodesByNameExact(0.98
    14. state.graph.nodes, 'manufacturing');0.99
    15. • Enforces which fields are editable vs. which0.98
    16. const [electricity] = findNodesByNameExact(1.00
    17. state.graph.nodes, 'grid electricity');1.00
    18. are derived from other fields0.99
    19. const [steel] = findNodesByNameExact(0.99
    20. state.graph.nodes, 'steel');0.99
    21. • Guarantees that it emits objects we expect0.99
    22. assert(mfg && electricity && steel,0.99
    23. • Requires teaching the agent how to use it0.98
    24. 'expected one of each - fail loud');0.98
    25. // cut electricity 15%0.99
    26. state = setRate(1.00
    27. state, mfg.identifier, electricity.identifier, 0.85);0.98
    28. // swap steel → stainless steel0.97
    29. state = editNode(state, steel.identifier, {0.99
    30. material: 'stainless steel'});0.98
    31. return state; // new, traceable state0.99
    32. },0.71
    33. );0.97
    34. 91.00
  • 9:37 #18 skipped

    shot 18·duplicate of #17

  • 10:08 #19 skipped

    shot 19·duplicate of #17

  • 10:38 #20 done16 line(s)

    shot 20·sharpness 3362.4

    1. docs.google.com1.00
    2. Deterministic execution to guarantee process1.00
    3. Even with the typed SDK as the entry point, we0.99
    4. run-executor.ts1.00
    5. only guided the agent towards our desired end0.99
    6. Lint the agent code0.99
    7. state1.00
    8. The real guarantee comes from final script that1.00
    9. Detect conflicts1.00
    10. we orchestrate on agent completion0.99
    11. The completion script calls the agent1.00
    12. Run agent edited code0.99
    13. generated code, validates the results, and send0.99
    14. any errors back to the agent0.98
    15. Validate output artifacts1.00
    16. Create review artifact1.00
  • 11:12 #21 skipped

    shot 21·duplicate of #20

  • 11:31 #22 skipped

    shot 22·duplicate of #20

  • 11:55 #23 done26 line(s)

    shot 23·sharpness 3074.2

    1. docs.google.com1.00
    2. 0.83
    3. 0.78
    4. Deterministic execution to guarantee process1.00
    5. 2. Overall Emissions Delta by Edit Function0.99
    6. Even with the typed SDK as the entry point, we0.98
    7. only guided the agent towards our desired end1.00
    8. 5.931.00
    9. -2.651.00
    10. state1.00
    11. The real guarantee comes from final script that0.99
    12. Er1.00
    13. we orchestrate on agent completion1.00
    14. Grapl0.99
    15. -0.053081.00
    16. 3.231.00
    17. The completion script calls the agent1.00
    18. OR0.81
    19. 51.00
    20. generated code, validates the results, and send1.00
    21. any errors back to the agent0.98
    22. Final1.00
    23. Baseline1.00
    24. option_b_rebalance1.00
    25. option_a_rebalance1.00
    26. Create review artifact0.99

Transcript

140 cues· 3,270 words· 17,810 chars

  1. 0:00 Hi everyone, my name is Andrew Dumit and I work on AI engineering at Watershed, the sustainability AI platform.
  2. 0:06 At Watershed, I work on AI for building product carbon footprints, and in turn measuring the emissions associated with all the things that a company buys and sells.
  3. 0:13 Sustainability, the vertical that Watershed is in, is one with a ton of expert judgment calls spread throughout it, and this has made building and deploying agents both exciting and challenging.
  4. 0:22 In this talk, I'll go deep on one of the tasks within sustainability that we spent quite a bit of time working on, and share learnings from our work deploying coding agents on that task, and why doing so has required that we respect the process.
  5. 0:35 Just to give a bit of context on sustainability as a vertical, we need to answer questions like what are the emissions attributable to one bottle of wine or which method should be used to allocate the emissions to various co-products when there are multiple sellable co-products that come from the same industrial process.
  6. 0:50 In these cases, there are many ways to get the right answer the wrong way and there are also many right answers that experts will disagree on.
  7. 0:58 And so you have to verify the process in addition to the answer because the answer is really only justified insofar as the process that produced that answer is correct.
  8. 1:08 One example of this comes from a study in 2020 where six experts were given the exact same data on the exact same bottle of wine.
  9. 1:17 And despite having access to the exact same things, they came to answers that varied by up to 50%.
  10. 1:24 Their expert judgment were all correct in a sense, but purely validating the answer here is not sufficient to know that the answer produced by a system trying to mimic those experts is itself correct.
  11. 1:39 So with that context on sustainability broadly and what kind of things that we need to answer, let's zoom into a specific task that we'll be talking about for the remainder of the talk.
  12. 1:47 That task is to help users edit complex graphs.
  13. 1:50 Each of these graphs represents the supply chain of a single product.
  14. 1:54 For example, on the right here, you can see the graph represents dark wash jeans.
  15. 1:59 Upstream of that is the assembly of that genes from the denim thread labels and zipper, along with the energy transportation and packaging that's required to move that through the supply chain and create and run those industrial processes.
  16. 2:14 Each graph here represents the entire flow of all of those things moving through the supply chain and is comprised of thousands of nodes, each with rich metadata describing the materials and processing at each step.
  17. 2:26 And so when we first tried to solve this problem a little over a year ago, it worked decently well on one graph.
  18. 2:31 It did have some problems, but we gave our React agent highly specified tools for exploring and interacting with the graph via function calls.
  19. 2:39 And even while it worked, a few problems were already apparent.
  20. 2:43 It lacked consistency where it could struggle to explore sufficiently.
  21. 2:47 And those tool calls on a single graph took a lot of context as it read deeply across these nodes with lots of metadata.
  22. 2:53 But then when we tried to scale it up to mini graphs or frankly, even just a few graphs, it absolutely broke.
  23. 3:00 The lack of consistency was greatly magnified with the agent taking one approach in one graph, a different approach in a second graph and completely forgetting about the third graph or even to handle it at all.
  24. 3:10 And exploration itself also became a huge bottleneck because it took many tool calls to explore and operate across these graphs at scale.
  25. 3:18 And it also worsened the context problem because these tool calls really gobbled up context, both on the edit side where it actually needed to make some changes and on the exploration side where it needed to figure out what changes to make.
  26. 3:29 And worse, the agent then really started to hallucinate different parts of the schema as those contexts got eaten.
  27. 3:35 And despite those specialized tools, this led to retries and ultimately errors.
  28. 3:40 In data terms, the task now comprised tens or hundreds of graphs and tens to hundreds of thousands of nodes that the agent had to operate over.
  29. 3:48 So time passed.
  30. 3:50 And coding agents got way better, and we thought, these will work great.
  31. 3:54 So naturally, we swapped them in.
  32. 3:56 And swapping in a coding agent gave us three really important outcomes that we were missing before.
  33. 4:01 First, it gave the agent the ability to delight users.
  34. 4:04 The agent could start to find clever ways to solve these underspecified and not fully formed problems.
  35. 4:11 It could also create fancy visualizations on the fly by writing code to do that.
  36. 4:14 And it could even answer related questions, because it had the full power of the coding agent in this environment.
  37. 4:20 And from our end, we saw the agent could explore and now edit way more efficiently than before.
  38. 4:26 It could write loops over graphs and nodes.
  39. 4:28 It could write scripts to unpack and summarize the node content underneath it all, basically following the same pattern of data exploration that goes on in kind of agentic data science workflows.
  40. 4:40 And beyond that, it also gave it the flexibility and power to do stuff outside what we even designed it for, which was really exciting.
  41. 4:45 And we're to this day consistently finding entirely new use cases via new questions users ask and new things that they want to try with this agent.
  42. 4:55 We were really excited, we put it out into the world, we started to write a bunch of evals for it, and we quickly learned that unconstrained code is quite scary.
  43. 5:04 If you're watching this, you've probably had that experience with cloud code, where it's gone a bit haywire towards achieving whatever goal you set it out to do.
  44. 5:11 It reached for something you didn't think it could or should reach for, or even had access or found a way to access something you didn't think it could.
  45. 5:19 For us, this looks like the agent will find creative ways to pretty much solve any problem.
  46. 5:24 In our case, we saw it write Python when we expected TypeScript and instructed it to write TypeScript because it found Python on the virtual machine that we had given it.
  47. 5:31 Or it would directly edit these graph artifacts underlying the data without leaving any lineage behind by just directly modifying parameters or data rather than actually writing code to effectively do that.
  48. 5:45 Second, the agent actually started to gaslight users sometimes, saying it had made edits when it hadn't.
  49. 5:50 It would write code that it thought was going to have the effect that it wanted to, and then it said, I'm done, everything's working, it's exactly what you said, and the edits were not actually made.
  50. 6:03 And so it kind of lied to users, or at least led them to believe that the edits were done as they expected it to be done.

Open at this second