Videos 3ZMUiFaQ3qg
Verifiable Environments for AI in Biology — Kenny Workman, LatchBio
Scene timeline
63 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 206
- whisperx 206
- chunks
- 30
- from 206 cues
- keyframes
- 27
- kept of 63 captured
- frames with text
- 26
- 472 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 7.9 MB
- word timings on 206 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:38 | 0s |
stt |
done | — | 2026-08-09 02:38 | 19s |
chunk |
done | — | 2026-08-09 02:39 | 0s |
text_embed |
done | — | 2026-08-10 19:37 | 0s |
keyframe |
done | — | 2026-08-09 02:39 | 2m 26s |
ocr |
done | — | 2026-08-09 02:41 | 11s |
frame_embed |
done | — | 2026-08-10 19:37 | 5s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.99
-
- AlEngineer0.98
- World'sFair1.00
- PRESENTED BY1.00
- Benchmarking Agents in Biology1.00
- Microsoft1.00
- Kenny Workman, LatchBio1.00
- World's Fair0.99
- Engineering the future of Al0.99
-
- AlEngineer0.98
- World'sFair1.00
- New measurement techniques drive data generation in biology1.00
- Growth of public biological sequence data archives (proxy for experimental dats nonaratinn0.97
- Single cells encapsulated1.00
- - GenBank (assembled/annotated)0.99
- - SRA/ENA (raw reads)0.99
- PRESENTED BY1.00
- 3D structure0.97
- Biomarker0.96
- formation1.00
- Target binding1.00
- Microsoft1.00
- Droplets0.99
- Target cel0.98
- Totce sas ated0.63
- 10150.80
- Aptamer sequence1.00
- Functional aptarner0.96
- Aptamer binding to targets0.98
- 10{140.84
- 2022.060.92
- 2023-060.99
- 2024111.00
- 10130.82
- Tissue Section0.96
- ONA Synthesis0.79
- Spatial0.99
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining0.97
-
- AlEngineer0.97
- World'sFair1.00
- “Experimental data is getting really big” (by the numbers):0.98
- Single Cell. 10X Flex; ~1M cells/2.5B reads per library; 2-6TB per run0.99
- Spatial. Vizgen's MERFISH Ultra; 7TB of raw image data per run0.99
- Proteomics. "Deep" LC-MS; 24 fractions yields ~300GB0.97
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining &Midtraining0.99
-
- AlEngineer0.99
- We (LatchBio) got started building data tools for biotech/pharma1.00
- World'sFair1.00
- 1. Data infrastructure. Storage + transformation of file data from large1.00
- experiments1.00
- 2. Agents. LLMs in loops that steer analysis with natural language0.99
- PIXELGEN1.00
- TECHNOLOGIES1.00
- Data1.00
- Metadata1.00
- Tracking1.00
- Workflow0.99
- axtflow0.87
- vizgen1.00
- Data Silos...1.00
- Disconnected1.00
- Stagnant1.00
- Hurdles to1.00
- TakaRa0.92
- Samples...1.00
- Workflows...0.99
- Analysis...1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.93
- Posttraining &Midtraining0.98
-
- AlEngineer0.97
- Last summer (~Sonnet4.5), agent prototypes started to work in bio0.99
- World'sFair1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.99
-
- AlEngineer0.98
- Last summer (~Sonnet4.5), agent prototypes started to work in bio1.00
- World's Fair0.98
- 10X Visium.0.99
- Submit1.00
- Powered by Claude 4.5 Opus0.99
- Built for0.95
- 101.00
- ential chromatin0.94
- Find cluster-specific marker genes1.00
- stomach cancer1.00
- for Alzheimer's mouse brain from1.00
- Characterize spatial cell types in1.00
- nics DBit-seq1.00
- 10X Xenium1.00
- colon cancer from 10X Visium1.00
- Pause1.00
- LatchBio0.97
- 0.02 / 0:350.93
- World's Fair0.99
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining0.99
-
- AlEngineer0.99
- Last summer (~Sonnet4.5), agent prototypes started to work in bio1.00
- World'sFair1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.94
- Posttraining &Midtraining0.99
-
- AlEngineer0.97
- Last summer (~Sonnet4.5), agent prototypes started to work in bio0.99
- World'sFair1.00
- Summary1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining0.99
-
- AlEngineer0.96
- World'sFair0.98
- Despite early promise, frontier models cannot yet be trusted to do real work1.00
- Missing a capability between “knowing biology" and “writing code":0.98
- PRESENTED BY0.99
- extracting scientific insight from messy, real-world data.1.00
- Microsoft1.00
- Struggle with blend of programming, data analysis, and domain1.00
- reasoning.1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1,20260.97
- Posttraining &Midtraining0.98
-
- AlEngineer0.99
- World'sFair1.00
- Spatial biology was a good place to start building agents0.99
- PRESENTED BY1.00
- Microsoft1.00
- 0M37640.80
- Technical greenfield. Many customers.1.00
- Also just a beautiful example of 'measurement drives progress'1.00
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining0.98
-
- AlEngineer0.98
- World'sFair1.00
- Sequencing-based spatial methods: solid-phase capture0.99
- Sequencing-based methods1.00
- Section1.00
- Tissue1.00
- Barcoded1.00
- Slide1.00
- General1.00
- H&E1.00
- Solld-phose capture0.94
- Deterministic barcoding1.00
- Overlay1.00
- Ca rea0.67
- Gene1.00
- (HDST)0.94
- Expression1.00
- Seq-Scope0.99
- LatchBio1.00
- https://blog.latch.bio/p/landscape-of-sequencing-based-spatial1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.98
-
- AlEngineer0.98
- World'sFair1.00
- What does the data end up looking like? And how do you analyze it?1.00
- niches1.00
- 80.56
- 述0.54
- • GM0.73
- ·LC0.87
- {..}0.79
- vare0.90
- oc0.86
- Normalization0.97
- Dimensional1.00
- Reduction1.00
- Clustering1.00
- Cell Typing0.98
- DE1.00
- Spatial Analysis1.00
- •LR0.80
- ·PPWM0.89
- • VI0.75
- LatchBio1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining0.97
Transcript
206 cues· 3,090 words· 17,535 chars
- 0:12 Thank you to the organizers for having me.
- 0:14 I'm one of the co-founders and CTO at Latch.
- 0:17 We are basically a vertical AI lab for benchmark and agent engineering, hoping to motivate and explain exactly what that means today.
- 0:29 Starting directly with motivation for agents in bio generally,
- 0:34 Many people in my domain are familiar with this curve, but this is basically the log linear curve of data generated over the years in biology.
- 0:42 And the reason I'm bringing it up, it will become directly important to the kinds of things we want to do in engineering.
- 0:47 This curve is driven by a very small handful of experimental classes.
- 0:52 One is called single cell biology.
- 0:54 This is where we split up cells, break them apart, measure their RNA.
- 0:58 The second is spatial biology, which will become the focus of the next segment of the talk.
- 1:02 Same thing as single cell, but you get spatial resolution.
- 1:04 You can look at how RNA is spread out geometrically over a tissue.
- 1:08 And the third thing is proteomics.
- 1:10 It's a broad category of different techniques.
- 1:11 They measure proteins.
- 1:13 Less abundant in ordering, like less data volume generated relative to the other two, but still important.
- 1:21 You guys are technical, and I always think it's good to ground things somewhat quantitatively, but these are really big numbers, and the experimental data from these techniques is growing quite rapidly, almost greater than any other domain of science other than particle collider machines.
- 1:39 Single cell experiments can yield two to six terabytes per run.
- 1:43 Spatial runs can yield seven terabytes of run, proteomics a few hundred gigs.
- 1:48 And the only reason I bring this up is to say, hey, the output of a single experiment can exceed what a scientist can safely store on a consumer laptop in many cases.
- 1:58 And the law is driving how the molecular capture works point to rapid gains in this throughput over the coming years.
- 2:05 One thing I like to do is when I read a new paper is decompose it and align it to this framework because it will become important in a second.
- 2:13 Modern biology research is centered around those experiments.
- 2:17 You basically choose a model, biological model, not the kind of models you guys are used to.
- 2:21 You generate data from that model, you process the data, you creatively think about the results in the context of prior literature, and you make a claim.
- 2:30 Almost all modern experiments papers that you see published follow this loose structure with a lot of nuance.
- 2:37 All's that to say is they become something of a panning experiment.
- 2:40 You're looking for signal using measurement in a sea of noise.
- 2:44 And so this is building up to the claim that, like code and SWE, data analysis scaffolds agentic biology.
- 2:50 It becomes this executable substrate that we can use to train things.
- 2:54 It induces a natural way to benchmark and climb capability.
- 2:58 I've written about this a lot at this blog, link here.
- 3:02 There's a lot more depth to this claim, so I wouldn't take the face value, but it's something to look into.
- 3:06 All you can take away from this is just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, data analysis might do the same thing in bio.
- 3:17 So how did we get started?
- 3:18 We were originally a data tool vendor for biotech and pharma.
- 3:25 where we started five years ago out of Berkeley.
- 3:28 I'm 25, we started when I was 20.
- 3:29 We store, transform file data from large experiments as a service.
- 3:33 We'll try to build products, export lots of things,
- 3:36 Over the last two years, we started moving away from biotech and pharma and more towards the people who build those kits I was talking about, package the software into kind of white-labeled things that they provide to the scientists themselves, help them analyze their data.
- 3:53 And then over time, there became this strong interaction with the agents using the infrastructure components as tools and the loop context you guys are familiar with.
- 4:03 Except in our domain, the tools can take days or weeks.
- 4:05 I'm serious.
- 4:09 What started to happen around last summer is agent prototypes started to work.
- 4:13 So we took coding models.
- 4:16 To our knowledge at the time, they were not seriously post-trained on any tasks in biology to this point.
- 4:21 And we started to build products that look a lot like all the other agent products.
- 4:25 They have a chat interface for you to ask questions to.
- 4:28 And they build dashboards and dispatch operations to external compute.
loading
Chapters
- 0:00 Terabytes of experimental data
- 1:41 Decomposing a new paper into tasks
- 2:44 A verifiable substrate for science
- 3:23 Five years in pharma
- 4:13 Coding models as biology tools
- 5:40 Why frontier models can't be trusted yet
- 6:44 Sequencing based spatial analysis
- 9:16 Reasoning, not memorized knowledge
- 10:07 Trajectory data from real scientists
- 12:14 Why biology tasks are messy
- 15:36 Biosecurity and refusals