Videos jRCpXUjz4CI
Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
Scene timeline
74 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 217
- whisperx 217
- chunks
- 37
- from 217 cues
- keyframes
- 48
- kept of 74 captured
- frames with text
- 48
- 1,032 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 8.5 MB
- word timings on 217 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:55 | 0s |
stt |
done | — | 2026-08-09 07:22 | 23s |
chunk |
done | — | 2026-08-09 07:22 | 0s |
text_embed |
done | — | 2026-08-10 19:40 | 0s |
keyframe |
done | — | 2026-08-09 07:22 | 2m 51s |
ocr |
done | — | 2026-08-09 07:25 | 25s |
frame_embed |
done | — | 2026-08-10 19:40 | 8s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.97
- Amazon AGI Lab0.99
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo0.99
- reducto1.00
- Sonar1.00
- Makers of1.00
- together.ai0.98
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- World's Fair0.96
-
- Ask Gemini1.00
- Relaunch to update i0.97
- AlEngineer0.98
- World's Fair0.97
- 40.78
- PRESENTED BY0.99
- You're going to see Avengers0.98
- Microsoft1.00
- Infinity War later today1.00
- World's Fair0.97
- Engineering tasf0.87
- füture of Al0.93
-
- 中D0.54
- Ask Gemini1.00
- Relaunch to update i0.98
- AlEngineer0.97
- World'sFair0.96
- 40.80
- PRESENTED BY1.00
- Musical.ly was about to rebrand0.98
- Microsoft1.00
- to“TikTok"0.97
- World's Fair0.95
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- 00.55
- Ask Gemini0.99
- Relaunch to update0.98
- AlEngineer0.97
- World'sFair1.00
- staysaasy1.00
- Follow1.00
- @staysaasy1.00
- It's 2018 and your coworker just sent you a 400 line pull request.0.99
- You get a cup of coffee and sit down to review it.1.00
- It's beautiful. Elegant micro-refactors. Crispy method names.1.00
- You catch a few things, but that's ok. It's part of the dance. They didn't1.00
- consider extensibility on part of their APl. Here's a comment buddy.0.98
- World's Fair0.95
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- 中p0.51
- Ask Gemini1.00
- Relaunch to update i0.95
- AlEngineer0.98
- World'sFair0.97
- 40.88
- Musical.ly was about to rebrand1.00
- to“TikTok"0.97
- World's Fair0.96
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- Ask Gemini1.00
- Relaunch to update i0.96
- AlEngineer0.98
- World'sFair1.00
- 40.87
- What will the history books say0.99
- about software engineering?1.00
- World'sFair0.98
- TRACK 5· JULY 1, 20260.93
- Evals1.00
-
- Ask Gemini1.00
- Relaunch to update i0.95
- AlEngineer0.97
- World'sFair1.00
- Q0.60
- "Software engineering was when1.00
- you knew what the code would do1.00
- before you ran it."0.99
- World's Fair0.95
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- Ask Gemini0.98
- Relaunch to update0.94
- AlEngineer0.96
- World'sFair1.00
- Q0.61
- François Chollet0.98
- Subscribe1.00
- @fchollet1.00
- Agentic coding is a form of machine learning. Generated code is best0.99
- treated as a blackbox artifact whose behavior and generalization should1.00
- be managed via empirical evaluation, like with any ML model.1.00
- World'sFair1.00
- TRACK 5· JULY 1,20260.97
- Evals1.00
-
- Ask Gemini1.00
- Relaunch to update i0.99
- AlEngineer0.97
- World'sFair1.00
- 40.58
- def extract_phone_number(text):1.00
- match = re.search(r"\d{3}-\d{3}-\d{4}", text)0.97
- return match.group()1.00
- def main():0.98
- text = "My number is 555-123-4567"0.98
- print(extract_phone_number(text))1.00
- World'sFair0.99
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- CD0.74
- 00.87
- 80.80
- Ask Gemini1.00
- Relaunch to update0.98
- AlEngineer0.98
- World'sFair1.00
- def extract_phone_number(text):1.00
- 40.99
- match = re.search(r"\d{3}-\d{3}-\d{4}", text)0.98
- return match.group()0.97
- response = OpenAI().responses.create(1.00
- model="gpt-5.5",0.99
- input=f"Extract the phone number:\n\ntext}",0.98
- return response.output_text0.99
- def main():1.00
- text = "My number is 555-123-4567"0.98
- print(extract_phone_number(text))1.00
- World's Fair0.99
- TRACK 5 • JULY 1, 20260.95
- Evals1.00
-
- CD0.57
- Ask Gemini0.96
- Relaunch to update0.95
- AlEngineer0.99
- World's Fair0.99
- 40.87
- Machine Learning1.00
- → Agent Building0.99
- Training data0.99
- PRESENTED BY0.99
- →0.81
- Microsoft1.00
- World's Fair0.97
- TRACK 5 · JULY 1, 20260.92
- Evals1.00
-
- D0.90
- 白0.55
- Ask Gemini0.99
- Relaunch to update i0.96
- AlEngineer0.99
- World's Fair0.99
- 40.90
- Machine Learning1.00
- → Agent Building0.97
- PRESENTED BY1.00
- Training data0.99
- → Environments0.99
- Test / validation set0.98
- → Evals (environments)0.98
- Microsoft1.00
- Weights1.00
- → Skills, prompts, tools, model0.99
- World's Fair0.97
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- 中D0.61
- Ask Gemini1.00
- Relaunch to update i0.95
- AlEngineer0.99
- World's Fair0.99
- 40.99
- Machine Learning1.00
- → Agent Building0.96
- Training data1.00
- → Environments0.99
- Test / validation set0.98
- → Evals (environments)0.98
- Weights1.00
- → Skills, prompts, tools, model0.99
- Loss function1.00
- → Environment rewards & feedback0.96
- Backprop / optimizer0.97
- → GEPA or text-based optimization0.98
- Gradient descent step1.00
- → Pull request0.98
- Overfitting1.00
- → Reward hacking (or overfitting)0.97
- World's Fair0.99
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- 中D0.59
- Ask Gemini1.00
- Relaunch to update i0.98
- AlEngineer0.99
- World's Fair0.96
- 40.77
- Agents >> ML0.99
- World's Fair0.98
- TRACK 5 · JULY 1, 20260.92
- Evals1.00
-
- CD0.78
- Ask Gemini1.00
- Relaunch to update i0.94
- AlEngineer1.00
- World's Fair0.99
- 40.92
- François Chollet1.00
- Subscribe1.00
- @fchollet1.00
- agents are1.00
- agent performance0.99
- Agentic coding is a form of machine learning. Generated code is best0.98
- treated as a blackbox artifact whose behavior and generalization should1.00
- be managed via empirical evaluation, like with any ML model.1.00
- World's Fair0.97
- TRACK 5· JULY 1,20260.96
- Evals1.00
-
- 中p0.59
- Ask Gemini0.97
- Relaunch to update0.98
- AlEngineer0.98
- World'sFair1.00
- Environments1.00
- 40.77
- Instruction1.00
- Sandbox1.00
- World's Fair0.96
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- CD0.67
- G0.63
- Ask Gemini0.99
- Relaunch to update0.93
- AlEngineer0.99
- World's Fair0.99
- Environments1.00
- 40.81
- Instruction1.00
- Sandbox1.00
- Verifier1.00
- World's Fair0.97
- TRACK 5 · JULY 1, 20260.92
- Evals1.00
Transcript
217 cues· 3,491 words· 19,050 chars
- 0:16 Awesome, thank you so much.
- 0:19 So yeah, like you said, my name's Alex Shaw.
- 0:22 I work at Lott Institute.
- 0:24 And I'll be speaking today about Harbor, which is an agent evaluation and RL environment framework.
- 0:31 And the title of my talk is Everything is a Rollout.
- 0:34 And I think you'll see as I get into it.
- 0:37 why we titled the talk that way.
- 0:39 But first, I want everybody to come travel back in time with me to the year 2018, so eight years ago.
- 0:49 And we're going to talk about some of the things that were going on in 2018.
- 0:53 So probably you're going to see Avengers Infinity War later today.
- 0:58 The second one is just getting released.
- 1:01 GPT-1 was just released, and you probably didn't even notice, although maybe some people did.
- 1:10 Musically, this random startup was about to rebrand to a product called TikTok.
- 1:17 And you might have just learned about AirPods when you saw somebody walking around with headphones that had no cord.
- 1:23 So what about software engineering?
- 1:25 What did software engineering look like in 2018?
- 1:30 We're gonna read this tweet from Stay Sassy Sassy with two A's about what it was like to write code in 2018.
- 1:38 So it says, it's 2018 and your coworker just sent you a 400 line pull request.
- 1:43 You get a cup of coffee and sit down to review it.
- 1:46 It's beautiful, elegant micro-refactors, crispy method names.
- 1:50 You catch a few things, but that's OK.
- 1:51 It's part of the dance.
- 1:53 They didn't consider extensibility on part of their API.
- 1:56 Here's a comment, buddy.
- 1:57 And this is actually just part of the tweet, so it keeps going.
- 2:00 You should look it up if you want.
- 2:02 And it's obviously written humorously.
- 2:06 But the thing is, it does feel a little bit nostalgic, just like some of these other things that were going on in 2018.
- 2:15 Now the thing is, this only changed maybe six or 12 or 18 if you're a very early adopter.
- 2:22 months ago, but it already feels like the distant past in some ways.
- 2:27 And I think it's time to start talking about, well, what will the history books say about software engineering?
- 2:34 And when I say software engineering, I mean that style of 2018 software engineering.
- 2:40 It will probably say, the history books will probably say a lot of things, but I think one thing for sure that they'll say is software engineering was when you knew what the code would do before you ran it.
- 2:53 So, and that brings me then to a different tweet from Francois Chalet, where he says, agentic coding is a form of machine learning.
- 3:02 Generated code is best treated as a black box artifact whose behavior and generalization should be managed via empirical evaluation, like with any ML model.
- 3:12 And that kind of brings us to the next part in this talk, which is to compare and contrast what agent development looks like versus what more traditional software engineering development looks like and why it demands a new set of tools to really understand what's going on and have confidence and trust.
- 3:31 So here's a 2018 program right here.
- 3:34 So you can tell already the purpose of the program is to extract phone numbers from text.
- 3:41 And we have a regex right here that looks for the phone number.
- 3:45 And I can say with 100% confidence what will happen if I run this program 1 million times in a row.
- 3:55 So now let's update it to the 2026 version.
- 3:59 So I swap out my regex.
- 4:02 And instead, obviously, I throw in my model call instead.
- 4:07 And I say, extract this phone number.
- 4:10 So in some ways, this is actually a more powerful program because the regex was actually a little bit brittle.
- 4:16 It would have missed any phone number that wasn't formatted exactly like how it was specified.
- 4:22 Whereas I'm pretty confident that this program with GPT 5.5 will catch a lot of the phone numbers that are formatted weirdly.
- 4:30 However, if I ran this exact program 1 million times, I'm not 100% confident that it will print the same thing every single time or that I know exactly what it will print.
- 4:44 It probably gets it right almost every time.
- 4:46 This is a pretty simple task.
loading