Videos O-CBZ3JtRvo
Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face
Scene timeline
85 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 194
- whisperx 194
- chunks
- 30
- from 194 cues
- keyframes
- 39
- kept of 85 captured
- frames with text
- 39
- 1,442 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 12.0 MB
- word timings on 194 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:59 | 0s |
stt |
done | — | 2026-08-09 13:54 | 23s |
chunk |
done | — | 2026-08-09 13:54 | 0s |
text_embed |
done | — | 2026-08-10 19:40 | 0s |
keyframe |
done | — | 2026-08-09 13:54 | 2m 45s |
ocr |
done | — | 2026-08-09 13:57 | 23s |
frame_embed |
done | — | 2026-08-10 19:40 | 7s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.97
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust1.00
- bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- neo4j0.98
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto0.99
- Sonar1.00
- Makers of0.98
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- AI.ENGINEER // July 1st, 20260.98
- World's Fair0.96
- Training Al models to0.99
- PRESENTED BY1.00
- out-execute attackers1.00
- Microsoft1.00
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber0.99
- back to defenders - and why high-quality data is the path.0.99
- ARITHMETIC1.00
- Uri Rolls0.98
- with1.00
- Thomas Wolf1.00
- Hugging Face1.00
- Co-Founder and CEO1.00
- Co-Founder and CSO1.00
- Al.Engineer // Training Frontier Models to Out-Think Hackers1.00
- World's Fair0.99
- Engineering the future of Al0.96
-
- AlEngineer0.97
- AI.ENGINEER // July 1st, 20260.98
- World'sFair1.00
- Training Al models to0.99
- PRESENTED BY1.00
- out-execute attackers1.00
- Microsoft1.00
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber0.99
- back to defenders - and why high-quality data is the path.0.99
- ARITHMETIC1.00
- Uri Rolls1.00
- with1.00
- Thomas Wolf0.97
- Hugging Face1.00
- Co-Founder and CEO1.00
- Co-Founder and CSO1.00
- Al.Engineer // Training Frontier Models to Out-Think Hackers1.00
- World's Fair0.98
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
-
- PRIZE0.79
- FOUNDATION0.95
- LEADERBOARDS0.98
- PRIZE0.99
- RESEARCH1.00
- CONTENT1.00
- AlEngineer0.99
- SERES0.98
- ARC-AGI-10.99
- ARC-AGI-20.94
- ARC-AGI-30.95
- World's Fair0.97
- 220.71
- ARC-AGI-31.00
- ARC-AGI-31.00
- START1.00
- measure human-like intelligence in Al agents.0.98
- The first interactive reasoning benchmark designed to0.98
- PRESENTED BY1.00
- Play [Humans]0.99
- Buld [AI]0.93
- Microsoft1.00
- LINKS0.88
- PUBLIC GMME SET0.97
- HUPAN LEADERBOARDS0.95
- DOCS + SOK0.88
- ARC PRIZE ZOZG TRACK0.97
- TECHNICAL PAPER0.91
- What is ARC-AGI-320.98
- ARC-AGI-3 is an interactive reasoning benchmark which challenges Al agents to explore novel0.98
- environments, acquire goals on the fy, build adaptable world models, and leam continuously.0.99
- A 100% score means Al agents can beat every game as efficiently as humans.0.98
- Instead of solving static puzzles, agents must leam from experience inside each environment—0.98
- perceiving what matters, selecting actions, and adapting their strategy without relying on natural-0.99
- language instructions.1.00
- World's Fair0.97
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining1.00
-
- PRIZE0.80
- FOUNDATION0.93
- LEADERBOARDS0.98
- BENCHSARK0.86
- PRIZE0.99
- RESEARCH1.00
- CONTENT1.00
- AlEngineer0.98
- SERES0.97
- ARC-AGI-10.99
- ARC-AGI-20.93
- ARC-AG-30.96
- World'sFair1.00
- 220.67
- ARC-AGI-31.00
- ARC-AGI-31.00
- START1.00
- measure human-ike intelligence in Al agents.0.97
- The first interactive reasoning bendhmark designed to0.98
- Play [Humara]0.96
- Buid [AI]0.90
- LNKS0.92
- PUBLIC GMME SET0.97
- HUPNN LEADERBOARDS0.94
- DOCS + SOK0.86
- ARC PRIZE ZOZG TRACK0.98
- TECHNICAL PAPER0.90
- What is ARC-AGI-320.98
- ARC-AGI-3 is an interactive reasoning benchmark which challenges Al agents to explore novel0.99
- environments, acquire goals on the fly, build adaptable world models, and leam continuousty.0.98
- A 100% score means Al agents can beat every game as efficiently as humans.0.97
- Instead of solving static puzzles, agents must leam from experience inside each environment—0.99
- perceiving what matters, selecting actions, and adapting their strategy without relying on natural-1.00
- language instructions.1.00
- World's Fair0.98
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- Format1.00
- Arrange1.00
- Slide Show0.98
- Window1.00
- Help1.00
- AIE WF26 - Uri Th..now0.88
- 00.67
- G0.99
- AutoSave0.98
- Lengineer_finaLno_opening - Saved to my Mac0.97
- Q Search (Cmd + Ctrl + U)0.96
- AlEngineer0.99
- Insert1.00
- Draw1.00
- Design1.00
- Transitions1.00
- Animations1.00
- Slide Show0.96
- Record1.00
- Review1.00
- View1.00
- Picture Format0.97
- Comments1.00
- Share1.00
- World's Fair0.98
- Layout0.96
- Shape FI0.77
- 03021.pdf0.99
- Add-ins0.96
- Training Al models to0.99
- out-execute attackers0.99
- ..07.26.280.97
- ai_engineer_v1.pd0.98
- AI.ENGINEER // July 1*, 20260.96
- Training Al models to1.00
- besketbull re the same sport.0.74
- out-execute attackers0.98
- 0.:20.07.500.93
- the_coodels_clean1.00
- 011.00
- Maskff lenchomark0.80
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber0.99
- back to defenders - and why high-quality data is the path.0.99
- 19.27.09 2026-0...t20.31.120.94
- Why acces control0.98
- 1..17.51.230.89
- s.5vg0.96
- ARITHMETIC1.00
- Uri Rolls0.99
- Co-Founder and CEO1.00
- with1.00
- Thomas Wolf1.00
- Co-Founder and CSO1.00
- Hugging Face1.00
- 40.91
- AlSmging// Training Frontier Models to Out-Think Hackeers0.87
- .19.530.91
- Why acces control0.98
- Click to add notes1.00
- Slide 1 of 270.96
- English (United States)1.00
- Accessibility: Imvestigate0.97
- Notes0.95
- Comments0.94
- 88 00.70
- 120%1.00
- The economics of yber offense are0.89
- shifting Attackers are becoming les0.90
- resource-constrained, and a vider0.88
- Click to add notes0.99
- 23.20.41.00
- 2026-0...09.25.200.98
- Screenshot1.00
- Slide 1 of 370.97
- English (United States)1.00
- Accessibility: Investigate0.98
- NotesComments880.87
- 123%1.00
- World's Fair0.96
- TRACK 9· JULY 1,20260.96
- Posttraining & Midtraining1.00
-
- Microsoft PowerPoint1.00
- Insert1.00
- Arrange1.00
- Slide Show0.98
- Window1.00
- Help1.00
- 00.85
- AlE WF26 - Uri Th..nowow0.90
- 00.68
- G0.91
- 00.61
- C0.82
- ai_engineer_fina_no_opening -Saved to my Mac0.96
- AlEngineer0.99
- Insert1.00
- Draw1.00
- Design1.00
- Transitions1.00
- Animations1.00
- Slide Show0.97
- Record1.00
- Review1.00
- View1.00
- Picture Format1.00
- World's Fair0.99
- Cut0.99
- Layout0.99
- 03021.pdf1.00
- Training Al models to0.98
- out-execute attackers0.94
- 0...07.26.281.00
- AI.ENGINEER // July 1*, 20260.94
- Training Al models to1.00
- er_tras...d (1)d.svg0.90
- Wecannot capture all'cyber ina0.75
- Thatse syin ming and0.56
- single benchmark.0.94
- besketbull re the same sport.0.81
- out-execute attackers1.00
- the_coodels_clean1.00
- 011.00
- Maskff eenchomark0.79
- back to defenders - and why high-quality data is the path.0.99
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber1.00
- Why access control0.98
- 白0.75
- ...t 17.51.2230.81
- vectorized_coodel0.99
- ARITHMETIC1.00
- Uri Rolls0.99
- Co-Founder and CEO1.00
- with1.00
- Thomas Wolf1.00
- Co-Founder and CSO1.00
- Hugging Face1.00
- 40.99
- Ang//Training Prontie Models to Out-Think Hackers0.74
- t 17.51.250.95
- 2026-0...19.53.270.98
- Screenshot0.97
- Why access control1.00
- Click to add notes1.00
- Slide 1 of 271.00
- English (United States)0.98
- Notes0.93
- Comments1.00
- 80.71
- shifting Attackers are becoming les0.92
- The economics of cyber offlense are0.96
- 2026-0...09.25.200.99
- Screenshot1.00
- Slide 1 of 370.99
- range of targets becomes feosible to hit0.94
- English (United States)0.98
- Accessibility: Investigate0.98
- Notes Comments880.91
- 123%1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.99
-
- Arrange0.99
- Slide Show0.98
- Help1.00
- AIE WF26 - Uri Th..now0.89
- 00.77
- AutoSave0.98
- C0.82
- al_engineer_final_no_opening -Saved to my Mac0.95
- AlEngineer0.99
- Animations1.00
- Side Show0.98
- Record0.93
- Review1.00
- View1.00
- Picture Format1.00
- Record1.00
- Share1.00
- World's Fair0.98
- 240.98
- 03021.pdf1.00
- AutoSave0.99
- C0.93
- ai_engineer_final — Saved to my Mac0.93
- Animations1.00
- Slide Show0.99
- Record1.00
- Review1.00
- View1.00
- 日0.51
- Add-ins1.00
- .07.26.280.96
- engineer_v1.pd1.00
- Training Al models to0.99
- out-execute attackers0.99
- AI.ENGINEER // July 1*, 20260.97
- Training Al models to1.00
- 0.:20.07.500.94
- the_coodels_clean0.99
- Why i hs happening?0.92
- out-execute attackers1.00
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber1.00
- Cyber is a hattle of shil and speed0.93
- back to defenders - and why high-quality data is the path.0.98
- 1...t17.51.230.88
- s.5vg0.96
- ARITHMETIC1.00
- Co-Founder and CEO1.00
- Uri Rolls0.99
- with1.00
- Co-Founder and CSO1.00
- Thomas Wolf1.00
- Hugging Face0.99
- AlEngineer // Training Frontier Medels to Out-Think Hackers0.96
- The economics of cyber offlense are0.92
- shifting Artackers are becoming less0.97
- 2026-0..09.25.200.96
- Screenshot1.00
- World's Fait0.92
- Side 1 of 370.98
- English (United States)0.98
- 日0.82
- 123%1.00
- TRACK 9 • JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- AlEngineer0.97
- AI.ENGINEER // July 1st, 20261.00
- World'sFair1.00
- Training Al models to0.99
- out-execute attackers1.00
- Why AI doesn't have to be the end of cybersecurity, but instead a chance to shift the cyber0.99
- back to defenders - and why high-quality data is the path.0.99
- ARITHMETIC1.00
- Uri Rolls0.99
- with1.00
- Thomas Wolf1.00
- Hugging Face1.00
- Co-Founder and CEO1.00
- Co-Founder and CSO1.00
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- World'sFa0.99
- G/Q0.50
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining1.00
-
- AlEngineer0.93
- WIRED1.00
- World'sFair1.00
- 'Dangerous' Al Models Are Coming No Matter What0.99
- The US govemment crackdown on Anthropic's Claude Fable 5 and Mythos 5 hides a0.99
- glaring truth: Al models with advanced hacking capabilities...0.97
- 1 week ago0.97
- 1Fathom Journal0.96
- AME CLAUDI1.00
- Claude Fable 5 & Mythos 5: The Too Dangerous' Al, Split In0.98
- Two Thurrock Council (7nUZqfAYSD)0.98
- In April, Anthropic unveiled Claude Mythos - a model so good at finding and exploiting0.99
- software vulnerabilities that the company said it wouldn't r...0.99
- 3 days ago0.98
- PCWorld0.99
- Claude's 'too dangerous' Al model is finally public. But0.98
- there's a catch1.00
- Claude Fable 5 is Anthropic's de-fanged Mythos-class model. Paid subscribers will only1.00
- have access until June 23rd without paying extra.0.99
- 2 weeks ago1.00
- nNature0.99
- Too dangerous to release: is Mythos the start of the0.98
- restricted-Al era?1.00
- Whathappens when Al companies produce models that they say the public can't have0.99
- - and how should users and governments react?1.00
- 1 month ago0.99
- Al.Engineer // Training Frontier Models to Out-Think Hackers1.00
- GQ0.76
- TRACK 9· JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- WIRED1.00
- World's Fair0.99
- 'Dangerous' Al Models Are Coming No Mat1.00
- The US govemment crackdown on Anthropic's Claude Fable 50.99
- CYBERSECURITY1.00
- glaring truth: Al models with advanced hacking capabilities..0.99
- Al speeds cybercrime by exposing flaws,0.99
- 1 week ago0.96
- and other cybersecurity news0.99
- 1Fathom Joumal0.96
- Claude Fable 5 & Mythos 5: The Too0.98
- Two Thurrock Council (7nUZqfAYSD)0.99
- software vulnerabilities that the company said it wou1.00
- In April, Anthropic unveiled Claude Mythos - a mode0.99
- Project0.96
- 3 days ago0.99
- Glasswing1.00
- W PCWorld0.94
- Claude's 'too dangerous' Al model is0.99
- Securing critical software0.99
- there's a catch1.00
- for the Al era0.99
- Contrue rading0.91
- Claude Fable 5 is Anthropic's de-fanged Mythos-clas:0.99
- have access until June 23rd without paying extra.1.00
- 2 weeks ago0.95
- nNature1.00
- Too dangerous to release: is Mythos the start of the0.99
- restricted-Al era?1.00
- What happens when Al companies produce models that they say the public can't have1.00
- - and how should users and governments react?1.00
- 1 month ago1.00
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- TECH0.96
- China Has Matched0.98
- AlEngineer1.00
- WIRED1.00
- Anthropie in Cybersecurity,1.00
- World'sFair1.00
- 'Dangerous' Al Models Are Coming No Mat0.99
- Resetting AI Race0.99
- The US govemment crackdown on Anthropic's Claude Fable 50.98
- CYBERSECURITY1.00
- Clampdowo on top U. artifial itllngence is0.79
- glaring truth: Al models with advanced hacking capabilities..0.99
- fueling concern that Washington is handing Bejing a0.97
- 1 week ago0.96
- and other cybersecurity news0.99
- 1Fathom Jourmal0.93
- Claude Fable 5 & Mythos 5: The Too0.99
- Two Thurrock Council (7nUZqfAYSD)0.98
- € Hi lm zai0.84
- software vulnerabilities that the company said it wou0.98
- In April, Anthropic unveiled Claude Mythos - a mode0.99
- Project0.96
- 3 days ago0.99
- Glasswing1.00
- PCWorld1.00
- Claude's 'too dangerous' Al model is0.99
- Securing critical software0.95
- there's a catch1.00
- for the Al era0.99
- Claude Fable 5 is Anthropic's de-fanged Mythos-clas:1.00
- orchestrated cyber espionage campaign0.99
- Disrupting the first reported Al-1.00
- have access until June 23rd without paying extra.0.99
- 2 weeks ago0.97
- nNature1.00
- Too dangerous to release: is Mythos the0.98
- restricted-Al era?1.00
- What happens when Al companies produce models that t1.00
- 80.99
- - and how should users and governments react?1.00
- 1 month ago0.98
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- TRACK 9· JULY 1, 20260.94
- Posttraining & Midtraining1.00
-
- TECH0.96
- China Has Matched0.98
- AlEngineer0.99
- WIRED1.00
- Anthropie in Cybersecurity,1.00
- World'sFair1.00
- 'Dangerous' Al Models Are Coming No Mat1.00
- Resetting AI Race0.99
- The US govemment crackdown on Anthropic's Claude Fable 51.00
- CYBERSECURITY1.00
- Clampdown on top U. artifial ntelligence is0.78
- glaring truth: Al models with advanced hacking capabilities..1.00
- fueling concern that Washingtonis handing Beijing a0.97
- 1 week ago0.96
- and other cybersecurity news0.99
- 1Fathom Jourmal0.93
- Claude Fable 5 & Mythos 5: The 'Too0.98
- Two Thurrock Council (7nUZqfAYSD)0.99
- Hi lm z ai0.81
- In April, Anthropic unveiled Claude Mythos - a mode0.99
- software vulnerabilities that the company said it wou0.99
- Project0.99
- 3 days ago0.99
- Glasswing1.00
- Daybreak1.00
- PCWorld1.00
- Claude's 'too dangerous' Al model is0.98
- Securing critical software0.95
- there's a catch1.00
- for the Al era0.99
- Claude Fable 5 is Anthropic's de-fanged Mythos-clas:1.00
- orchestrated cyber espionage campaign gle says hackers used Al to0.99
- Disrupting the first reported Al-1.00
- have access until June 23rd without paying extra.0.99
- 'elop a major security flaw0.96
- 2 weeks ago0.99
- n Nature0.95
- Too dangerous to release: is Mythos the0.99
- Google0.98
- restricted-Al era?1.00
- Whathappens when Al companies produce models that t0.98
- 80.99
- - and how should users and governments react?1.00
- 1 month ago1.00
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- World's Fair0.94
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- Why is this happening?0.98
- World'sFair1.00
- 自n0.51
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
-
- AlEngineer0.97
- Why is this happening?0.98
- World'sFair1.00
- Cyber is a battle of skill and speed.0.99
- PRESENTED BY0.97
- Z0.94
- Microsoft1.00
- GLM-5.21.00
- Skilled hacker1.00
- World'sFair0.98
- Al.Engineer // Training Frontier Models to Out-Think Hackers0.99
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.99
Transcript
194 cues· 3,395 words· 18,010 chars
- 0:13 Hello, everyone.
- 0:14 Thanks for showing up at this data quality.
- 0:17 So as you saw, we'll probably talk a little bit about other things than data quality.
- 0:21 And first, I actually tell you about why I'm very excited about this talk and why actually I accepted to come present it with Yuri.
- 0:31 There's two reasons for that.
- 0:32 And the first reason is, I think, and what you will see today is that cybersecurity is a much wider field for AI, a much wider playing field and exploration field than you might think.
- 0:45 And in particular, what we'll show you today is a benchmark
- 0:49 that Arithmetic and URI has been developing, and I've been a little bit advising, which I think is even close to things like Arc AGI 3 for people who have been following progress around AGI, in which that, by which I mean that this benchmark asks for people, oh yeah, cool, to do, if you have been playing with Arc AGI 3, some of you, who knows here Arc AGI 3, or Arc AGI in general?
- 1:15 One person?
- 1:17 OK, we're in a very data quality field.
- 1:20 Basically, this asks models to try to understand what's happening in the world.
- 1:24 And it's actually small games that the models need to understand basically what's the current state, how we can play with it, and how we can actually change the state of the game.
- 1:35 So it's actually something you can play yourself.
- 1:38 So basically, it asks a model to understand when you click somewhere, something happening at another place.
- 1:43 And you may think this is very simple and that we should be past that on our way to AGI.
- 1:50 The thing you will discover if you play with this benchmark is no.
- 1:54 Models have 1% to 2% success rate on this generic benchmark.
- 1:59 And the reason is the current model, even though they're really good, they can't really build a dynamic model of what's happening in the world or what's happening in any type of world.
- 2:08 And I think that the benchmark that
- 2:11 Arithmetic has been developed as a benchmark that's also extremely challenging for a model in that they need to understand what's happening and to act accordingly to the world model they've been building on the fly.
- 2:21 So that's the first reason I think this benchmark is really interesting and why I'm actually very happy to show you that.
- 2:26 And the second reason is I think there's a lot of things that open source model can bring in cybersecurity.
- 2:33 And we tend to have this very binary view of closed source model are good for cyber, open source model are bad.
- 2:39 And I think what we want to show today is that open source model are one part of the solution to cybersecurity challenges today.
- 2:48 And in particular, if you think in terms of attack and defense and how it will be in the future, this balance, we think that cyber and open source model use in cybersecurity will be key to actually be able to solve the defense solution.
- 3:04 So there is a future, I think, where cyber is alive and everyone is well protected.
- 3:09 And I'm pretty sure this future involves open source model.
- 3:12 Now, Yuri is also kind of an impressive person, so I'm really happy I met him.
- 3:18 He was studying at Harvard, dropped to build this idea of what the future of cyber should be, and it's an honor to have you on stage with me.
- 3:28 Thank you so much, Thomas.
- 3:29 And thank you everyone who came.
- 3:30 I'm really grateful.
- 3:32 And I'm also grateful for the work we've done together to build this benchmark, which is, I think, incredibly difficult for the models and actually shows some really, really interesting leaps in places that I think we still have way to go.
- 3:44 I guess the way we've been thinking about the problem and the reason we set out to do this is it's very clear that the economics of cyber are fundamentally shifting.
- 3:52 There's this inherent thing that is inherent to cyber, which is that attackers need to choose their resource really wisely.
- 3:59 And if you sort of think about cyber as a house in a way, then my job is to block every door and close every window and make sure that there's no way in.
- 4:07 And the attacker's job is to find at least one scene, one crack, one thing I missed.
- 4:11 And then once inside, my job is to put sensors and anything I can to keep them out.
- 4:17 And the whole stack, the entire world of cyber that we've been building for the past 20 years has been based on this economics that the attackers have to choose their targets, and we do everything we can across it to protect ourselves.
- 4:29 It is true that that is changing in really dramatic ways.
- 4:31 The models are incredibly powerful.
- 4:34 They're able to find a ton of primitives.
- 4:36 They're able to find a bunch of zero-day exploits.
- 4:38 We're seeing this.
- 4:38 There's so many news and chaos around this point.
- 4:42 And on the other hand, it seems like we as defenders don't seem to be prepared for this world and the way that it's coming.
- 4:49 And I actually think that I might be on the wrong one.
- 4:54 What I want to show is this idea of
- 4:58 So if you think about the way cyber has been, this is definitely, I think, the way we've been thinking about AI and cyber for the past many, many, many months.
- 5:08 And it's freaky, and it's getting really scary.
loading
Chapters
- 0:00 Why cyber is a wide new field for AI
- 1:34 The ARC-AGI-3 parallel: models can't model the world
- 2:24 Open source models as part of the defense
- 3:45 The shifting economics of cyber
- 5:52 The optimistic thesis: models are the solution
- 7:06 The first benchmark: access control
- 8:20 Data quality: finding your own zero days
- 10:01 A real solve: the Keycloak name versus ID exploit
- 11:57 Live demo: one solve at K1
- 14:22 Only models can replace the old stack
- 15:00 The speed challenge and specialized defenders