Videos XLEYtv3cMlw
Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait
Scene timeline
114 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 121
- whisperx 121
- chunks
- 35
- from 121 cues
- keyframes
- 109
- kept of 114 captured
- frames with text
- 109
- 3,343 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 15.6 MB
- word timings on 121 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:02 | 1m 29s |
stt |
done | — | 2026-08-11 00:03 | 19s |
chunk |
done | — | 2026-08-11 00:03 | 0s |
text_embed |
done | — | 2026-08-11 00:03 | 1s |
keyframe |
done | — | 2026-08-11 00:03 | 1m 26s |
ocr |
done | — | 2026-08-11 00:05 | 1m 21s |
frame_embed |
done | — | 2026-08-11 00:06 | 22s |
Frames, and what the machine read
-
- AlE_conference presentation2.tldr0.97
- Autonomous Agents for Scientific Tasks1.00
- autor1.00
- Sina Shahandeh,1.00
- Radicait1.00
- 1.0001.00
- Abstract:1.00
- 0.9951.00
- There has been much w0.99
- problems, or static supe0.99
- discovery task, the prob0.97
- very large and complex0.99
- BPB (ower0.65
- 0.9901.00
- physical model, a meas0.98
- clinic, observatory, or fi0.98
- 0.9851.00
- In this talk, we show sce1.00
- classes, and hyperpara0.98
- performance comes fro1.00
- behaves, implementing0.98
- 0.9801.00
- real data.1.00
- We show how an ontolo0.98
- generation that is key to1.00
- 0.9750.98
- industrial and applied r0.97
- TLDERRAW0.61
-
- AlE_conference presentation2.tldr0.96
- Autonomous Agents for Scientific Tasks0.97
- autor1.00
- Sina Shahandeh,1.00
- Radicait1.00
- 1.0000.96
- Abstract:1.00
- 0.9950.92
- There has been much work on Autoresearch where the objectives are coding puzzles, toy optimization1.00
- problems, or static supervised-learning ML tasks. However, for an autonomous agent to assist with a scientific1.00
- (ette sr)0.51
- discovery task, the problems must come from real measurement data of the world, the search space must be1.00
- very large and complex, and the objective must have scientific meaning. A better solution would improve a1.00
- Val g (er0.74
- 0.9901.00
- physical model, a measurement or characterization method, or experimental decision-making in an actual lab,1.00
- clinic, observatory, or field setting.1.00
- 0.9851.00
- In this talk, we show scenarios where the agent has to search over methods, priors, data preprocessing, model1.00
- classes, and hyperparameters while learning from intermediate failures. Often, a step change in agent0.99
- behaves, implementing that hypothesis correctly in a mathematical model, and executing it on the existing0.99
- performance comes from forming an appropriate scientific hypothesis about how the physical system0.99
- 0.9801.00
- real data.1.00
- We show how an ontology-based memory system is used in the harness to assist with the hypothesis0.99
- generation that is key to the agent's success. All demonstrations come from real scientific problems solved in0.99
- 0.9751.00
- industrial and applied research settings.0.99
- MLDERWAW0.57
-
- AlE_conference presentation2.tldr0.95
- Humar1.00
- Revisit1.00
- autoresearch1.00
- June 16,2026·Q0.96
- Alvin Cheung1.00
- 'UC Berkeley.20.92
- 2026#test0.98
- Autoresearch Progress: 83 Experiments, 15 Kept Improvements1.00
- 1.0000.90
- Discarded0.97
- Running best0.99
- 0.9950.91
- V(al s eter0.59
- 0.9901.00
- Elorting0.63
- 0.9851.00
- 0.9801.00
- 0.9751.00
- 00.99
- 201.00
- Experiment #1.00
- 401.00
- 601.00
- 801.00
- TLDRAW0.97
-
- AlE_conferencepresentation2.tldr0.98
- Humans Still Beat AI in the Long Horizon:1.00
- Revisiting Test-Time Scaling in the Agent Era1.00
- June 16, 2026- Qiuyang Mang1, Kaiyuan Liu2, Bo Peng3. hreyas Pimpalgaonkar4, Luke ZEettlemoyer2, Alex Dimakis1.4,0.96
- Alvin Cheung1.00
- UC Berkeley - 2 University of Washington - 3 Princeton Universit - Bespoke Labs0.95
- 2026#test-time-scaling # LLM-agents # scaling-lawsresearch0.97
- ss: 83 Experiments, 15 Kept Improvements0.99
- Running best1.00
- Kept0.99
- Discarded1.00
- Human vs Agent0.99
- fop10-humans0.95
- 18530.96
- 18001.00
- fop50-humans0.99
- Cloude Code Opus-4.60.93
- Codex GPT-5.50.96
- humons keep climbing0.99
- for days1.00
- then plafeou by 24h0.99
- 1%000.87
- 15870.99
- Eotng0.70
- ogents sprint early0.98
- 1000.97
- 12000.99
- 10921.00
- 10000.97
- Experiment #0.98
- 401.00
- 601.00
- 801.00
- 2h0.90
- 8h0.90
- th0.91
- time1.00
- 2h0.97
- 2d0.98
- 40.63
- 8d0.72
- TLD RRAW0.59
-
- AlE_conferencepresentation2.tldr0.98
- The bottlene1.00
- Humans Still Beat AI in the Long Horizon:0.99
- Revisiting Test-Time Scaling in the Agent Era1.00
- June 16, 2026- Qiuyang Mang,. Kaiyuan Lu2, Bo Peng3. Shreyas Pimpalgaonkar4, Luke Zettlemoyer2, Alex Dimakis14,0.93
- Alvin Cheung1.00
- Hypotl0.96
- 1 UC Berkeley - 2 University of Washington · 3 Princeton University · 4 Bespoke Labs0.92
- 2026 #test-time-scaling # LLM-agents # scaling-lawsresearch0.96
- Discarded1.00
- Running best0.97
- Kept1.00
- Human vs Agent0.98
- top10-humons0.96
- 18531.00
- 18001.00
- Cloude Code Opus-4.60.99
- Codex GPT-5.51.00
- top50-humans0.98
- humans keep climbing0.97
- for days0.99
- then plateou by 24h0.98
- 16000.88
- 15871.00
- Eotng0.74
- agents sprint eorly0.96
- 1000.90
- 13681.00
- 12000.95
- 10921.00
- 10001.00
- 2h0.78
- 18h0.73
- time1.00
- 24h0.84
- 2d0.93
- 40.86
- 8d0.97
- Nd0.56
- TLD RAW0.60
-
- AlE_conferencepresentation2.tldr0.98
- The1.00
- Humans Still Beat AI in the Long Horizon:0.99
- Revisiting Test-Time Scaling in the Agent Era1.00
- June 16, 2026- Qiuyang Mang1, Kaiyuan Liu2, Bo Peng3, Shreyas Pimpalgaonkar4, Luke Zettlemoyer2, Alex Dimakis1.4,0.98
- Alvin Cheung10.99
- 1 UC Berkeley · 2 University of Washington - 3 Princeton University · 4 Bespoke Labs0.93
- 2026· # test-time-scaling # LLM-agents #scaling-lawsresearch0.97
- Discarded1.00
- Kept1.00
- Running best1.00
- Human vs Agent0.99
- top10-humans0.97
- 18531.00
- 18001.00
- fop50-humans0.98
- Cloude Code Opus-4.60.98
- Codex GPT-5.50.96
- humans keep climbing1.00
- for days1.00
- then plateau by 24h1.00
- 16001.00
- 15871.00
- Elokting0.75
- agents sprint early0.99
- 14000.95
- 12000.97
- d42-1370.90
- 10921.00
- 10001.00
- 2h1.00
- 5h0.71
- 8h1.00
- 16h0.96
- 24h0.75
- 2d1.00
- 4d0.72
- 8d0.96
- nd0.72
- time1.00
- TLDERRAW0.63
-
- AlE_conference presentation2.tldr0.94
- The bottleneck for Autonomous Scientific Agents is0.97
- Hypothesis Generation.1.00
- Observation0.92
- 18531.00
- Question1.00
- Conclusion1.00
- 15871.00
- Ansjysis0.64
- Exirnt0.63
- Hyhtsis0.78
- NLDERATW0.60
-
- AlE_conferencepresentation2.tldr0.97
- The bottleneck for Autonomous Scientific Agents is0.99
- Hypothesis Generation.1.00
- Observation0.96
- Question1.00
- Conclusion1.00
- Ansjvsis0.65
- Exirrnt0.67
- Hyptsis0.63
- Q0.70
- NLDERAW0.54
-
- AlE_conference presentation2.tldr0.95
- CT → In Silico PET0.95
- Coronal View1.00
- Diagnostic CT1.00
- In Silico PET1.00
- Lung Nodule1.00
- Decompose the problem into1.00
- SLDRAW0.55
-
- AlE_conferencepresentation2.tldr0.97
- CT → In Silico PET0.96
- Coronal View1.00
- Diagnostic CT1.00
- In Silico PET1.00
- Lung Nodule0.98
- 40.95
- AI0.79
- Decompose the problem into /g0.97
-
- AlE_conference blog_ct_to_insilico_pet.gi0.98
- aU0.51
-
- AlE_conferenceblog_ct_to_insilico_pet.gif0.99
- CT → In Silico PET0.98
- Diagnostic CT1.00
- In Silico PET1.00
- AI0.87
- Posterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif0.99
- CT → In Silico PET0.96
- Diagnostic CT0.99
- In SilicoPET1.00
- AI0.88
- Posterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif1.00
- CT → In Silico PET0.97
- Diagnostic CT1.00
- In Silico PET1.00
- Lung Nodule0.99
- Lung Nodule0.99
- AI0.93
- Posterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif1.00
- CT → In Silico PET0.95
- Diagnostic CT1.00
- In Silico PET1.00
- Lung Nodule0.95
- Lung Nodule0.99
- AI0.96
- Posterior1.00
-
- AlE_conference blog_ct_to_insilico_pet.gif0.99
- CT → In Silico PET0.99
- Diagnostic CT0.99
- In Silico PET1.00
- →1.00
- AI0.93
- Posterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif1.00
- CT → In Silico PET0.97
- Diagnostic CT1.00
- In Silico PET1.00
- Al0.83
- Posterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif1.00
- CT → In Silico PET0.97
- Diagnostic CT1.00
- In Silico PET1.00
- 10.72
- AI0.78
- sterior1.00
-
- AlE_conferenceblog_ct_to_insilico_pet.gif1.00
- CT → In Silico PET0.99
- Diagnostic CT0.98
- In Silico PET1.00
- AI0.87
- Posterior1.00
-
- AlE_conference presentation2.tldr0.94
- CT → In Silico PET0.99
- Coronal View1.00
- Diagnostic CT1.00
- In Silico PET1.00
- Lung No0.99
- Lung Nodule1.00
- AI0.75
- 40.99
- compose the problem into /goa0.99
-
- AlE_conference presentation2.tldr0.94
- Coronal View1.00
- Diagnostic CT1.00
- In Silico PET1.00
- Lung No1.00
- Lung Nodule1.00
- Al0.75
- 41.00
- ecompose the problem into /goa0.98
- /1000.91
-
- AlE_conference presentation2.tIdr0.94
- c Agents is1.00
- CT → In Silico PET0.99
- Coronal View0.98
- Diagnostic CT0.97
- In Silico PET1.00
- Observation1.00
- Question1.00
- Eer0.84
- Decompose the problem into /goal0.99
- /1oop0.99
- TLDRAW0.99
Transcript
121 cues· 2,878 words· 15,441 chars
- 0:00 Hello everyone, my name is Sina Shahandeh.
- 0:04 My pleasure to present you this talk about running autonomous agents for scientific tasks.
- 0:12 Let's dive in.
- 0:14 So here, I think everyone is quite familiar with the concept of auto researcher or auto research.
- 0:23 This is the original André Karpathis
- 0:27 GitHub repo where we have an ML model and we ask a coding agent to find a particular matrix and then optimize the code in order to minimize the error.
- 0:39 Basically, do a hell climb over model optimization.
- 0:45 Now, for many of these coding tasks, this works very well.
- 0:50 But when the problems become very much open-ended and sometimes long horizon, like most of the scientific tasks, you have this case where AI agents usually kind of saturate to a certain level.
- 1:06 simply they are very good at implementation of the of the code or changing the running the experiments over lots of data and so on but the problem is they ran out of ideas or you know what people call them research taste now you can see that in this situations you know good humans keep going higher and the top 1% humans you know they keep even improving better and better over time and
- 1:33 Now, the difference from here is the way that good ideas or good hypothesis on how the model could be improved or how the problem could be solved.
- 1:48 keep coming up, humans keep coming up.
- 1:50 So in the scientific task you have this scientific method where we observe a situation, we make a question, we come up with a good hypothesis and come up with a hypothesis on how to solve this problem and then we do experiment and implement and do the experiment, do the loop and each of the iterations we learn something and we improve.
- 2:13 Now
- 2:15 The components, the learning components, I think those are all questions of memory and implementation from learning the mistakes, which is one of the bottlenecks of using coding agents.
- 2:25 But I think this is quite solved by just simply organizing patterns of activity.
- 2:31 I think what is much more difficult is coming up with a hypothesis.
- 2:34 So how can we come up with a good hypothesis, good ideas,
- 2:37 for our coding agents to keep improving better the process and this is something core things that i would like to kind of focus on this talk and plus a little bonus at the end so let's look at the problem here we're trying to achieve as an example so you have a good idea of what we're trying to do so here is what we're doing at radicade we're building in silicopete meaning you have a city and we want to generate
- 3:06 PET image PET scan from the CT scan example this you can see here we have CT scans and of slices of scans of the patient and they might have a nodule in the lung and the question is is this cancerous or not is it a lung cancer so they do a PET scan which is a difficult process and very time-consuming and to do and
- 3:30 but here we do an ml to ml model image translation one can kind of change the modality learn the structure of the body and infer what would happen in a PET scan if the hyper the activity of the tissue so you know certain teachers absorb more radioactive tracer and they shine up in this PET scan and the tumors usually that's the case now to generate this relationship we need
- 3:59 we need a model to do the translation but the problem itself has many components so really this like any other scientific task the problem is decomposing that problem entire long-term horizon two years ten years research process into steps and each of those steps is really fundamentally are a goal or a loop
- 4:21 So I'm going to focus on one of these particular ones right now, and that is on training of machine learning model.
- 4:27 So we have here decoder encoder type of situations.
- 4:33 and so encoding the CT and then decoding it into PET.
- 4:38 So that's the typical GAN model, which kind of generates the image.
- 4:41 Now for this, we can kind of define these kind of, you know, the architecture and we try it, we'll capture data and do all the 80% of work to basically bringing the good data set.
- 4:52 And here we create the metrics and so on, and around the image, you know, fidelity of synthetic PET to real PET and so on.
- 5:03 so on but the challenge is how can we improve this situation given a certain initial point and we'll go back to our idea of hill climb around this optimization so you can see an example of iterations coming from a real run in codex
- 5:21 where we improve the data and so on, and then the model goes around and tries to do the optimization.
- 5:28 And you can see there's a range of possibilities, and some of them become dead end, some of them don't improve anything, but we kind of desaturate at a certain point.
- 5:38 And then you really need a good idea.
- 5:40 A good idea has to come up.
- 5:41 So in this case, we have slices of CT, and we feed these as a channel.
- 5:47 into the model so initial problem initial model that was trained was two and a half d so treating each ct slice as a it'd be the 2d convolutions but stack over channel now if you give this to a typical ml model it would not think about it as you know go through hyperparameters or you know some sort of you know playing around with problems that it knows but it wouldn't do a very radical change
- 6:13 for example, to come up with a 3D idea of compositions or change the whole problem upside down.
- 6:20 So to create those ideas for the model to try, I had to kind of
- 6:29 in the midst of the codex loop say, what about this idea?
- 6:32 What about that idea?
- 6:33 Go read papers out there and see what the papers are saying, what other people are trying.
- 6:39 To induce that hypothesis generation, we need to do something about our ML model, LLM models.
- 6:49 this is a trick that i found that working very efficiently and it's very similar to that chain of thoughts step-by-step problem first is to decompose the problem into its subcomponents but it's an explicit act um action so you know you can ask it actually go through your problem here in this case you know create a translated polymer nodule ct patches into equivalent pet
- 7:16 that's our top problem that I just explained to you and then it has components into it so different domain in this case you can see you know the the data component the
- 7:28 actual core, the learning, the architecture, the training laws, the operational part of the ML modeling, the metrics and evidence, and the peripheral scripts that kind of run the model.
- 7:43 And data preparation itself is very important pieces.
- 7:45 Now what we have here is this hierarchy of components of this model that's induced.
- 7:52 So this itself can be induced very easily using a prompt.
- 7:56 So basically,
- 8:00 a coding agent can itself go in with this prompt of going through this code base and create this series of hyper documents that are linked to each other.
- 8:10 Now what we're trying to do is to give our coding agents ability to look at this problem as component where it might not do so and then induce
loading