read-only demo

Videos XLEYtv3cMlw

Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait

index_state ready data_status ok

AI Engineer· published 2026-07-18· 0:19:22· en-US· indexed 2026-08-11 00:07

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:14, 1 of 1 keyframes kept
  2. Shot 1, 0:14 to 0:19, 1 of 1 keyframes kept
  3. Shot 2, 0:19 to 0:46, 1 of 1 keyframes kept
  4. Shot 3, 0:46 to 0:47, 1 of 1 keyframes kept
  5. Shot 4, 0:47 to 0:50, 1 of 1 keyframes kept
  6. Shot 5, 0:50 to 1:20, 1 of 1 keyframes kept
  7. Shot 6, 1:20 to 1:51, 0 of 1 keyframes kept
  8. Shot 7, 1:51 to 2:20, 1 of 1 keyframes kept
  9. Shot 8, 2:20 to 2:48, 0 of 1 keyframes kept
  10. Shot 9, 2:48 to 2:52, 1 of 1 keyframes kept
  11. Shot 10, 2:52 to 2:53, 1 of 1 keyframes kept
  12. Shot 11, 2:53 to 3:07, 1 of 1 keyframes kept
  13. Shot 12, 3:07 to 3:09, 1 of 1 keyframes kept
  14. Shot 13, 3:09 to 3:19, 1 of 1 keyframes kept
  15. Shot 14, 3:19 to 3:24, 1 of 1 keyframes kept
  16. Shot 15, 3:24 to 3:25, 1 of 1 keyframes kept
  17. Shot 16, 3:25 to 3:30, 1 of 1 keyframes kept
  18. Shot 17, 3:30 to 3:46, 1 of 1 keyframes kept
  19. Shot 18, 3:46 to 3:49, 1 of 1 keyframes kept
  20. Shot 19, 3:49 to 3:51, 1 of 1 keyframes kept
  21. Shot 20, 3:51 to 3:58, 1 of 1 keyframes kept
  22. Shot 21, 3:58 to 4:00, 1 of 1 keyframes kept
  23. Shot 22, 4:00 to 4:02, 1 of 1 keyframes kept
  24. Shot 23, 4:02 to 4:06, 1 of 1 keyframes kept
  25. Shot 24, 4:06 to 4:20, 1 of 1 keyframes kept
  26. Shot 25, 4:20 to 4:23, 1 of 1 keyframes kept
  27. Shot 26, 4:23 to 4:24, 1 of 1 keyframes kept
  28. Shot 27, 4:24 to 4:26, 1 of 1 keyframes kept
  29. Shot 28, 4:26 to 4:39, 1 of 1 keyframes kept
  30. Shot 29, 4:39 to 4:41, 1 of 1 keyframes kept
  31. Shot 30, 4:41 to 4:42, 1 of 1 keyframes kept
  32. Shot 31, 4:42 to 5:12, 1 of 1 keyframes kept
  33. Shot 32, 5:12 to 5:13, 1 of 1 keyframes kept
  34. Shot 33, 5:13 to 5:15, 1 of 1 keyframes kept
  35. Shot 34, 5:15 to 5:27, 1 of 1 keyframes kept
  36. Shot 35, 5:27 to 5:44, 1 of 1 keyframes kept
  37. Shot 36, 5:44 to 6:12, 1 of 1 keyframes kept
  38. Shot 37, 6:12 to 6:13, 1 of 1 keyframes kept
  39. Shot 38, 6:13 to 6:42, 1 of 1 keyframes kept
  40. Shot 39, 6:42 to 6:45, 1 of 1 keyframes kept
  41. Shot 40, 6:45 to 6:46, 1 of 1 keyframes kept
  42. Shot 41, 6:46 to 6:47, 1 of 1 keyframes kept
  43. Shot 42, 6:47 to 6:55, 1 of 1 keyframes kept
  44. Shot 43, 6:55 to 6:58, 1 of 1 keyframes kept
  45. Shot 44, 6:58 to 6:59, 1 of 1 keyframes kept
  46. Shot 45, 6:59 to 7:09, 1 of 1 keyframes kept
  47. Shot 46, 7:09 to 7:54, 1 of 1 keyframes kept
  48. Shot 47, 7:54 to 8:34, 1 of 1 keyframes kept
  49. Shot 48, 8:34 to 8:44, 1 of 1 keyframes kept
  50. Shot 49, 8:44 to 8:45, 1 of 1 keyframes kept
  51. Shot 50, 8:45 to 8:48, 1 of 1 keyframes kept
  52. Shot 51, 8:48 to 8:50, 1 of 1 keyframes kept
  53. Shot 52, 8:50 to 8:57, 1 of 1 keyframes kept
  54. Shot 53, 8:57 to 9:08, 1 of 1 keyframes kept
  55. Shot 54, 9:08 to 9:09, 1 of 1 keyframes kept
  56. Shot 55, 9:09 to 9:11, 1 of 1 keyframes kept
  57. Shot 56, 9:11 to 9:14, 1 of 1 keyframes kept
  58. Shot 57, 9:14 to 9:27, 1 of 1 keyframes kept
  59. Shot 58, 9:27 to 9:28, 1 of 1 keyframes kept
  60. Shot 59, 9:28 to 9:42, 1 of 1 keyframes kept
  61. Shot 60, 9:42 to 9:45, 1 of 1 keyframes kept
  62. Shot 61, 9:45 to 9:52, 1 of 1 keyframes kept
  63. Shot 62, 9:52 to 9:54, 0 of 1 keyframes kept
  64. Shot 63, 9:54 to 9:56, 1 of 1 keyframes kept
  65. Shot 64, 9:56 to 10:32, 1 of 1 keyframes kept
  66. Shot 65, 10:32 to 10:39, 1 of 1 keyframes kept
  67. Shot 66, 10:39 to 10:42, 1 of 1 keyframes kept
  68. Shot 67, 10:42 to 10:49, 1 of 1 keyframes kept
  69. Shot 68, 10:49 to 10:51, 1 of 1 keyframes kept
  70. Shot 69, 10:51 to 11:08, 1 of 1 keyframes kept
  71. Shot 70, 11:08 to 11:11, 1 of 1 keyframes kept
  72. Shot 71, 11:11 to 11:29, 1 of 1 keyframes kept
  73. Shot 72, 11:29 to 11:31, 0 of 1 keyframes kept
  74. Shot 73, 11:31 to 11:40, 0 of 1 keyframes kept
  75. Shot 74, 11:40 to 11:44, 1 of 1 keyframes kept
  76. Shot 75, 11:44 to 11:46, 1 of 1 keyframes kept
  77. Shot 76, 11:46 to 11:48, 1 of 1 keyframes kept
  78. Shot 77, 11:48 to 12:09, 1 of 1 keyframes kept
  79. Shot 78, 12:09 to 12:11, 1 of 1 keyframes kept
  80. Shot 79, 12:11 to 12:19, 1 of 1 keyframes kept
  81. Shot 80, 12:19 to 12:22, 1 of 1 keyframes kept
  82. Shot 81, 12:22 to 12:44, 1 of 1 keyframes kept
  83. Shot 82, 12:44 to 13:12, 1 of 1 keyframes kept
  84. Shot 83, 13:12 to 13:54, 1 of 1 keyframes kept
  85. Shot 84, 13:54 to 13:56, 1 of 1 keyframes kept
  86. Shot 85, 13:56 to 13:59, 1 of 1 keyframes kept
  87. Shot 86, 13:59 to 14:00, 1 of 1 keyframes kept
  88. Shot 87, 14:00 to 14:01, 1 of 1 keyframes kept
  89. Shot 88, 14:01 to 14:04, 1 of 1 keyframes kept
  90. Shot 89, 14:04 to 14:34, 1 of 1 keyframes kept
  91. Shot 90, 14:34 to 14:38, 1 of 1 keyframes kept
  92. Shot 91, 14:38 to 14:40, 1 of 1 keyframes kept
  93. Shot 92, 14:40 to 14:42, 1 of 1 keyframes kept
  94. Shot 93, 14:42 to 14:49, 1 of 1 keyframes kept
  95. Shot 94, 14:49 to 14:51, 1 of 1 keyframes kept
  96. Shot 95, 14:51 to 14:53, 1 of 1 keyframes kept
  97. Shot 96, 14:53 to 15:31, 1 of 1 keyframes kept
  98. Shot 97, 15:31 to 15:35, 1 of 1 keyframes kept
  99. Shot 98, 15:35 to 15:36, 1 of 1 keyframes kept
  100. Shot 99, 15:36 to 15:41, 1 of 1 keyframes kept
  101. Shot 100, 15:41 to 16:18, 1 of 1 keyframes kept
  102. Shot 101, 16:18 to 16:21, 1 of 1 keyframes kept
  103. Shot 102, 16:21 to 16:43, 1 of 1 keyframes kept
  104. Shot 103, 16:43 to 16:46, 1 of 1 keyframes kept
  105. Shot 104, 16:46 to 16:50, 1 of 1 keyframes kept
  106. Shot 105, 16:50 to 17:24, 1 of 1 keyframes kept
  107. Shot 106, 17:24 to 17:28, 1 of 1 keyframes kept
  108. Shot 107, 17:28 to 17:30, 1 of 1 keyframes kept
  109. Shot 108, 17:30 to 17:54, 1 of 1 keyframes kept
  110. Shot 109, 17:54 to 18:03, 1 of 1 keyframes kept
  111. Shot 110, 18:03 to 18:04, 1 of 1 keyframes kept
  112. Shot 111, 18:04 to 18:30, 1 of 1 keyframes kept
  113. Shot 112, 18:30 to 18:33, 1 of 1 keyframes kept
  114. Shot 113, 18:33 to 19:22, 1 of 1 keyframes kept

114 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
121
whisperx 121
chunks
35
from 121 cues
keyframes
109
kept of 114 captured
frames with text
109
3,343 lines read
chapters
0
from the source metadata
keyframe bytes
15.6 MB
word timings on 121 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 00:02 1m 29s
stt done 2026-08-11 00:03 19s
chunk done 2026-08-11 00:03 0s
text_embed done 2026-08-11 00:03 1s
keyframe done 2026-08-11 00:03 1m 26s
ocr done 2026-08-11 00:05 1m 21s
frame_embed done 2026-08-11 00:06 22s

Frames, and what the machine read

  • 0:08 #0 done28 line(s)

    shot 0·sharpness 1719.2

    1. AlE_conference presentation2.tldr0.97
    2. Autonomous Agents for Scientific Tasks1.00
    3. autor1.00
    4. Sina Shahandeh,1.00
    5. Radicait1.00
    6. 1.0001.00
    7. Abstract:1.00
    8. 0.9951.00
    9. There has been much w0.99
    10. problems, or static supe0.99
    11. discovery task, the prob0.97
    12. very large and complex0.99
    13. BPB (ower0.65
    14. 0.9901.00
    15. physical model, a meas0.98
    16. clinic, observatory, or fi0.98
    17. 0.9851.00
    18. In this talk, we show sce1.00
    19. classes, and hyperpara0.98
    20. performance comes fro1.00
    21. behaves, implementing0.98
    22. 0.9801.00
    23. real data.1.00
    24. We show how an ontolo0.98
    25. generation that is key to1.00
    26. 0.9750.98
    27. industrial and applied r0.97
    28. TLDERRAW0.61
  • 0:18 #1 done29 line(s)

    shot 1·sharpness 4339.7

    1. AlE_conference presentation2.tldr0.96
    2. Autonomous Agents for Scientific Tasks0.97
    3. autor1.00
    4. Sina Shahandeh,1.00
    5. Radicait1.00
    6. 1.0000.96
    7. Abstract:1.00
    8. 0.9950.92
    9. There has been much work on Autoresearch where the objectives are coding puzzles, toy optimization1.00
    10. problems, or static supervised-learning ML tasks. However, for an autonomous agent to assist with a scientific1.00
    11. (ette sr)0.51
    12. discovery task, the problems must come from real measurement data of the world, the search space must be1.00
    13. very large and complex, and the objective must have scientific meaning. A better solution would improve a1.00
    14. Val g (er0.74
    15. 0.9901.00
    16. physical model, a measurement or characterization method, or experimental decision-making in an actual lab,1.00
    17. clinic, observatory, or field setting.1.00
    18. 0.9851.00
    19. In this talk, we show scenarios where the agent has to search over methods, priors, data preprocessing, model1.00
    20. classes, and hyperparameters while learning from intermediate failures. Often, a step change in agent0.99
    21. behaves, implementing that hypothesis correctly in a mathematical model, and executing it on the existing0.99
    22. performance comes from forming an appropriate scientific hypothesis about how the physical system0.99
    23. 0.9801.00
    24. real data.1.00
    25. We show how an ontology-based memory system is used in the harness to assist with the hypothesis0.99
    26. generation that is key to the agent's success. All demonstrations come from real scientific problems solved in0.99
    27. 0.9751.00
    28. industrial and applied research settings.0.99
    29. MLDERWAW0.57
  • 0:35 #2 done26 line(s)

    shot 2·sharpness 818.4

    1. AlE_conference presentation2.tldr0.95
    2. Humar1.00
    3. Revisit1.00
    4. autoresearch1.00
    5. June 16,2026·Q0.96
    6. Alvin Cheung1.00
    7. 'UC Berkeley.20.92
    8. 2026#test0.98
    9. Autoresearch Progress: 83 Experiments, 15 Kept Improvements1.00
    10. 1.0000.90
    11. Discarded0.97
    12. Running best0.99
    13. 0.9950.91
    14. V(al s eter0.59
    15. 0.9901.00
    16. Elorting0.63
    17. 0.9851.00
    18. 0.9801.00
    19. 0.9751.00
    20. 00.99
    21. 201.00
    22. Experiment #1.00
    23. 401.00
    24. 601.00
    25. 801.00
    26. TLDRAW0.97
  • 0:46 #3 done42 line(s)

    shot 3·sharpness 1639.5

    1. AlE_conferencepresentation2.tldr0.98
    2. Humans Still Beat AI in the Long Horizon:1.00
    3. Revisiting Test-Time Scaling in the Agent Era1.00
    4. June 16, 2026- Qiuyang Mang1, Kaiyuan Liu2, Bo Peng3. hreyas Pimpalgaonkar4, Luke ZEettlemoyer2, Alex Dimakis1.4,0.96
    5. Alvin Cheung1.00
    6. UC Berkeley - 2 University of Washington - 3 Princeton Universit - Bespoke Labs0.95
    7. 2026#test-time-scaling # LLM-agents # scaling-lawsresearch0.97
    8. ss: 83 Experiments, 15 Kept Improvements0.99
    9. Running best1.00
    10. Kept0.99
    11. Discarded1.00
    12. Human vs Agent0.99
    13. fop10-humans0.95
    14. 18530.96
    15. 18001.00
    16. fop50-humans0.99
    17. Cloude Code Opus-4.60.93
    18. Codex GPT-5.50.96
    19. humons keep climbing0.99
    20. for days1.00
    21. then plafeou by 24h0.99
    22. 1%000.87
    23. 15870.99
    24. Eotng0.70
    25. ogents sprint early0.98
    26. 1000.97
    27. 12000.99
    28. 10921.00
    29. 10000.97
    30. Experiment #0.98
    31. 401.00
    32. 601.00
    33. 801.00
    34. 2h0.90
    35. 8h0.90
    36. th0.91
    37. time1.00
    38. 2h0.97
    39. 2d0.98
    40. 40.63
    41. 8d0.72
    42. TLD RRAW0.59
  • 0:49 #4 done40 line(s)

    shot 4·sharpness 1587.1

    1. AlE_conferencepresentation2.tldr0.98
    2. The bottlene1.00
    3. Humans Still Beat AI in the Long Horizon:0.99
    4. Revisiting Test-Time Scaling in the Agent Era1.00
    5. June 16, 2026- Qiuyang Mang,. Kaiyuan Lu2, Bo Peng3. Shreyas Pimpalgaonkar4, Luke Zettlemoyer2, Alex Dimakis14,0.93
    6. Alvin Cheung1.00
    7. Hypotl0.96
    8. 1 UC Berkeley - 2 University of Washington · 3 Princeton University · 4 Bespoke Labs0.92
    9. 2026 #test-time-scaling # LLM-agents # scaling-lawsresearch0.96
    10. Discarded1.00
    11. Running best0.97
    12. Kept1.00
    13. Human vs Agent0.98
    14. top10-humons0.96
    15. 18531.00
    16. 18001.00
    17. Cloude Code Opus-4.60.99
    18. Codex GPT-5.51.00
    19. top50-humans0.98
    20. humans keep climbing0.97
    21. for days0.99
    22. then plateou by 24h0.98
    23. 16000.88
    24. 15871.00
    25. Eotng0.74
    26. agents sprint eorly0.96
    27. 1000.90
    28. 13681.00
    29. 12000.95
    30. 10921.00
    31. 10001.00
    32. 2h0.78
    33. 18h0.73
    34. time1.00
    35. 24h0.84
    36. 2d0.93
    37. 40.86
    38. 8d0.97
    39. Nd0.56
    40. TLD RAW0.60
  • 1:11 #5 done41 line(s)

    shot 5·sharpness 1913.0

    1. AlE_conferencepresentation2.tldr0.98
    2. The1.00
    3. Humans Still Beat AI in the Long Horizon:0.99
    4. Revisiting Test-Time Scaling in the Agent Era1.00
    5. June 16, 2026- Qiuyang Mang1, Kaiyuan Liu2, Bo Peng3, Shreyas Pimpalgaonkar4, Luke Zettlemoyer2, Alex Dimakis1.4,0.98
    6. Alvin Cheung10.99
    7. 1 UC Berkeley · 2 University of Washington - 3 Princeton University · 4 Bespoke Labs0.93
    8. 2026· # test-time-scaling # LLM-agents #scaling-lawsresearch0.97
    9. Discarded1.00
    10. Kept1.00
    11. Running best1.00
    12. Human vs Agent0.99
    13. top10-humans0.97
    14. 18531.00
    15. 18001.00
    16. fop50-humans0.98
    17. Cloude Code Opus-4.60.98
    18. Codex GPT-5.50.96
    19. humans keep climbing1.00
    20. for days1.00
    21. then plateau by 24h1.00
    22. 16001.00
    23. 15871.00
    24. Elokting0.75
    25. agents sprint early0.99
    26. 14000.95
    27. 12000.97
    28. d42-1370.90
    29. 10921.00
    30. 10001.00
    31. 2h1.00
    32. 5h0.71
    33. 8h1.00
    34. 16h0.96
    35. 24h0.75
    36. 2d1.00
    37. 4d0.72
    38. 8d0.96
    39. nd0.72
    40. time1.00
    41. TLDERRAW0.63
  • 1:24 #6 skipped

    shot 6·duplicate of #5

  • 2:02 #7 done12 line(s)

    shot 7·sharpness 1034.8

    1. AlE_conference presentation2.tldr0.94
    2. The bottleneck for Autonomous Scientific Agents is0.97
    3. Hypothesis Generation.1.00
    4. Observation0.92
    5. 18531.00
    6. Question1.00
    7. Conclusion1.00
    8. 15871.00
    9. Ansjysis0.64
    10. Exirnt0.63
    11. Hyhtsis0.78
    12. NLDERATW0.60
  • 2:23 #8 skipped

    shot 8·duplicate of #7

  • 2:49 #9 done11 line(s)

    shot 9·sharpness 1029.4

    1. AlE_conferencepresentation2.tldr0.97
    2. The bottleneck for Autonomous Scientific Agents is0.99
    3. Hypothesis Generation.1.00
    4. Observation0.96
    5. Question1.00
    6. Conclusion1.00
    7. Ansjvsis0.65
    8. Exirrnt0.67
    9. Hyptsis0.63
    10. Q0.70
    11. NLDERAW0.54
  • 2:52 #10 done8 line(s)

    shot 10·sharpness 1117.5

    1. AlE_conference presentation2.tldr0.95
    2. CT → In Silico PET0.95
    3. Coronal View1.00
    4. Diagnostic CT1.00
    5. In Silico PET1.00
    6. Lung Nodule1.00
    7. Decompose the problem into1.00
    8. SLDRAW0.55
  • 3:03 #11 done9 line(s)

    shot 11·sharpness 1013.3

    1. AlE_conferencepresentation2.tldr0.97
    2. CT → In Silico PET0.96
    3. Coronal View1.00
    4. Diagnostic CT1.00
    5. In Silico PET1.00
    6. Lung Nodule0.98
    7. 40.95
    8. AI0.79
    9. Decompose the problem into /g0.97
  • 3:08 #12 done2 line(s)

    shot 12·sharpness 80.1

    1. AlE_conference blog_ct_to_insilico_pet.gi0.98
    2. aU0.51
  • 3:17 #13 done6 line(s)

    shot 13·sharpness 306.2

    1. AlE_conferenceblog_ct_to_insilico_pet.gif0.99
    2. CT → In Silico PET0.98
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. AI0.87
    6. Posterior1.00
  • 3:23 #14 done6 line(s)

    shot 14·sharpness 440.1

    1. AlE_conferenceblog_ct_to_insilico_pet.gif0.99
    2. CT → In Silico PET0.96
    3. Diagnostic CT0.99
    4. In SilicoPET1.00
    5. AI0.88
    6. Posterior1.00
  • 3:24 #15 done8 line(s)

    shot 15·sharpness 504.3

    1. AlE_conferenceblog_ct_to_insilico_pet.gif1.00
    2. CT → In Silico PET0.97
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. Lung Nodule0.99
    6. Lung Nodule0.99
    7. AI0.93
    8. Posterior1.00
  • 3:27 #16 done8 line(s)

    shot 16·sharpness 507.3

    1. AlE_conferenceblog_ct_to_insilico_pet.gif1.00
    2. CT → In Silico PET0.95
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. Lung Nodule0.95
    6. Lung Nodule0.99
    7. AI0.96
    8. Posterior1.00
  • 3:32 #17 done7 line(s)

    shot 17·sharpness 471.2

    1. AlE_conference blog_ct_to_insilico_pet.gif0.99
    2. CT → In Silico PET0.99
    3. Diagnostic CT0.99
    4. In Silico PET1.00
    5. 1.00
    6. AI0.93
    7. Posterior1.00
  • 3:48 #18 done6 line(s)

    shot 18·sharpness 298.3

    1. AlE_conferenceblog_ct_to_insilico_pet.gif1.00
    2. CT → In Silico PET0.97
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. Al0.83
    6. Posterior1.00
  • 3:50 #19 done7 line(s)

    shot 19·sharpness 359.4

    1. AlE_conferenceblog_ct_to_insilico_pet.gif1.00
    2. CT → In Silico PET0.97
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. 10.72
    6. AI0.78
    7. sterior1.00
  • 3:55 #20 done6 line(s)

    shot 20·sharpness 288.3

    1. AlE_conferenceblog_ct_to_insilico_pet.gif1.00
    2. CT → In Silico PET0.99
    3. Diagnostic CT0.98
    4. In Silico PET1.00
    5. AI0.87
    6. Posterior1.00
  • 4:00 #21 done10 line(s)

    shot 21·sharpness 1129.8

    1. AlE_conference presentation2.tldr0.94
    2. CT → In Silico PET0.99
    3. Coronal View1.00
    4. Diagnostic CT1.00
    5. In Silico PET1.00
    6. Lung No0.99
    7. Lung Nodule1.00
    8. AI0.75
    9. 40.99
    10. compose the problem into /goa0.99
  • 4:02 #22 done10 line(s)

    shot 22·sharpness 1107.6

    1. AlE_conference presentation2.tldr0.94
    2. Coronal View1.00
    3. Diagnostic CT1.00
    4. In Silico PET1.00
    5. Lung No1.00
    6. Lung Nodule1.00
    7. Al0.75
    8. 41.00
    9. ecompose the problem into /goa0.98
    10. /1000.91
  • 4:05 #23 done12 line(s)

    shot 23·sharpness 1216.3

    1. AlE_conference presentation2.tIdr0.94
    2. c Agents is1.00
    3. CT → In Silico PET0.99
    4. Coronal View0.98
    5. Diagnostic CT0.97
    6. In Silico PET1.00
    7. Observation1.00
    8. Question1.00
    9. Eer0.84
    10. Decompose the problem into /goal0.99
    11. /1oop0.99
    12. TLDRAW0.99

Transcript

121 cues· 2,878 words· 15,441 chars

  1. 0:00 Hello everyone, my name is Sina Shahandeh.
  2. 0:04 My pleasure to present you this talk about running autonomous agents for scientific tasks.
  3. 0:12 Let's dive in.
  4. 0:14 So here, I think everyone is quite familiar with the concept of auto researcher or auto research.
  5. 0:23 This is the original André Karpathis
  6. 0:27 GitHub repo where we have an ML model and we ask a coding agent to find a particular matrix and then optimize the code in order to minimize the error.
  7. 0:39 Basically, do a hell climb over model optimization.
  8. 0:45 Now, for many of these coding tasks, this works very well.
  9. 0:50 But when the problems become very much open-ended and sometimes long horizon, like most of the scientific tasks, you have this case where AI agents usually kind of saturate to a certain level.
  10. 1:06 simply they are very good at implementation of the of the code or changing the running the experiments over lots of data and so on but the problem is they ran out of ideas or you know what people call them research taste now you can see that in this situations you know good humans keep going higher and the top 1% humans you know they keep even improving better and better over time and
  11. 1:33 Now, the difference from here is the way that good ideas or good hypothesis on how the model could be improved or how the problem could be solved.
  12. 1:48 keep coming up, humans keep coming up.
  13. 1:50 So in the scientific task you have this scientific method where we observe a situation, we make a question, we come up with a good hypothesis and come up with a hypothesis on how to solve this problem and then we do experiment and implement and do the experiment, do the loop and each of the iterations we learn something and we improve.
  14. 2:13 Now
  15. 2:15 The components, the learning components, I think those are all questions of memory and implementation from learning the mistakes, which is one of the bottlenecks of using coding agents.
  16. 2:25 But I think this is quite solved by just simply organizing patterns of activity.
  17. 2:31 I think what is much more difficult is coming up with a hypothesis.
  18. 2:34 So how can we come up with a good hypothesis, good ideas,
  19. 2:37 for our coding agents to keep improving better the process and this is something core things that i would like to kind of focus on this talk and plus a little bonus at the end so let's look at the problem here we're trying to achieve as an example so you have a good idea of what we're trying to do so here is what we're doing at radicade we're building in silicopete meaning you have a city and we want to generate
  20. 3:06 PET image PET scan from the CT scan example this you can see here we have CT scans and of slices of scans of the patient and they might have a nodule in the lung and the question is is this cancerous or not is it a lung cancer so they do a PET scan which is a difficult process and very time-consuming and to do and
  21. 3:30 but here we do an ml to ml model image translation one can kind of change the modality learn the structure of the body and infer what would happen in a PET scan if the hyper the activity of the tissue so you know certain teachers absorb more radioactive tracer and they shine up in this PET scan and the tumors usually that's the case now to generate this relationship we need
  22. 3:59 we need a model to do the translation but the problem itself has many components so really this like any other scientific task the problem is decomposing that problem entire long-term horizon two years ten years research process into steps and each of those steps is really fundamentally are a goal or a loop
  23. 4:21 So I'm going to focus on one of these particular ones right now, and that is on training of machine learning model.
  24. 4:27 So we have here decoder encoder type of situations.
  25. 4:33 and so encoding the CT and then decoding it into PET.
  26. 4:38 So that's the typical GAN model, which kind of generates the image.
  27. 4:41 Now for this, we can kind of define these kind of, you know, the architecture and we try it, we'll capture data and do all the 80% of work to basically bringing the good data set.
  28. 4:52 And here we create the metrics and so on, and around the image, you know, fidelity of synthetic PET to real PET and so on.
  29. 5:03 so on but the challenge is how can we improve this situation given a certain initial point and we'll go back to our idea of hill climb around this optimization so you can see an example of iterations coming from a real run in codex
  30. 5:21 where we improve the data and so on, and then the model goes around and tries to do the optimization.
  31. 5:28 And you can see there's a range of possibilities, and some of them become dead end, some of them don't improve anything, but we kind of desaturate at a certain point.
  32. 5:38 And then you really need a good idea.
  33. 5:40 A good idea has to come up.
  34. 5:41 So in this case, we have slices of CT, and we feed these as a channel.
  35. 5:47 into the model so initial problem initial model that was trained was two and a half d so treating each ct slice as a it'd be the 2d convolutions but stack over channel now if you give this to a typical ml model it would not think about it as you know go through hyperparameters or you know some sort of you know playing around with problems that it knows but it wouldn't do a very radical change
  36. 6:13 for example, to come up with a 3D idea of compositions or change the whole problem upside down.
  37. 6:20 So to create those ideas for the model to try, I had to kind of
  38. 6:29 in the midst of the codex loop say, what about this idea?
  39. 6:32 What about that idea?
  40. 6:33 Go read papers out there and see what the papers are saying, what other people are trying.
  41. 6:39 To induce that hypothesis generation, we need to do something about our ML model, LLM models.
  42. 6:49 this is a trick that i found that working very efficiently and it's very similar to that chain of thoughts step-by-step problem first is to decompose the problem into its subcomponents but it's an explicit act um action so you know you can ask it actually go through your problem here in this case you know create a translated polymer nodule ct patches into equivalent pet
  43. 7:16 that's our top problem that I just explained to you and then it has components into it so different domain in this case you can see you know the the data component the
  44. 7:28 actual core, the learning, the architecture, the training laws, the operational part of the ML modeling, the metrics and evidence, and the peripheral scripts that kind of run the model.
  45. 7:43 And data preparation itself is very important pieces.
  46. 7:45 Now what we have here is this hierarchy of components of this model that's induced.
  47. 7:52 So this itself can be induced very easily using a prompt.
  48. 7:56 So basically,
  49. 8:00 a coding agent can itself go in with this prompt of going through this code base and create this series of hyper documents that are linked to each other.
  50. 8:10 Now what we're trying to do is to give our coding agents ability to look at this problem as component where it might not do so and then induce

Open at this second