read-only demo

Videos iCj_ATyThvc

How Autoresearch is changing ML research — Zhengyao Jiang, Weco

index_state ready data_status ok

AI Engineer· published 2026-07-16· 0:16:16· en-US· indexed 2026-08-11 02:48

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 1:06, 1 of 1 keyframes kept
  6. Shot 5, 1:06 to 1:35, 1 of 1 keyframes kept
  7. Shot 6, 1:35 to 2:10, 1 of 1 keyframes kept
  8. Shot 7, 2:10 to 2:45, 0 of 1 keyframes kept
  9. Shot 8, 2:45 to 3:05, 1 of 1 keyframes kept
  10. Shot 9, 3:05 to 3:32, 1 of 1 keyframes kept
  11. Shot 10, 3:32 to 3:59, 0 of 1 keyframes kept
  12. Shot 11, 3:59 to 4:23, 1 of 1 keyframes kept
  13. Shot 12, 4:23 to 4:51, 1 of 1 keyframes kept
  14. Shot 13, 4:51 to 5:19, 0 of 1 keyframes kept
  15. Shot 14, 5:19 to 5:47, 1 of 1 keyframes kept
  16. Shot 15, 5:47 to 6:15, 0 of 1 keyframes kept
  17. Shot 16, 6:15 to 6:43, 1 of 1 keyframes kept
  18. Shot 17, 6:43 to 7:11, 1 of 1 keyframes kept
  19. Shot 18, 7:11 to 7:39, 0 of 1 keyframes kept
  20. Shot 19, 7:39 to 8:09, 1 of 1 keyframes kept
  21. Shot 20, 8:09 to 8:39, 0 of 1 keyframes kept
  22. Shot 21, 8:39 to 9:06, 1 of 1 keyframes kept
  23. Shot 22, 9:06 to 9:34, 1 of 1 keyframes kept
  24. Shot 23, 9:34 to 10:03, 0 of 1 keyframes kept
  25. Shot 24, 10:03 to 10:36, 0 of 1 keyframes kept
  26. Shot 25, 10:36 to 11:08, 1 of 1 keyframes kept
  27. Shot 26, 11:08 to 11:41, 1 of 1 keyframes kept
  28. Shot 27, 11:41 to 12:10, 0 of 1 keyframes kept
  29. Shot 28, 12:10 to 12:39, 1 of 1 keyframes kept
  30. Shot 29, 12:39 to 13:08, 1 of 1 keyframes kept
  31. Shot 30, 13:08 to 13:37, 0 of 1 keyframes kept
  32. Shot 31, 13:37 to 14:06, 0 of 1 keyframes kept
  33. Shot 32, 14:06 to 14:35, 0 of 1 keyframes kept
  34. Shot 33, 14:35 to 15:03, 1 of 1 keyframes kept
  35. Shot 34, 15:03 to 15:31, 0 of 1 keyframes kept
  36. Shot 35, 15:31 to 15:59, 1 of 1 keyframes kept
  37. Shot 36, 15:59 to 16:15, 0 of 1 keyframes kept

37 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
149
whisperx 149
chunks
29
from 149 cues
keyframes
23
kept of 37 captured
frames with text
23
497 lines read
chapters
13
from the source metadata
keyframe bytes
4.6 MB
word timings on 149 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 02:44 1m 12s
stt done 2026-08-11 02:45 15s
chunk done 2026-08-11 02:46 0s
text_embed done 2026-08-11 02:46 0s
keyframe done 2026-08-11 02:46 1m 39s
ocr done 2026-08-11 02:47 11s
frame_embed done 2026-08-11 02:48 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 456.7

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 663.3

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2742.3

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.93
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.90
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:27 #3 done17 line(s)

    shot 3·sharpness 2715.5

    1. AlEngineer0.98
    2. World'sFair1.00
    3. This year, OpenAl ran a hiring challenge, a competition called1.00
    4. Parameter Golf.1.00
    5. PRESENTED BY0.97
    6. Microsoft1.00
    7. The top contributor was the one candidate they couldn't hire.1.00
    8. It wasn't a person. It's an agent we built, Aiden.0.98
    9. d'sFair1.00
    10. Opē0.87
    11. AlEngir0.95
    12. icrosoft1.00
    13. World'0.94
    14. d'sF0.93
    15. sny0.99
    16. Engineering the future of Al0.98
    17. inhtr0.89
  • 1:01 #4 done37 line(s)

    shot 4·sharpness 2359.7

    1. AlEngineer0.97
    2. OPENAI· MODEL CRAFT CHALLENGE0.97
    3. World'sFair1.00
    4. Parameter Golf0.99
    5. Code golf for neural networks. Train the best language model you can, then shrink it.0.99
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. 16 MB0.99
    9. 10 min1.00
    10. BPB↓0.98
    11. Total artifact1.00
    12. Training budget1.00
    13. Lower is better1.00
    14. Model weights + training code, combined.1.00
    15. Wall-clock cap on 8×H100 GPUs.1.00
    16. Bits-per-byte on the FineWeb val set1.00
    17. (tokenizer-agnostic).1.00
    18. Fair1.00
    19. OpenA0.97
    20. BY THE NUMBERS0.99
    21. THE FIELD,1.00
    22. 1,0161.00
    23. 2,0481.00
    24. 471.00
    25. $249,5501.00
    26. participants1.00
    27. PRs filed0.96
    28. merged records · 31 contributors0.97
    29. RunPod credits bumed0.99
    30. AlEngineer1.00
    31. soft0.99
    32. lorld's Fa0.95
    33. Mar 18–Apr 30, 20260.98
    34. Fair0.97
    35. yk0.84
    36. Engineering the future of Al0.99
    37. tuiin0.51
  • 1:15 #5 done16 line(s)

    shot 5·sharpness 2231.0

    1. AlEngineer0.95
    2. World'sFair1.00
    3. You've seen a lot of autoresearch today, agents hill-climbing benchmarks.0.99
    4. PRESENTED BY0.96
    5. Can an autoresearch agent produce work0.99
    6. Microsoft0.96
    7. a community recognizes?1.00
    8. Not a good score on your own machine. Something other engineers will merge, fork,1.00
    9. and build on.0.97
    10. Fair1.00
    11. OpenA0.94
    12. AlEngineer0.99
    13. osoft0.98
    14. World'sF0.99
    15. Fair0.99
    16. Engineering the future of Al0.99
  • 2:02 #6 done52 line(s)

    shot 6·sharpness 2473.6

    1. AlEngineer1.00
    2. PARAMETER GOLF· THE METHOD0.99
    3. World's Fair0.97
    4. Aiden, an agent that publishes its own work0.99
    5. WHAT IT READS0.97
    6. PRESENTED.BY1.00
    7. 0.94
    8. from the literature1.00
    9. Public papers1.00
    10. AIDEN· THE PRIVATE LOOP0.96
    11. Ideation1.00
    12. Microsoft1.00
    13. Forms new experiment ideas1.00
    14. 80.99
    15. others' attempts1.00
    16. Repository PRs1.00
    17. WHAT IT PUBLISHES1.00
    18. A0.92
    19. Local experimentation1.00
    20. Trains & evaluates candidates locally1.00
    21. ?0.61
    22. Leaderboard1.00
    23. others cite, fork & build on1.00
    24. 0.98
    25. its own past runs1.00
    26. Experiment logs0.95
    27. Quality gate1.00
    28. 0.68
    29. reviewed &1.00
    30. reproduced1.00
    31. Follows all the rules0.98
    32. passes gate1.00
    33. 0.52
    34. opens a public PR1.00
    35. Publish1.00
    36. d'sFair1.00
    37. Upel0.72
    38. G0.74
    39. Reflection1.00
    40. rewrites prompts & tools1.00
    41. The gain is real, not noise0.99
    42. AlEngin0.99
    43. → Reads in0.91
    44. → Internal step0.96
    45. Publishes1.00
    46. -→ Self-improvement0.97
    47. crosof1.00
    48. World'1.00
    49. igineer0.98
    50. d's0.91
    51. sny1.00
    52. Engineering the future of Al0.99
  • 2:24 #7 skipped

    shot 7·duplicate of #4

  • 3:02 #8 done24 line(s)

    shot 8·sharpness 1980.2

    1. AlEngineer0.99
    2. PARAMETER GOLF·THE RESULTS1.00
    3. World's Fair0.97
    4. It set the most records on the board0.99
    5. dexhunter (Aiden)0.97
    6. PRESENTED BY0.99
    7. aquariouseworkman1.00
    8. 31.00
    9. Microsoft1.00
    10. abaybektursun1.00
    11. aryanbhosale1.00
    12. 21.00
    13. + 7 more0.99
    14. 2-11.00
    15. sFair1.00
    16. Open1.00
    17. 7 official leaderboard records, merged PRs that set a new best. More than twice the top human's 3.0.99
    18. AlEnginee1.00
    19. rosoft1.00
    20. World's0.95
    21. 'sFi0.78
    22. neer1.00
    23. snyl1.00
    24. Engineering the future of Al0.99
  • 3:08 #9 done33 line(s)

    shot 9·sharpness 2590.3

    1. AlEngineer0.99
    2. PARAMETER GOLF· THE RESULTS0.99
    3. World'sFair1.00
    4. And other people built on its work1.00
    5. CCCCCCCCCCCO1.00
    6. Baseline0.99
    7. #50 MatthewL0.81
    8. PRESENTED BY1.00
    9. Microsoft1.00
    10. Its records kept getting forked, cited, and1.00
    11. #493 abaybektursu0.98
    12. built on more than anyone's.1.00
    13. #10601.00
    14. KevinClark0.99
    15. a1334 aryanbhosale0.99
    16. #14131.00
    17. #1529 msisovic0.87
    18. #15140.93
    19. was 7.0.99
    20. PR-citation h-index of 10, the highest of any contributor. The next0.99
    21. air1.00
    22. OpenAl0.96
    23. #17691.00
    24. nprime060.93
    25. #1855 codemath30000.99
    26. #2135 codemath30000.99
    27. soft1.00
    28. orld's Fa0.92
    29. AlEngineer0.98
    30. Each row is a leaderboard record; lines trace which0.99
    31. built on which. The highlighted rows are Aiden's.1.00
    32. :air0.96
    33. Engineering the future of Al1.00
  • 3:35 #10 skipped

    shot 10·duplicate of #9

  • 4:16 #11 done16 line(s)

    shot 11·sharpness 1789.2

    1. AlEngineer0.99
    2. PARAMETER GOLF· THE RESULTS0.98
    3. World'sFair1.00
    4. Of course an Al has high throughput.1.00
    5. PRESENTED BY0.99
    6. 1,3001.00
    7. Microsoft1.00
    8. experiments over 22 days, on a single node.1.00
    9. Fair1.00
    10. Open/0.98
    11. AlEngineer0.99
    12. osoft1.00
    13. World'sI0.92
    14. ;Fai0.86
    15. yh0.97
    16. Engineering the future of Al1.00
  • 4:37 #12 done24 line(s)

    shot 12·sharpness 2126.4

    1. AlEngineer0.99
    2. PARAMETER GOLF· THE RESULTS0.99
    3. World's Fair0.99
    4. But it ran better than the field, not just more1.00
    5. Compute used vs. records set0.99
    6. PR acceptance ratehigher is better0.99
    7. PRESENTED BY0.99
    8. 4%1.00
    9. Aiden1.00
    10. 28%1.00
    11. Microsoft1.00
    12. → 15% of the records (7 of 47)0.97
    13. Field average1.00
    14. 4.7%1.00
    15. ~6× the field's acceptance rate.1.00
    16. ;Fair0.95
    17. Open/0.95
    18. AlEngineer0.98
    19. 'osoft0.95
    20. World'sI0.95
    21. Fi0.84
    22. ber0.91
    23. snyh1.00
    24. Engineering the future of Al1.00
  • 5:13 #13 skipped

    shot 13·duplicate of #12

  • 5:36 #14 done19 line(s)

    shot 14·sharpness 2405.0

    1. AlEngineer0.98
    2. Did the Al beat the humans?0.99
    3. World'sFair1.00
    4. Where the ideas came from1.00
    5. research papers1.00
    6. PRESENTEDBY1.00
    7. other participants1.00
    8. Microsoft1.00
    9. other communities, like nanoGPT0.99
    10. the file-size constraint, its most original1.00
    11. ideas1.00
    12. air1.00
    13. Google Deepl0.95
    14. Almost every idea came from people. Aiden's most original came from the constraint.0.99
    15. AlEngineer0.99
    16. to1.00
    17. World'sFail0.98
    18. ayPa!0.93
    19. Engineering the future of Al0.99
  • 6:12 #15 skipped

    shot 15·duplicate of #14

  • 6:29 #16 done18 line(s)

    shot 16·sharpness 2302.6

    1. AlEngineer0.98
    2. Did the Al beat the humans?1.00
    3. World'sFair1.00
    4. Where the ideas came from0.99
    5. research papers1.00
    6. other participants1.00
    7. other communities, like nanoGPT1.00
    8. the file-size constraint, its most original1.00
    9. ideas1.00
    10. ;Fair0.91
    11. Google Dee0.99
    12. Almost every idea came from people. Aiden's most original came from the constraint.0.99
    13. AlEngineer0.99
    14. rcto0.94
    15. World'sF1.00
    16. ;Fai0.87
    17. yPc0.83
    18. Engineering the future of Al0.99
  • 7:08 #17 done33 line(s)

    shot 17·sharpness 1720.6

    1. AlEngineer0.98
    2. World's Fair0.97
    3. 3 · Merging two ideas0.98
    4. The community's CaseOps tokenizer is the real improvement — a valid leaderboard record.1.00
    5. 1.07191.00
    6. 1.07181.00
    7. 1.0721.00
    8. (ot s ter)0.64
    9. 1.0701.00
    10. #1729·community0.99
    11. Valin BBPpB0.74
    12. 1.068-0.93
    13. 1.06781.00
    14. 1.06801.00
    15. over the 16 MB cap - invalid0.98
    16. 1.0661.00
    17. 1.06551.00
    18. air0.99
    19. Gooale Deepl0.95
    20. 1.064-0.95
    21. #16261.00
    22. + Gated attention1.00
    23. + Quant gate1.00
    24. + CaseOps0.99
    25. base1.00
    26. (Qwen paper)1.00
    27. (Aiden)1.00
    28. (community)1.00
    29. Engineer1.00
    30. to1.00
    31. Id'sFaii0.88
    32. air1.00
    33. Engineering the future of Al0.98
  • 7:20 #18 skipped

    shot 18·duplicate of #17

  • 7:51 #19 done19 line(s)

    shot 19·sharpness 2429.6

    1. AlEngineer0.97
    2. What Aiden is good0.99
    3. World'sFair1.00
    4. at1.00
    5. 1 · Finding and implementing ideas0.96
    6. It brought a recent paper's idea into the competition, and mined good ingredients out of a noisy community.0.99
    7. 2· The logically-obvious moves0.98
    8. Add parameters, break the file-size limit, so reach for quantization.0.99
    9. 3 · Efficient combination search0.98
    10. Quickly finding the right combinations across a huge search space.1.00
    11. d'sFair0.96
    12. Google D0.94
    13. AlEngir0.96
    14. ducto1.00
    15. World'0.92
    16. ngineer1.00
    17. Payl1.00
    18. d'sF0.99
    19. Engineering the future of Al1.00
  • 8:13 #20 skipped

    shot 20·duplicate of #19

  • 8:53 #21 done14 line(s)

    shot 21·sharpness 2434.0

    1. AlEngineer0.96
    2. World'sFair1.00
    3. None of these sound very sexy. But good0.99
    4. execution here is what's pivotal.0.99
    5. What we're really looking at is a large group of humans and one Al system.0.99
    6. Does that mean a single human engineer's marginal contribution gets smaller?1.00
    7. air1.00
    8. Google Deepl0.96
    9. AlEngineer0.98
    10. :to0.84
    11. Norld'sFai0.98
    12. air1.00
    13. ayPa!0.93
    14. Engineering the future of Al0.99
  • 9:31 #22 done14 line(s)

    shot 22·sharpness 2514.4

    1. AlEngineer0.97
    2. World'sFair1.00
    3. None of these sound very sexy. But good1.00
    4. PRESENTED BY0.99
    5. execution here is what's pivotal.0.99
    6. Microsoft0.95
    7. What we're really looking at is a large group of humans and one Al system.0.99
    8. Does that mean a single human engineer's marginal contribution gets smaller?1.00
    9. air1.00
    10. GoogleDeepl1.00
    11. to1.00
    12. 'sFai0.97
    13. air1.00
    14. Engineering the future of Al0.99
  • 9:43 #23 skipped

    shot 23·duplicate of #3

Transcript

149 cues· 1,763 words· 10,103 chars

  1. 0:12 This April, OpenAI ran a hiring challenge, a competition called Parameter Golf.
  2. 0:19 The top contributor was one candidate that they couldn't hire.
  3. 0:25 It wasn't a person.
  4. 0:26 It's an agent we built called Aiden.
  5. 0:31 In Parameter Golf, the goal is to train the best language model you can under size and the computation constraints.
  6. 0:41 About 1,000 machine learning engineers, researchers participate.
  7. 0:47 They filed 2,000 submissions.
  8. 0:51 Only 47 passed OpenAI's review and made into the leaderboard.
  9. 0:57 Seven of those are actually aidants.
  10. 1:01 More than twice what any human contributed.
  11. 1:07 You've seen a lot of auto research today.
  12. 1:09 Agents are here climbing benchmarks.
  13. 1:12 Those are really impressive results.
  14. 1:15 The question I want to ask is a bit different here.
  15. 1:18 Can the auto research agent produce work that a human community actually recognize?
  16. 1:26 Beyond a good score, agent is optimizing for something that other engineers can merge, fork, and build on.
  17. 1:36 So instead of having an agent just here climbing locally, we build one that publishes its own work, and that's Aden.
  18. 1:47 Quick context on us.
  19. 1:49 WeCo is an auto research company that founded about two and a half years ago.
  20. 1:54 I'm co-founder and the CEO, Zheng Yao.
  21. 1:57 Got my PhD at UCL on reinforcement learning.
  22. 2:01 About two years ago, we viewed AID, the top auto research agent independently evaluated by OpenAI in their MLE bench paper.
  23. 2:14 even though back then there's no such name called auto research.
  24. 2:18 People call it machine learning engineering agent.
  25. 2:22 ADAN is the next step
  26. 2:25 and an experimental prototype.
  27. 2:29 It's a multi-agent, self-improving system that can read public information, like research papers and other PRs, run its own experiments, and submit a PR once the findings pass a quality gate.
  28. 2:46 We send Aiden to parameter golf competition.
  29. 2:50 And it ran for about 22 days.
  30. 2:53 By the end, Aiden has set seven leaderboard records.
  31. 2:57 Each one is a new best for the competition, sampled by OpenAI.
  32. 3:02 And the best human only made three.
  33. 3:07 Passing the host review is one signal for the quality.
  34. 3:11 A second, maybe more important one,
  35. 3:15 is whether other participants would build on your work.
  36. 3:20 And it turns out Aidan's work had the highest impact within the whole community.
  37. 3:27 Here we are using a inference measure that is used widely in academia.
  38. 3:34 It's called the H-index.
  39. 3:36 Roughly, if you have X papers get cited X times, then your H-index is X.
  40. 3:44 Computed over PRs, Aiden was 10, and the next human was seven.
  41. 3:50 The whole community was building on an AI system's work, including many of other leaderboard entries.
  42. 4:01 To break it down a little bit, why can an autonomous AI system be so powerful?
  43. 4:08 One obvious reason is that it's an AI.
  44. 4:12 It can run tirelessly.
  45. 4:14 Over 22 days, it ran about 1,300 experiments on a single H100 node.
  46. 4:25 But the throughput isn't the whole picture.
  47. 4:27 A well-tuned AI system can also keep its output quality high.
  48. 4:34 On the compute side, it uses at most 4% of competition's total compute.
  49. 4:44 And it made about 15% of the records.
  50. 4:49 Also, 28% of its submissions made the leaderboard, roughly six times higher hit rate than the community average.

Chapters

  1. 0:00 Introduction to Parameter Golf and the Aiden agent
  2. 1:06 Defining the challenge: Auto-research vs. human community
  3. 1:47 About Weco AI and the development of Aiden
  4. 3:07 Evaluating Aiden's impact and H-index in the community
  5. 4:01 Why autonomous AI is powerful: Throughput and efficiency
  6. 5:21 Human-AI collaboration: How ideas move the frontier
  7. 6:32 Case study: Combining research, architecture, and tokenization
  8. 7:41 Summary of auto-research strengths: Execution and search
  9. 9:06 The role of human design in competition
  10. 10:04 The Andrej Karpathy metaphor: Gradient descent and coding
  11. 11:19 Auto-research as training a model: Evals and abstractions
  12. 13:36 Case study: Improving data pipelines via strict API abstractions
  13. 14:38 Conclusion: The new craft of the AI engineer

Open at this second