Videos iCj_ATyThvc
How Autoresearch is changing ML research — Zhengyao Jiang, Weco
Scene timeline
37 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 149
- whisperx 149
- chunks
- 29
- from 149 cues
- keyframes
- 23
- kept of 37 captured
- frames with text
- 23
- 497 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 4.6 MB
- word timings on 149 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 02:44 | 1m 12s |
stt |
done | — | 2026-08-11 02:45 | 15s |
chunk |
done | — | 2026-08-11 02:46 | 0s |
text_embed |
done | — | 2026-08-11 02:46 | 0s |
keyframe |
done | — | 2026-08-11 02:46 | 1m 39s |
ocr |
done | — | 2026-08-11 02:47 | 11s |
frame_embed |
done | — | 2026-08-11 02:48 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.93
- OpenAI0.92
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.90
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- World'sFair1.00
- This year, OpenAl ran a hiring challenge, a competition called1.00
- Parameter Golf.1.00
- PRESENTED BY0.97
- Microsoft1.00
- The top contributor was the one candidate they couldn't hire.1.00
- It wasn't a person. It's an agent we built, Aiden.0.98
- d'sFair1.00
- Opē0.87
- AlEngir0.95
- icrosoft1.00
- World'0.94
- d'sF0.93
- sny0.99
- Engineering the future of Al0.98
- inhtr0.89
-
- AlEngineer0.97
- OPENAI· MODEL CRAFT CHALLENGE0.97
- World'sFair1.00
- Parameter Golf0.99
- Code golf for neural networks. Train the best language model you can, then shrink it.0.99
- PRESENTED BY1.00
- Microsoft1.00
- 16 MB0.99
- 10 min1.00
- BPB↓0.98
- Total artifact1.00
- Training budget1.00
- Lower is better1.00
- Model weights + training code, combined.1.00
- Wall-clock cap on 8×H100 GPUs.1.00
- Bits-per-byte on the FineWeb val set1.00
- (tokenizer-agnostic).1.00
- Fair1.00
- OpenA0.97
- BY THE NUMBERS0.99
- THE FIELD,1.00
- 1,0161.00
- 2,0481.00
- 471.00
- $249,5501.00
- participants1.00
- PRs filed0.96
- merged records · 31 contributors0.97
- RunPod credits bumed0.99
- AlEngineer1.00
- soft0.99
- lorld's Fa0.95
- Mar 18–Apr 30, 20260.98
- Fair0.97
- yk0.84
- Engineering the future of Al0.99
- tuiin0.51
-
- AlEngineer0.95
- World'sFair1.00
- You've seen a lot of autoresearch today, agents hill-climbing benchmarks.0.99
- PRESENTED BY0.96
- Can an autoresearch agent produce work0.99
- Microsoft0.96
- a community recognizes?1.00
- Not a good score on your own machine. Something other engineers will merge, fork,1.00
- and build on.0.97
- Fair1.00
- OpenA0.94
- AlEngineer0.99
- osoft0.98
- World'sF0.99
- Fair0.99
- Engineering the future of Al0.99
-
- AlEngineer1.00
- PARAMETER GOLF· THE METHOD0.99
- World's Fair0.97
- Aiden, an agent that publishes its own work0.99
- WHAT IT READS0.97
- PRESENTED.BY1.00
- 目0.94
- from the literature1.00
- Public papers1.00
- AIDEN· THE PRIVATE LOOP0.96
- Ideation1.00
- Microsoft1.00
- Forms new experiment ideas1.00
- 80.99
- others' attempts1.00
- Repository PRs1.00
- WHAT IT PUBLISHES1.00
- A0.92
- Local experimentation1.00
- Trains & evaluates candidates locally1.00
- ?0.61
- Leaderboard1.00
- others cite, fork & build on1.00
- 目0.98
- its own past runs1.00
- Experiment logs0.95
- Quality gate1.00
- →0.68
- reviewed &1.00
- reproduced1.00
- Follows all the rules0.98
- passes gate1.00
- ①0.52
- opens a public PR1.00
- Publish1.00
- d'sFair1.00
- Upel0.72
- G0.74
- Reflection1.00
- rewrites prompts & tools1.00
- The gain is real, not noise0.99
- AlEngin0.99
- → Reads in0.91
- → Internal step0.96
- Publishes1.00
- -→ Self-improvement0.97
- crosof1.00
- World'1.00
- igineer0.98
- d's0.91
- sny1.00
- Engineering the future of Al0.99
-
- AlEngineer0.99
- PARAMETER GOLF·THE RESULTS1.00
- World's Fair0.97
- It set the most records on the board0.99
- dexhunter (Aiden)0.97
- PRESENTED BY0.99
- aquariouseworkman1.00
- 31.00
- Microsoft1.00
- abaybektursun1.00
- aryanbhosale1.00
- 21.00
- + 7 more0.99
- 2-11.00
- sFair1.00
- Open1.00
- 7 official leaderboard records, merged PRs that set a new best. More than twice the top human's 3.0.99
- AlEnginee1.00
- rosoft1.00
- World's0.95
- 'sFi0.78
- neer1.00
- snyl1.00
- Engineering the future of Al0.99
-
- AlEngineer0.99
- PARAMETER GOLF· THE RESULTS0.99
- World'sFair1.00
- And other people built on its work1.00
- CCCCCCCCCCCO1.00
- Baseline0.99
- #50 MatthewL0.81
- PRESENTED BY1.00
- Microsoft1.00
- Its records kept getting forked, cited, and1.00
- #493 abaybektursu0.98
- built on more than anyone's.1.00
- #10601.00
- KevinClark0.99
- a1334 aryanbhosale0.99
- #14131.00
- #1529 msisovic0.87
- #15140.93
- was 7.0.99
- PR-citation h-index of 10, the highest of any contributor. The next0.99
- air1.00
- OpenAl0.96
- #17691.00
- nprime060.93
- #1855 codemath30000.99
- #2135 codemath30000.99
- soft1.00
- orld's Fa0.92
- AlEngineer0.98
- Each row is a leaderboard record; lines trace which0.99
- built on which. The highlighted rows are Aiden's.1.00
- :air0.96
- Engineering the future of Al1.00
-
- AlEngineer0.99
- PARAMETER GOLF· THE RESULTS0.98
- World'sFair1.00
- Of course an Al has high throughput.1.00
- PRESENTED BY0.99
- 1,3001.00
- Microsoft1.00
- experiments over 22 days, on a single node.1.00
- Fair1.00
- Open/0.98
- AlEngineer0.99
- osoft1.00
- World'sI0.92
- ;Fai0.86
- yh0.97
- Engineering the future of Al1.00
-
- AlEngineer0.99
- PARAMETER GOLF· THE RESULTS0.99
- World's Fair0.99
- But it ran better than the field, not just more1.00
- Compute used vs. records set0.99
- PR acceptance ratehigher is better0.99
- PRESENTED BY0.99
- 4%1.00
- Aiden1.00
- 28%1.00
- Microsoft1.00
- → 15% of the records (7 of 47)0.97
- Field average1.00
- 4.7%1.00
- ~6× the field's acceptance rate.1.00
- ;Fair0.95
- Open/0.95
- AlEngineer0.98
- 'osoft0.95
- World'sI0.95
- Fi0.84
- ber0.91
- snyh1.00
- Engineering the future of Al1.00
-
- AlEngineer0.98
- Did the Al beat the humans?0.99
- World'sFair1.00
- Where the ideas came from1.00
- research papers1.00
- PRESENTEDBY1.00
- other participants1.00
- Microsoft1.00
- other communities, like nanoGPT0.99
- the file-size constraint, its most original1.00
- ideas1.00
- air1.00
- Google Deepl0.95
- Almost every idea came from people. Aiden's most original came from the constraint.0.99
- AlEngineer0.99
- to1.00
- World'sFail0.98
- ayPa!0.93
- Engineering the future of Al0.99
-
- AlEngineer0.98
- Did the Al beat the humans?1.00
- World'sFair1.00
- Where the ideas came from0.99
- research papers1.00
- other participants1.00
- other communities, like nanoGPT1.00
- the file-size constraint, its most original1.00
- ideas1.00
- ;Fair0.91
- Google Dee0.99
- Almost every idea came from people. Aiden's most original came from the constraint.0.99
- AlEngineer0.99
- rcto0.94
- World'sF1.00
- ;Fai0.87
- yPc0.83
- Engineering the future of Al0.99
-
- AlEngineer0.98
- World's Fair0.97
- 3 · Merging two ideas0.98
- The community's CaseOps tokenizer is the real improvement — a valid leaderboard record.1.00
- 1.07191.00
- 1.07181.00
- 1.0721.00
- (ot s ter)0.64
- 1.0701.00
- #1729·community0.99
- Valin BBPpB0.74
- 1.068-0.93
- 1.06781.00
- 1.06801.00
- over the 16 MB cap - invalid0.98
- 1.0661.00
- 1.06551.00
- air0.99
- Gooale Deepl0.95
- 1.064-0.95
- #16261.00
- + Gated attention1.00
- + Quant gate1.00
- + CaseOps0.99
- base1.00
- (Qwen paper)1.00
- (Aiden)1.00
- (community)1.00
- Engineer1.00
- to1.00
- Id'sFaii0.88
- air1.00
- Engineering the future of Al0.98
-
- AlEngineer0.97
- What Aiden is good0.99
- World'sFair1.00
- at1.00
- 1 · Finding and implementing ideas0.96
- It brought a recent paper's idea into the competition, and mined good ingredients out of a noisy community.0.99
- 2· The logically-obvious moves0.98
- Add parameters, break the file-size limit, so reach for quantization.0.99
- 3 · Efficient combination search0.98
- Quickly finding the right combinations across a huge search space.1.00
- d'sFair0.96
- Google D0.94
- AlEngir0.96
- ducto1.00
- World'0.92
- ngineer1.00
- Payl1.00
- d'sF0.99
- Engineering the future of Al1.00
-
- AlEngineer0.96
- World'sFair1.00
- None of these sound very sexy. But good0.99
- execution here is what's pivotal.0.99
- What we're really looking at is a large group of humans and one Al system.0.99
- Does that mean a single human engineer's marginal contribution gets smaller?1.00
- air1.00
- Google Deepl0.96
- AlEngineer0.98
- :to0.84
- Norld'sFai0.98
- air1.00
- ayPa!0.93
- Engineering the future of Al0.99
-
- AlEngineer0.97
- World'sFair1.00
- None of these sound very sexy. But good1.00
- PRESENTED BY0.99
- execution here is what's pivotal.0.99
- Microsoft0.95
- What we're really looking at is a large group of humans and one Al system.0.99
- Does that mean a single human engineer's marginal contribution gets smaller?1.00
- air1.00
- GoogleDeepl1.00
- to1.00
- 'sFai0.97
- air1.00
- Engineering the future of Al0.99
Transcript
149 cues· 1,763 words· 10,103 chars
- 0:12 This April, OpenAI ran a hiring challenge, a competition called Parameter Golf.
- 0:19 The top contributor was one candidate that they couldn't hire.
- 0:25 It wasn't a person.
- 0:26 It's an agent we built called Aiden.
- 0:31 In Parameter Golf, the goal is to train the best language model you can under size and the computation constraints.
- 0:41 About 1,000 machine learning engineers, researchers participate.
- 0:47 They filed 2,000 submissions.
- 0:51 Only 47 passed OpenAI's review and made into the leaderboard.
- 0:57 Seven of those are actually aidants.
- 1:01 More than twice what any human contributed.
- 1:07 You've seen a lot of auto research today.
- 1:09 Agents are here climbing benchmarks.
- 1:12 Those are really impressive results.
- 1:15 The question I want to ask is a bit different here.
- 1:18 Can the auto research agent produce work that a human community actually recognize?
- 1:26 Beyond a good score, agent is optimizing for something that other engineers can merge, fork, and build on.
- 1:36 So instead of having an agent just here climbing locally, we build one that publishes its own work, and that's Aden.
- 1:47 Quick context on us.
- 1:49 WeCo is an auto research company that founded about two and a half years ago.
- 1:54 I'm co-founder and the CEO, Zheng Yao.
- 1:57 Got my PhD at UCL on reinforcement learning.
- 2:01 About two years ago, we viewed AID, the top auto research agent independently evaluated by OpenAI in their MLE bench paper.
- 2:14 even though back then there's no such name called auto research.
- 2:18 People call it machine learning engineering agent.
- 2:22 ADAN is the next step
- 2:25 and an experimental prototype.
- 2:29 It's a multi-agent, self-improving system that can read public information, like research papers and other PRs, run its own experiments, and submit a PR once the findings pass a quality gate.
- 2:46 We send Aiden to parameter golf competition.
- 2:50 And it ran for about 22 days.
- 2:53 By the end, Aiden has set seven leaderboard records.
- 2:57 Each one is a new best for the competition, sampled by OpenAI.
- 3:02 And the best human only made three.
- 3:07 Passing the host review is one signal for the quality.
- 3:11 A second, maybe more important one,
- 3:15 is whether other participants would build on your work.
- 3:20 And it turns out Aidan's work had the highest impact within the whole community.
- 3:27 Here we are using a inference measure that is used widely in academia.
- 3:34 It's called the H-index.
- 3:36 Roughly, if you have X papers get cited X times, then your H-index is X.
- 3:44 Computed over PRs, Aiden was 10, and the next human was seven.
- 3:50 The whole community was building on an AI system's work, including many of other leaderboard entries.
- 4:01 To break it down a little bit, why can an autonomous AI system be so powerful?
- 4:08 One obvious reason is that it's an AI.
- 4:12 It can run tirelessly.
- 4:14 Over 22 days, it ran about 1,300 experiments on a single H100 node.
- 4:25 But the throughput isn't the whole picture.
- 4:27 A well-tuned AI system can also keep its output quality high.
- 4:34 On the compute side, it uses at most 4% of competition's total compute.
- 4:44 And it made about 15% of the records.
- 4:49 Also, 28% of its submissions made the leaderboard, roughly six times higher hit rate than the community average.
loading
Chapters
- 0:00 Introduction to Parameter Golf and the Aiden agent
- 1:06 Defining the challenge: Auto-research vs. human community
- 1:47 About Weco AI and the development of Aiden
- 3:07 Evaluating Aiden's impact and H-index in the community
- 4:01 Why autonomous AI is powerful: Throughput and efficiency
- 5:21 Human-AI collaboration: How ideas move the frontier
- 6:32 Case study: Combining research, architecture, and tokenization
- 7:41 Summary of auto-research strengths: Execution and search
- 9:06 The role of human design in competition
- 10:04 The Andrej Karpathy metaphor: Gradient descent and coding
- 11:19 Auto-research as training a model: Evals and abstractions
- 13:36 Case study: Improving data pipelines via strict API abstractions
- 14:38 Conclusion: The new craft of the AI engineer