Videos kiqubc5b5Yo
Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo
Scene timeline
46 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 169
- whisperx 169
- chunks
- 35
- from 169 cues
- keyframes
- 27
- kept of 46 captured
- frames with text
- 27
- 647 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 4.9 MB
- word timings on 169 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:00 | 0s |
stt |
done | — | 2026-08-09 05:06 | 27s |
chunk |
done | — | 2026-08-09 05:07 | 0s |
text_embed |
done | — | 2026-08-10 19:38 | 1s |
keyframe |
done | — | 2026-08-09 05:07 | 2m 28s |
ocr |
done | — | 2026-08-09 05:09 | 13s |
frame_embed |
done | — | 2026-08-10 19:38 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- World'sFair1.00
- MORGAN STANLEY0.99
- ALPHALAB1.00
- AlphaLab1.00
- Building an auto-research agent — and what it0.99
- taught us about evals & environments.1.00
- Brendan Rappazzo Hogan0.98
- AI Engineer World's Fair1.00
- Engineering the future of Al0.98
- World'sFair1.00
-
- AlEngineer0.97
- 01· THE SETUP0.99
- World'sFair1.00
- The team & the opportunity1.00
- PRESENTED BY1.00
- 01 ~30 PhD Al/ML researchers — half academic, half0.98
- applied work.1.00
- Microsoft1.00
- THE CATCH1.00
- 02 Our problems are often well-posed — e.g. given0.99
- An auto-research system felt out of0.99
- some time-series data, predict future values,0.99
- often under constraints like staying well-1.00
- reach -0.91
- calibrated.1.00
- until December 2025.1.00
- 03 And highly homogeneous across desks — the0.99
- same shapes recur. Lots of low-hanging fruit.1.00
- AUTO-RESEARCH MORGAN STANLEY0.99
- Engineering the future of Al1.00
- World'sFair1.00
-
- AlEngineer0.98
- 01· THE SETUP0.99
- World'sFair1.00
- The team & the opportunity1.00
- PRESENTED BY1.00
- 01 ~30 PhD Al/ML researchers — half academic, half0.98
- applied work.1.00
- Microsoft1.00
- THE CATCH1.00
- 02 Our problems are often well-posed — e.g. given0.99
- An auto-research system felt out of1.00
- some time-series data, predict future values,0.99
- often under constraints like staying well-0.99
- reach -0.93
- calibrated.1.00
- until December 2025.1.00
- 03 And highly homogeneous across desks — the0.99
- same shapes recur. Lots of low-hanging fruit.1.00
- ALPHALAB1.00
- AUTO-RESEARCH @ MORGAN STANLEY0.99
- TRACK 3• JULY 2, 20260.96
- Al in Finance0.98
- World'sFair1.00
-
- AlEngineer0.97
- 01· THE SETUP0.99
- World'sFair1.00
- The team & the opportunity1.00
- 01 ~30 PhD Al/ML researchers — half academic, half0.98
- applied work.1.00
- THE CATCH1.00
- 02 Our problems are often well-posed — e.g. given0.98
- An auto-research system felt out of1.00
- some time-series data, predict future values,1.00
- often under constraints like staying well-0.99
- reach -0.93
- calibrated.1.00
- until December 2025.1.00
- 03 And highly homogeneous across desks — the0.98
- same shapes recur. Lots of low-hanging fruit.1.00
- ALPHALAB1.00
- AUTO-RESEARCH @ MORGAN STANLEY0.98
- 02 / 180.90
- TRACK 3· JULY 2, 20260.95
- Al in Finance0.97
- World'sFair1.00
-
- AlEngineer0.99
- 02 · GOALS0.93
- World's Fair0.99
- What we wanted from auto-research1.00
- ABOVE ALL ELSE0.98
- Maximize new algorithms in production.1.00
- data1.00
- spectrum0.97
- agnostic1.00
- Plug into all our data & prebuilt1.00
- Span: improve existing → transfer →0.98
- Model-agnostic — and improve as0.97
- scaffolding.1.00
- run from scratch.1.00
- the models improve.1.00
- expertise1.00
- compute1.00
- Encode our human & enterprise0.99
- Keep our GPUs churning.1.00
- knowledge.1.00
- ALPHALAB· AUTO-RESEARCH @ MORGAN STANLEY0.97
- 03 / 180.97
- TRACK 3· JULY 2, 20260.96
- Al in Finance1.00
- World's Fair0.99
-
- AlEngineer0.98
- 03· ROADMAP0.98
- World'sFair1.00
- Roadmap1.00
- 011.00
- The team & the idea1.00
- 021.00
- AlphaLab 1.0 — how we built it0.99
- 031.00
- Did it work? — and what broke0.99
- 041.00
- Evals & environments1.00
- 051.00
- Climbing to recursive self-improvement1.00
- 061.00
- AlphaLab 2.01.00
- AUTO-RESEARCH @ MORGAN STANLEY0.98
- TRACK 3• JULY 2, 20260.96
- Al in Finance0.99
- World's Fair0.98
-
- AlEngineer0.99
- 03 · ROADMAP0.91
- World's Fair0.98
- Roadmap1.00
- 011.00
- The team & the idea0.98
- PRESENTED BY1.00
- 020.99
- AlphaLab 1.0 — how we built it0.98
- Microsoft1.00
- 030.99
- Did it work? — and what broke0.98
- 041.00
- Evals & environments1.00
- 051.00
- Climbing to recursive self-improvement1.00
- 061.00
- AlphaLab 2.01.00
- AUTO-RESEARCH @ MORGAN STANLEY0.99
- TRACK 3· JULY 2, 20260.95
- Al in Finance1.00
- World's Fair1.00
-
- AlEngineer0.99
- ALPHALAB 1.01.00
- World'sFair1.00
- The agentic harness1.00
- One harness. Input: where the data is & what to predict. Output: a suite of trained models.0.99
- PRESENTED BY1.00
- Microsoft1.00
- INPUT1.00
- PHASE 11.00
- PHASE 21.00
- PHASE 30.99
- OUTPUT1.00
- data_path,1.00
- Research1.00
- Eval building1.00
- Experiment1.00
- A suite of1.00
- trained ML1.00
- prediction_task1.00
- explore the data1.00
- build the scoreboard1.00
- run at scale1.00
- models.1.00
- ALPHALAB· AUTO-RESEARCH MORGAN STANLEY0.97
- 05 / 180.98
- TRACK 3·JULY 2,20260.96
- Alin Finance0.96
- World'sFair1.00
-
- AlEngineer0.98
- ALPHALAB 1.0· UNDER THE HOOD0.98
- World's Fair0.98
- How the harness works0.97
- It's all homegrown — the harness, the scaffolding, every tool, and the orchestration across many0.99
- models. No off-the-shelf agent framework.0.98
- CONTEXT - notes, files, to-do0.98
- tool_call1.00
- The Model1.00
- "tool": "slurm.submit",1.00
- orchestrated - many models0.96
- "args": { "script": "train.py", "gpus": 4 }}0.97
- shell1.00
- web1.00
- slurm1.00
- emits a tool call1.00
- full access1.00
- research1.00
- GPU cluster1.00
- The model only ever reads and writes text. Everything around it — the loop,0.98
- harness executes it → result appended0.98
- the tools, the multi-model orchestration — is code we wrote.0.98
- ALPHALAB· AUTO-RESEARCH MORGAN STANLEY0.98
- 06 / 180.98
- TRACK 3· JULY 2,20260.94
- Al in Finance0.98
- World's Fair0.98
-
- AlEngineer0.99
- PHASE 1 OF 31.00
- World's Fair1.00
- Research1.00
- AlphaLab0.97
- Phase 1·plan.md0.98
- Alpha Lab PHASE10.93
- Thinking.0.95
- Build up the context & boilerplate that make the1.00
- FILES1.00
- data_report1.00
- plan.md1.00
- Files0.99
- Board1.00
- next phases possible.1.00
- Scaffolding is built around a single to-do list item.0.99
- logs0.91
- nd schema.md0.96
- nd statistics.md0.96
- Forecast exchange_rate for each country at multiple horizons (sh0.98
- test whether exogenous features (ex_1...ex_4) add value.0.99
- To-Do (will be checked off as completed)1.00
- UUaI0.54
- Every stage writes detailed markdown notes.0.99
- notes1.00
- nd 00_data_overview.md0.98
- conversation.jsonl1.00
- Data Loading & Schema1.00
- Workspace & Reproducibility0.99
- [] Verify environment/packages; ensure analysis is fully reproducible0.98
- [] Create a data loading utility and consistent plotting style.0.99
- Web search is most important at this stage.1.00
- nd 01_statistical_profiling.md0.99
- [] Load CSV; parse dates; inspect dtypes, missingness, duplicates.0.98
- ~3-4 hrs1.00
- plots1.00
- img00_country_counts.png0.97
- img01_ex_features_hist.png1.00
- Statistical Profiling (All Columns)0.98
- [] Distribution plots for exchange_rate and ex_1...ex_4 (overall &0.99
- [] Per-column summary stats (overall + by country): mean, std, quar0.99
- Validate date range and expected countries (should be 8).0.99
- Confirm entity keys: (date, country) uniqueness.0.99
- img01_target_box_by_country...0.97
- Temporal Structure0.97
- scripts0.95
- img01_target_hist_by_country...0.86
- [] Seasonality diagnostics (weekly/monthly/annual).0.99
- Determine frequency and gaps per country; identify non-trading0.99
- Detect structural breaks / regime changes (1990-2010 includes c0.97
- py 00_setup_and_load.py0.99
- Target Variable Deep Dive1.00
- ALPHALAB· AUTO-RESEARCH @ MORGAN STANLEY0.98
- 07 / 180.92
- TRACK 3· JULY 2, 20260.92
- Al in Finance1.00
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer0.99
- PHASE 2 OF 31.00
- World'sFair1.00
- Eval building1.00
- Builder0.94
- Build an evaluation / backtesting framework — the0.97
- bullds eval framework1.00
- scoreboard everything else plays against.0.99
- LLMs aren't malicious, but they're prone to mistakes. So the build1.00
- issues found1.00
- Critic (fresh agent)0.99
- checks leakage & bias0.95
- itself is adversarial:0.98
- Critical issues?0.99
- tests fall0.98
- Builder1.00
- writes the eval framework1.00
- no issues0.99
- Tester1.00
- runs automated tests0.99
- Critic1.00
- fresh agent — checks leakage & bias0.99
- Tests pass?0.99
- Tester1.00
- runs automated tests1.00
- or max iterations1.00
- all pass0.95
- Exits once there are no critical issues left.1.00
- Valldated harness0.96
- passed to Phase 30.99
- ALPHALAB· AUTO-RESEARCH @ MORGAN STANLEY0.96
- 08 / 180.92
- TRACK 3· JULY 2, 20260.93
- Al in Finance0.97
- AlEngineer1.00
- World'sFair1.00
Transcript
169 cues· 3,277 words· 17,457 chars
- 0:12 Well, thank you everyone for coming.
- 0:15 Today I'll be presenting what we've been building at Morgan Stanley, a sort of auto research agent to try and automate quant research.
- 0:25 And I wanted to just take the first couple of minutes to sort of explain, give some context on our team and also why we're even pursuing trying to build this auto research agent.
- 0:37 And so our group were relatively small.
- 0:39 We're about 30 PhD AI researchers and we kind of,
- 0:44 operate both like kind of half academic, so we're encouraged to publish papers, open source code, share our research, and then the other half is sort of more applied internal work.
- 0:56 And I think a lot of our problems, or at least some of them can be fairly well posed.
- 1:02 They kind of have a similar shape to Kaggle where we have
- 1:06 an input time series data set and our task is to just predict future values and maybe with some other constraints of wanting the model to be well calibrated.
- 1:17 And so I think being on the sales side we kind of have maybe less sort of adversarial selection and so it kind of lends itself naturally to this auto research framework.
- 1:29 And also, I think we have a lot of algorithms in production where we have a feeling if we could just put in more cycles, there might be some improvements to kind of squeeze out, whether it's just better hyperparameter tuning or exploring a lot of different kind of ensembling these different methods together.
- 1:47 Also, we work with a lot of different desks, and there's a feeling that something we have built for, say, credit bonds could really transfer well to Mooney bonds, and that kind of translation process should be something that can be automatable with agents.
- 2:03 And even going to new trading desks and trying to build them new algorithms, there's sort of a lot of low-hanging fruit where they're not really using a lot of machine learning or AI automation.
- 2:15 even like a reasonably trained model should be able to have a big impact.
- 2:19 And so sort of in the last year and a half where these agents became able to do long horizon tasks, we've been interested in trying to do this automation.
- 2:30 But it wasn't really until December of 2025, and I think this is a pretty common sentiment now that
- 2:37 It really felt possible for the first time with Opus 4.5 and with these harnesses like Cloud Code and Codex where it really felt like the models were at a point where they could do these long horizon tasks and also kind of the idea of putting them in these harnesses that allow them to do so.
- 2:55 And so really starting this year, we made it a big effort to build this auto research agent.
- 3:01 And of course, the top line goal is just maximizing P&L and maximizing the number of new algorithms we can put in production.
- 3:11 But we also had these design concerns.
- 3:13 Of course, we wanted to play nicely and integrate with all of our data and the pre-built scaffolding we have around back testing and evals.
- 3:23 We also wanted it to be able to run the full spectrum.
- 3:27 You know, on one side, maybe it's a new data set, and we just want to say, like, here's the path to the data, a natural language description of what we hope to predict, and the thing should go off and do research, build its own eval, and start doing experimentation.
- 3:43 And then sort of on the other side of the spectrum, maybe it's like, no, we have our data scripts, we have our eval, we actually have a few good models we've already produced, and we just wanted to kind of churn and do more cycles and see if it can find an improvement.
- 3:57 Also, we wanted to build this to be really model agnostic, so we could use any of the frontier providers or increasingly any open source model.
- 4:07 And also, we wanted to build it in a way that as these models get better and better, it's not sort of consuming what we've built.
- 4:16 We kind of rise with the tide of the models.
- 4:19 And lastly, really carefully think about how do we encode our enterprise knowledge as Morgan Stanley and then our human expertise as quant researchers.
- 4:30 And so for the rest of the talk, I want to start with what I'm calling Alpha Lab 1.0, which is our first version.
- 4:38 We released, I think it was early April, and we put out a full
- 4:43 like 40 page tech report going through all the details and results.
- 4:46 We also open source all the code on GitHub.
- 4:49 So I kind of want to cover that more at a high level because all the details are so public, but I'm happy to talk afterwards in depth about any part.
- 5:00 But then I really want to cover sort of what's happened since then.
- 5:03 So what were the initial results?
- 5:06 What were the kind of failure cases since then?
- 5:08 Because we have encountered a lot of failures.
- 5:11 And then talk about how we're really addressing those by building kind of our own rich set of evals and environments and how that's sort of allowing us to improve the harness and climb towards this sort of self-recursive improvement and sort of our grand vision now for Alpha Lab 2.0.
- 5:30 And so to start, you know, Alpha Lab is an agentic harness and kind of going towards that first side of the spectrum.
- 5:39 The goal is you have some data set, you know, let's say it's like a exchange rate data set or something and you can just provide the path to the file or maybe it lives on an API and you can just say here's the API access and the API spec.
- 5:54 And then just in natural language say what you want to predict.
- 5:57 So maybe it's as simple as
- 6:00 the simple exchange rate, I just am curious in predicting one day out what the rate will be.
- 6:06 And the harness then works in these three phases.
- 6:08 So the first phase is research, and I'll cover these all more in depth.
- 6:12 The second phase is then actually building its own evaluation or back testing.
- 6:17 And the third phase is sort of the mass experimentation, which is really the heart of the harness.
- 6:24 And then as output, you get a suite of trained machine learning models that are trying to do the prediction you care about.
- 6:34 And one design choice we made, so the harness is actually all our own code.
- 6:40 So we decided not to use any off-the-shelf agent framework.
loading
Chapters
- 0:00 Introduction: a thirty person research group
- 1:30 What changed when coding agents arrived
- 2:55 Building AlphaLab 1.0, open sourced
- 4:34 Encoding enterprise standards into the system
- 6:40 Why they skipped off the shelf harnesses
- 8:21 How the research loop runs
- 9:36 Strategist and worker agents
- 10:40 Guarding against a bad eval
- 13:23 Finding real improvements
- 16:12 Verifiable environments, like Kaggle
- 18:45 The human job in the limit