Videos AVMr9PMINyo
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Dan Fu and Olive Song
Scene timeline
51 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 250
- whisperx 250
- chunks
- 36
- from 250 cues
- keyframes
- 48
- kept of 51 captured
- frames with text
- 48
- 118 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 5.3 MB
- word timings on 250 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:47 | 0s |
stt |
done | — | 2026-08-09 02:52 | 24s |
chunk |
done | — | 2026-08-09 02:53 | 0s |
text_embed |
done | — | 2026-08-10 19:37 | 1s |
keyframe |
done | — | 2026-08-09 02:53 | 2m 53s |
ocr |
done | — | 2026-08-09 02:55 | 7s |
frame_embed |
done | — | 2026-08-10 19:37 | 8s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer1.00
- World's Fair0.99
-
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.97
- World's Fair1.00
-
- AlEngineer1.00
- World's Fair0.99
-
- AlEngineer0.99
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.97
- World's Fair1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- AlEngineer0.98
- World's Fair1.00
-
- AlEngineer0.98
- World's Fair0.97
-
- AlEngineer0.98
- World's Fair0.99
-
- AlEngineer0.97
- World's Fair1.00
-
- AlEngineer0.97
- World's Fair1.00
-
- AlEngineer0.99
- World's Fair1.00
Transcript
250 cues· 3,628 words· 19,400 chars
- 0:12 This is a discussion that I'm particularly excited about because the field is moving so fast and we have two people that kind of have this unique vantage point on the field.
- 0:22 And what I want to do is kind of ask questions to see if we can learn from that.
- 0:26 So I want to start off with intros, talk a little bit about your role, what you're thinking about, what you're working on.
- 0:31 Maybe Dan, if you can go first.
- 0:33 Hey, everyone.
- 0:34 I'm Dan.
- 0:34 I'm the VP of Kernels at Together AI.
- 0:37 I lead inference, GPU optimization, trying to figure out how to use GPUs most effectively to serve AI models.
- 0:45 Yeah, so one of the things that I wanted to dive in with Dan about is new model drops, what is everything that goes on behind the scenes to serve it so that everybody here, there's a lot of builders here that can use it.
- 0:57 Olive, I want to throw it over to you to talk about your role and what you're focusing on.
- 1:01 Yeah, I'm Olive, and I am the research lead of RL at Minimax.
- 1:06 And I am responsible for the final training of the model and the shipping of the model, so basically everything before the infrastructure.
- 1:13 Awesome.
- 1:14 So maybe I want to start off, this panel is focusing on open source.
- 1:18 I wanted to start off with, this is your strongest model yet, Minimax M3.
- 1:23 Why open source it?
- 1:24 What's the idea behind that as a company as you're releasing these models?
- 1:29 We do believe that the open source community as a whole is very strong and powerful.
- 1:35 While we open source the model, everyone can use it.
- 1:38 So it aligns with our mission that we want to have intelligence with everyone.
- 1:42 And also different developers can contribute to the model through feedback, through their own PRs, and we can build the models even stronger.
- 1:50 And also, for example,
- 1:53 then you will be able to optimize on our open-weight model and make it inference faster and then serve better for everyone.
- 2:00 Yeah.
- 2:01 Yeah, we're big believers in open-source set together.
- 2:03 And yeah, I think we've been following you guys for a while, I think, from way older Minimax models.
- 2:11 So seeing M3 and seeing how far it's come is really impressive and really great.
- 2:16 So I wanted to kind of pick on this a little bit more.
- 2:20 Can you explain?
- 2:21 So we've got the model creators themselves, Minimax.
- 2:23 We've got experts on the inference side of things.
- 2:26 How did this partnership come to be?
- 2:28 So they launch an open source model and we're now distributing it.
- 2:31 I checked this morning.
- 2:32 We have the lion's share of token usage for Minimax M3.
- 2:38 How does this partnership come to be and how do we serve a model like this at scale?
- 2:41 Yeah, yeah, great question.
- 2:42 So at Together, I think one of the things that we're really interested in is how do you make intelligence abundant?
- 2:49 So how do you get more tokens for more people to do more useful things and get all these capabilities into more people's hands?
- 2:58 So we follow all the open models very closely.
- 3:02 I don't remember when exactly we started partnering.
- 3:04 Oh, actually, I think I do know this.
- 3:06 We had a car event in Las Vegas sometime last year, and someone from Minimax came, and there he was like, guys, you really gotta serve our next model.
- 3:14 It's gonna be really, really great.
- 3:16 So I think from there, we started talking.
- 3:18 We were serving Minimax 2.5.
- 3:21 and I think 2.7 for a while.
- 3:24 And then when leading up to the launch of M3, we were quite excited about it, I think.
- 3:30 We were seeing the usage and what people were doing with it.
- 3:34 It was really quite exciting.
loading
Chapters
- 0:00 Introducing the RL lead at MiniMax
- 1:17 Why open source and open weights
- 3:45 What builders are doing with the model
- 4:11 A model that builds games
- 5:12 Computer use and OS World
- 5:49 Writing GPU kernels
- 6:25 Parallel kernel bench
- 7:28 A day zero inference stack
- 9:34 Optimization across the stack
- 10:47 Multimodality and training collapse
- 14:10 Replicating a twelve hour run
- 17:22 Where open models go next