Videos 3jGAU2sbAyY
Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
Scene timeline
61 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 193
- whisperx 193
- chunks
- 39
- from 193 cues
- keyframes
- 25
- kept of 61 captured
- frames with text
- 25
- 621 lines read
- chapters
- 17
- from the source metadata
- keyframe bytes
- 5.6 MB
- word timings on 193 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 22:19 | 1m 22s |
stt |
done | — | 2026-08-10 22:21 | 22s |
chunk |
done | — | 2026-08-10 22:21 | 0s |
text_embed |
done | — | 2026-08-10 22:21 | 0s |
keyframe |
done | — | 2026-08-10 22:21 | 1m 52s |
ocr |
done | — | 2026-08-10 22:23 | 11s |
frame_embed |
done | — | 2026-08-10 22:23 | 4s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Mainstream TTS architecture0.98
- Samuel Humeau - The Mistral Al Team0.99
-
- Intro1.00
- 4mistralai/Voxtral-4B-TTS-2603like 698FollowingMistral Al_ 16.6k0.96
- Text-to-Speech1.00
- 9 languagesvllm0.98
- mistral-common1.00
- arxiv:2603.255511.00
- License: cc-by-nc-4.00.99
- AIE1.00
- ★1.00
- Model card0.97
- Files and versions xet0.97
- Community360.96
- ★1.00
- ★1.00
- ★1.00
- Voxtral 4B TTS 26030.98
- Voxtral TTS is a frontier, open-weights text-to-speech model that's fast, instantly adaptable, and1.00
- produces lifelike speech for voice agents. The model is released with BF16 weights and a set of1.00
- reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model0.99
- inherits.1.00
- For more details, see our:1.00
- Demo1.00
- Blog.post0.96
- Research Paper0.99
- Google DeepMind1.00
-
- Intro1.00
- This talk is addressed to builders interested in knowing more about how0.99
- TTS models work.1.00
- *★0.63
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- Over the last 3 years, a dominant trend has emerged in TTS0.99
- ★1.00
- ★1.00
- ★1.00
- architectures, which we'll review.1.00
- This talk happens as we (Mistral) released our first TTS model, which1.00
- we'll take as example.0.96
- Engineering the future of Al0.99
-
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- # Braintrust0.98
- WorkOS OpenAI0.96
-
- About Me1.00
- Samuel Humeau1.00
- AIE1.00
- ★1.00
- ★1.00
- ★0.98
- Al Scientist, Mistral0.98
- Previously Facebook FAIR, Nabla0.99
- Likes: S2T, Research, product0.99
- AlEngineer0.97
- EUROPE1.00
-
- About Mistral0.99
- Frontier lab founded in 2023.1.00
- Enable people and organizations1.00
- AIE1.00
- to unlock the transformational1.00
- ★0.99
- ★1.00
- value of frontier Al with open and1.00
- flexible solutions.1.00
- Produces frontier models but also1.00
- help organizations with their1.00
- custom needs via our products:0.99
- Forge, Mistral compute, Vibe for1.00
- Work.1.00
- AlEngineer0.97
- EUROPE1.00
-
- Text-To-Speech1.00
- AIE1.00
- ★1.00
- ★1.00
- A0.79
- Listen to the article0.99
- Talk to our agent1.00
- AlEngineer0.97
- EUROPE1.00
-
- Voxtral TTS and perspective for agent building1.00
- Local demo: [here]1.00
- Fallback videos1.00
- AIE1.00
- Vaxtral0.95
- ★1.00
- ★1.00
- ★1.00
- Teet io ech0.82
- = Pad =.o0.63
- Voice Agent0.98
- ASSISTANT CONTEXT0.84
- ave knowledge of the Abbey Schedule tor April 8:0.89
- -12. MTak by Luke Hae (EveLan. Orowth0.53
- 1220 PM Reacfy mini ghing a boly e A by0.69
- AlEngineer0.95
- EUROPE1.00
-
- Chrome1.00
- File1.00
- Edit0.86
- History1.00
- Bookmarks1.00
- Profiles1.00
- Tab1.00
- Window Help0.97
- Thu Apr 9 11:450.97
- mistral-tts0.97
- 41.00
- External Headphones0.99
- localhost:51731.00
- a☆0.60
- Unear - Perso0.93
- Calendar0.98
- Deskare1.00
- Lucca0.99
- WaB0.80
- OMy PRs0.93
- M Malbox0.95
- Le Chat0.93
- Al Studio - Mistral0.95
- OGithub0.91
- Linkediln0.89
- Sheets1.00
- Docs0.98
- Paragraph Viewer1.00
- Chunk process0.91
- Voxtral1.00
- VOICE1.00
- Al Engineer London0.98
- Paul on_us0.95
- Neutral1.00
- Text to Speech1.00
- TEXT1.00
- Stop previene0.79
- Voice Agent0.98
- And so with the sunshine and the great bursts of leaves growing0.99
- **0.84
- ★1.00
- u0.74
- Bitrate1.00
- on the trees, just as things grow in fast movies, I had that1.00
- familiar conviction that life was beginning over again with the0.99
- AIE1.00
- ★1.00
- Checks1.00
- summer.1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Output will appear here1.00
- 196 chars1.00
- PCM (fastest)1.00
- Generate0.94
- Powered by Mistral Al0.96
- Z0.89
- AlEngineer0.97
- EUROPE1.00
-
- Chrome1.00
- File1.00
- Edit0.98
- View1.00
- History1.00
- Blookmarks0.97
- Profiles1.00
- Tab1.00
- Window Help0.98
- q0.74
- Thu Apr 9 11:460.99
- mistral-ts0.98
- localhost:51731.00
- ☆0.98
- Work0.99
- New Chrome available0.99
- 器0.56
- Linear - Perso0.93
- Calendar0.97
- Deskare0.99
- Lucca0.99
- waB0.82
- My PRs0.99
- M Malbox0.93
- Le Chat0.94
- Al Studio - Mistral0.96
- OGithub0.87
- Linkedln0.85
- Sheets0.97
- Docs0.97
- Paragraph Viewer1.00
- Chunk process1.00
- Voxtral1.00
- VOICE1.00
- Start session0.94
- Al Engineer London0.99
- Paul en_us0.93
- Neutral1.00
- Text to Speech0.98
- MODEL1.00
- Voice Agent1.00
- mistral-small-25060.98
- Start a session and speak — your words will appear here.0.99
- **0.87
- ★1.00
- u0.58
- Bitrate1.00
- ASSISTANT CONTEXT1.00
- AIE1.00
- ★1.00
- Checks1.00
- No lists, no markdown, no lengthy explanations.1.00
- You are a concise voice assistant. Reply in 1–2 short sentences.0.98
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- You have knowledge of the Abbey Schedule for April 9:0.98
- - 11:15 AM — "Beyond Transcription: Building Voice Al That0.97
- Actually Understands Conversations* by Hervé Bredin0.98
- - 11:40 AM — Tak by Samuel Humeau (Mistral, Al Scientist)0.98
- - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.99
- - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.97
- Marafioti (Hugging Face, Multimodal Research)0.99
- - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.98
- OCR Pipelines to Multimodal...* (speaker not available)0.97
- Reset conversation1.00
- Z0.94
- Engineering the future of Al0.99
-
- Chrome1.00
- File0.99
- Edit0.99
- View1.00
- History1.00
- Bookmarks1.00
- Profiles0.99
- Tab1.00
- Window Help0.98
- 00.59
- Thu Apr 9 11:460.97
- mistral-ts0.99
- a0.67
- localhost:51730.99
- @☆0.75
- Work0.98
- New Chrome avallable0.96
- 器0.89
- Unear - Perso0.98
- Calendar0.95
- Deskare1.00
- Lucca0.98
- waB0.62
- My PRs0.99
- M Malbox0.90
- Le Chat0.95
- AI Studio - Mistral0.97
- OGithub0.93
- Linkedln0.97
- Sheets0.92
- Docs0.96
- Paragraph Viewer0.97
- Chunk process0.99
- Voxtral1.00
- VOICE1.00
- Stop1.00
- Assistant speaking..0.97
- Al Engineer London1.00
- Paul an_us0.92
- Neutral1.00
- Text to Speech1.00
- MODEL1.00
- Hey, Paul, can you tell me what's the title of the session at 1220, please?0.97
- ★1.00
- Voice Agent1.00
- mistral-small-25061.00
- The session at 12:20 PM is titled *Reachy mini: giving a body to Al* by Andrés Marafioti.0.98
- ★1.00
- Bitrate1.00
- ASSISTANT CONTEXT1.00
- And what was the session at 11.15, please?1.00
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Checks1.00
- No lists, no markdown, no lengthy explanations.0.98
- You are a concise voice assistant. Reply in 1-2 short sentences.0.98
- You have knowledge of the Abbey Schedule for April 9:0.98
- - 11:15 AM — "Beyond Transcription: Building Voice Al That0.98
- The session at 11:15 AM is *Beyond Transcription: Building Voice Al That Actually Understands0.99
- Conversations" by Hervé Bredin.0.99
- Actually Understands Conversations* by Hervé Bredin0.98
- - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.98
- - 11:40 AM — Talk by Samuel Humeau (Mistral, Al Scientist)0.99
- 0:04/0:080.94
- - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.96
- Marafioti (Hugging Face, Multimodal Research)0.99
- - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.97
- OCR Pipelines to Multimodal...* (speaker not available)0.97
- Reset conversation0.97
- Powered by Mistral Al0.95
- Z0.91
- AlEngineer0.96
- EUROPE1.00
-
- Chrome1.00
- File0.98
- Edit0.96
- View1.00
- History1.00
- Bookmarks0.99
- Profiles1.00
- Tab1.00
- Window Help0.96
- Thu Apr 9 11:461.00
- mistral-tts0.89
- a0.86
- ①0.88
- localhost:51730.99
- a☆0.74
- Work1.00
- New Chrome avalable0.99
- Unear - Perso0.96
- Calendar0.95
- Deskare1.00
- Lucca0.99
- W&B0.67
- My PRs0.99
- M Maibox0.87
- Le Chat0.90
- AIl Studio - Mistral.0.89
- GithubUinkedin0.95
- Sheets1.00
- Docs0.97
- Paragraph Viewer0.98
- Chunk process0.99
- Voxtral1.00
- VOICE1.00
- Start session0.97
- Al Engineer London1.00
- Paul on_us0.92
- Neutral1.00
- Text to Speech1.00
- MODEL1.00
- Hey, Paul, can you tell me what's the title of the session at 1220, please?0.97
- Voice Agent1.00
- mistral-small-25061.00
- The session at 12:20 PM is titled *Reachy mini: giving a body to Al* by Andrés Marafioti.0.97
- ★1.00
- Bitrate0.98
- ASSISTANT CONTEXT0.99
- And what was the session at 11.15, please?1.00
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Checks1.00
- No lists, no markdown, no lengthy explanations.1.00
- Actually Understands Conversations* by Hervé Bredin0.99
- You are a concise voice assistant. Reply in 1–2 short sentences.0.98
- You have knowledge of the Abbey Schedule for April 9:0.98
- - 11:15 AM — "Beyond Transcription: Building Voice Al That0.98
- The session at 11:15 AM is *Beyond Transcription: Building Voice Al That Actually Understands0.98
- Conversations* by Hervé Bredin.0.99
- Did you enjoy it as much as I did?0.99
- - 11:40 AM — Talk by Samuel Humeau (Mistral, AI Scientist)0.98
- - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.99
- I don't have personal experiences or emotions, but I'm glad you enjoyed it!0.99
- - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.98
- Marafioti (Hugging Face, Multimodal Research)0.99
- - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.97
- OCR Pipelines to Multimodal...* (speaker not available)0.97
- 0:00/0:030.94
- Reset conversation1.00
- Powered by Mistral Al0.95
- Z0.92
- AlEngineer0.96
- EUROPE1.00
-
- Digression 1: vocal identity1.00
- Most large companies already have their vocal identity (advertisement,0.99
- the “voice of the company”)0.98
- AIE1.00
- ★0.99
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Vocal identity will be part of branding of a larger number of companies.1.00
- AlEngineer0.97
- EUROPE1.00
Transcript
193 cues· 2,986 words· 15,884 chars
- 0:14 So I'm from Mistral AI, and we are going to talk about speech generation and text-to-speech.
- 0:22 There is an occasion.
- 0:23 We released last week our first text-to-speech model, and it's open source, so I really encourage you to check it out.
- 0:29 It's an extremely strong text-to-speech model.
- 0:32 We are very proud of it.
- 0:35 And for this occasion, I thought we could review some of the recent trend in text-to-speech architecture, since there is a dominant trend emerging these days, although this can change very quickly.
- 0:54 And so this talk is slightly academic and addressed to people who want to know a bit more about how you do text-to-speech.
- 1:01 This being said, we have a few years before the machines do all the science for us, so we might enjoy it today.
- 1:10 I'm Sam.
- 1:12 Yeah, I work at Mistral as an AI scientist.
- 1:14 Before, I was at Facebook Fair when it was called Facebook.
- 1:19 And Mistral, a few words about the company.
- 1:22 We are a frontier lab.
- 1:24 We've been founded a couple of years ago.
- 1:28 We produce frontier model, but we're also a B2B business.
- 1:32 We help organization in their AI transformation, which is kind of a buzzword, but literally every company is transforming with AI.
- 1:39 We help them by providing them tools, product, and dedicated people to help them in their custom needs.
- 1:49 Back to the text-to-speech.
- 1:50 So there are a few offline use case of speech generation, like the famous listen to the blog or listen to the article.
- 1:59 But nowadays, the king use case for text-to-speech is its usage within agents.
- 2:06 And in particular, it's used to interface with a chat agent, typically in a pipe like this, where you have a central chat agent that does text-to-text, but does it extremely well, and you want to talk to it, so you add a speech-to-text, and you want it to speak to you, so you add a text-to-speech.
- 2:26 As everybody in this conf will tell you the latency is key here So you can reduce the latency on the left by having the speech-to-text done in real time So that when you detect the end of term you already have the transcript.
- 2:39 It's already done And we are gonna focus a bit on the right side today it's also very important that as soon as you have the first audio packets you you start to to voice them out this way the perceived latency is lower and
- 2:57 In fact, since your LLM can stream some text to you, actually what you ultimately want is something like this, if you're going to interface a chat assistant, which is a real-time text input, text-to-speech, where as soon as you have the first token of the LLM, the machine starts to speak.
- 3:19 We're going to talk about it in the end of the talk.
- 3:23 I want to focus a bit at the beginning, at the output side, and what it means to stream audio.
- 3:30 So to illustrate this, I have this app that I vibe-coded for the occasion.
- 3:38 And so we are going to use this text-to-speech model that we released, that I mentioned.
- 3:44 And we are going to hear Paul.
- 3:45 So Paul is an actual human being that sounds like this.
- 3:50 The persistent anxiety that fills the rest of my life is calmed for as long as I have the flavor of something good in my mouth.
- 4:00 So this is like some actual recording on some actual person named Paul and we're copying his voice.
- 4:08 So with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies, I had that familiar conviction that life was beginning over again with the summer.
- 4:18 So let's focus first on what's happening here.
- 4:22 As you can see, we are copying the voice, and the first audio packet happens first, and we can start to emit audio, which greatly reduces the perceived latency, even though the full computation of the audio happens like a few seconds later.
- 4:40 So if you use it in an agent, so here I crafted a small agent using a speech-to-text, one of our LLM, and this very text-to-speech, so we can speak to Paul.
- 4:55 And hey, Paul, can you tell me what's the title of the session at 12.20, please?
- 5:04 The session at 12.20 PM is titled Ritchie Mini, Giving a Body to AI by Andres Marafioti.
- 5:11 And what was the session at 11.15, please?
- 5:19 The session at 11.15 AM is Beyond Transcription, Building Voice AI That Actually Understands Conversations by Hervé Bredin.
- 5:29 Did you enjoy it as much as I did?
- 5:33 I don't have personal experiences or emotions, but I'm glad you enjoyed it.
- 5:38 That's all I can do.
- 5:41 So the important thing here is that since the audio packet arrived first, you still have a decent latency, and you can enjoy the conversation with the agent, despite the fact that the audio is still not generated fully.
- 5:56 And so we're going to dig, oh, yeah, sorry.
- 5:59 I want to make one digression.
- 6:00 So I mentioned the voice cloning here.
- 6:03 This model can only need a few seconds to clone the voice of someone.
- 6:10 So again, this is how it sounded.
- 6:13 And this is how we generate text.
loading
Chapters
- 0:00 Introduction and Mistral's new open-source TTS model
- 2:06 Text-to-speech in AI agents and latency
- 3:33 Live demo: Voice cloning with 'Paul'
- 6:00 Voice cloning capabilities and multilingual examples
- 8:01 Historical context of audio generation
- 8:55 Transformer-based architecture for TTS
- 10:00 Challenges of information density in audio
- 10:55 Comparison of bit rates: text vs. audio
- 11:39 Using neural audio codecs
- 13:10 Backbone transformer and frame-based generation
- 14:56 Text conditioning and model architecture
- 16:08 Latency performance metrics
- 16:22 Future outlook: Streaming text input
- 17:35 Q&A: Generating text and audio simultaneously
- 18:24 Q&A: Availability of voice cloning features
- 19:35 Q&A: Philosophical take on speech interfaces
- 20:44 Q&A: Next steps for streaming audio and text input