read-only demo

Videos 3jGAU2sbAyY

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

index_state ready data_status ok

AI Engineer· published 2026-05-09· 0:22:26· en-US· indexed 2026-08-10 22:23

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:21, 1 of 1 keyframes kept
  5. Shot 4, 0:21 to 0:34, 1 of 1 keyframes kept
  6. Shot 5, 0:34 to 0:53, 1 of 1 keyframes kept
  7. Shot 6, 0:53 to 0:54, 1 of 1 keyframes kept
  8. Shot 7, 0:54 to 1:09, 0 of 1 keyframes kept
  9. Shot 8, 1:09 to 1:17, 1 of 1 keyframes kept
  10. Shot 9, 1:17 to 1:47, 1 of 1 keyframes kept
  11. Shot 10, 1:47 to 2:10, 1 of 1 keyframes kept
  12. Shot 11, 2:10 to 2:37, 0 of 1 keyframes kept
  13. Shot 12, 2:37 to 3:04, 0 of 1 keyframes kept
  14. Shot 13, 3:04 to 3:31, 0 of 1 keyframes kept
  15. Shot 14, 3:31 to 3:32, 1 of 1 keyframes kept
  16. Shot 15, 3:32 to 4:01, 1 of 1 keyframes kept
  17. Shot 16, 4:01 to 4:29, 0 of 1 keyframes kept
  18. Shot 17, 4:29 to 4:58, 1 of 1 keyframes kept
  19. Shot 18, 4:58 to 5:27, 1 of 1 keyframes kept
  20. Shot 19, 5:27 to 5:56, 1 of 1 keyframes kept
  21. Shot 20, 5:56 to 5:58, 1 of 1 keyframes kept
  22. Shot 21, 5:58 to 6:01, 0 of 1 keyframes kept
  23. Shot 22, 6:01 to 6:27, 0 of 1 keyframes kept
  24. Shot 23, 6:27 to 6:53, 0 of 1 keyframes kept
  25. Shot 24, 6:53 to 7:18, 0 of 1 keyframes kept
  26. Shot 25, 7:18 to 7:56, 0 of 1 keyframes kept
  27. Shot 26, 7:56 to 7:57, 1 of 1 keyframes kept
  28. Shot 27, 7:57 to 8:25, 1 of 1 keyframes kept
  29. Shot 28, 8:25 to 8:53, 0 of 1 keyframes kept
  30. Shot 29, 8:53 to 9:22, 0 of 1 keyframes kept
  31. Shot 30, 9:22 to 9:51, 0 of 1 keyframes kept
  32. Shot 31, 9:51 to 10:20, 0 of 1 keyframes kept
  33. Shot 32, 10:20 to 10:49, 0 of 1 keyframes kept
  34. Shot 33, 10:49 to 11:21, 0 of 1 keyframes kept
  35. Shot 34, 11:21 to 11:49, 0 of 1 keyframes kept
  36. Shot 35, 11:49 to 12:17, 0 of 1 keyframes kept
  37. Shot 36, 12:17 to 12:45, 0 of 1 keyframes kept
  38. Shot 37, 12:45 to 13:13, 0 of 1 keyframes kept
  39. Shot 38, 13:13 to 13:44, 0 of 1 keyframes kept
  40. Shot 39, 13:44 to 14:15, 0 of 1 keyframes kept
  41. Shot 40, 14:15 to 14:51, 1 of 1 keyframes kept
  42. Shot 41, 14:51 to 15:18, 0 of 1 keyframes kept
  43. Shot 42, 15:18 to 15:48, 1 of 1 keyframes kept
  44. Shot 43, 15:48 to 16:03, 0 of 1 keyframes kept
  45. Shot 44, 16:03 to 16:21, 0 of 1 keyframes kept
  46. Shot 45, 16:21 to 16:46, 0 of 1 keyframes kept
  47. Shot 46, 16:46 to 17:12, 0 of 1 keyframes kept
  48. Shot 47, 17:12 to 17:54, 0 of 1 keyframes kept
  49. Shot 48, 17:54 to 18:40, 1 of 1 keyframes kept
  50. Shot 49, 18:40 to 19:06, 1 of 1 keyframes kept
  51. Shot 50, 19:06 to 19:33, 1 of 1 keyframes kept
  52. Shot 51, 19:33 to 20:00, 0 of 1 keyframes kept
  53. Shot 52, 20:00 to 20:01, 0 of 1 keyframes kept
  54. Shot 53, 20:01 to 20:31, 0 of 1 keyframes kept
  55. Shot 54, 20:31 to 20:53, 0 of 1 keyframes kept
  56. Shot 55, 20:53 to 21:00, 0 of 1 keyframes kept
  57. Shot 56, 21:00 to 21:29, 0 of 1 keyframes kept
  58. Shot 57, 21:29 to 21:30, 0 of 1 keyframes kept
  59. Shot 58, 21:30 to 21:57, 0 of 1 keyframes kept
  60. Shot 59, 21:57 to 22:10, 1 of 1 keyframes kept
  61. Shot 60, 22:10 to 22:25, 1 of 1 keyframes kept

61 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
193
whisperx 193
chunks
39
from 193 cues
keyframes
25
kept of 61 captured
frames with text
25
621 lines read
chapters
17
from the source metadata
keyframe bytes
5.6 MB
word timings on 193 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 22:19 1m 22s
stt done 2026-08-10 22:21 22s
chunk done 2026-08-10 22:21 0s
text_embed done 2026-08-10 22:21 0s
keyframe done 2026-08-10 22:21 1m 52s
ocr done 2026-08-10 22:23 11s
frame_embed done 2026-08-10 22:23 4s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 827.9

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:17 #3 done2 line(s)

    shot 3·sharpness 307.6

    1. Mainstream TTS architecture0.98
    2. Samuel Humeau - The Mistral Al Team0.99
  • 0:22 #4 done25 line(s)

    shot 4·sharpness 1861.5

    1. Intro1.00
    2. 4mistralai/Voxtral-4B-TTS-2603like 698FollowingMistral Al_ 16.6k0.96
    3. Text-to-Speech1.00
    4. 9 languagesvllm0.98
    5. mistral-common1.00
    6. arxiv:2603.255511.00
    7. License: cc-by-nc-4.00.99
    8. AIE1.00
    9. 1.00
    10. Model card0.97
    11. Files and versions xet0.97
    12. Community360.96
    13. 1.00
    14. 1.00
    15. 1.00
    16. Voxtral 4B TTS 26030.98
    17. Voxtral TTS is a frontier, open-weights text-to-speech model that's fast, instantly adaptable, and1.00
    18. produces lifelike speech for voice agents. The model is released with BF16 weights and a set of1.00
    19. reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model0.99
    20. inherits.1.00
    21. For more details, see our:1.00
    22. Demo1.00
    23. Blog.post0.96
    24. Research Paper0.99
    25. Google DeepMind1.00
  • 0:49 #5 done16 line(s)

    shot 5·sharpness 2708.7

    1. Intro1.00
    2. This talk is addressed to builders interested in knowing more about how0.99
    3. TTS models work.1.00
    4. *★0.63
    5. 1.00
    6. AIE1.00
    7. 1.00
    8. 1.00
    9. Over the last 3 years, a dominant trend has emerged in TTS0.99
    10. 1.00
    11. 1.00
    12. 1.00
    13. architectures, which we'll review.1.00
    14. This talk happens as we (Mistral) released our first TTS model, which1.00
    15. we'll take as example.0.96
    16. Engineering the future of Al0.99
  • 0:54 #6 done7 line(s)

    shot 6·sharpness 836.0

    1. AIE1.00
    2. 1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. # Braintrust0.98
    7. WorkOS OpenAI0.96
  • 1:01 #7 skipped

    shot 7·duplicate of #5

  • 1:12 #8 done11 line(s)

    shot 8·sharpness 2026.7

    1. About Me1.00
    2. Samuel Humeau1.00
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. 0.98
    7. Al Scientist, Mistral0.98
    8. Previously Facebook FAIR, Nabla0.99
    9. Likes: S2T, Research, product0.99
    10. AlEngineer0.97
    11. EUROPE1.00
  • 1:24 #9 done16 line(s)

    shot 9·sharpness 2775.3

    1. About Mistral0.99
    2. Frontier lab founded in 2023.1.00
    3. Enable people and organizations1.00
    4. AIE1.00
    5. to unlock the transformational1.00
    6. 0.99
    7. 1.00
    8. value of frontier Al with open and1.00
    9. flexible solutions.1.00
    10. Produces frontier models but also1.00
    11. help organizations with their1.00
    12. custom needs via our products:0.99
    13. Forge, Mistral compute, Vibe for1.00
    14. Work.1.00
    15. AlEngineer0.97
    16. EUROPE1.00
  • 2:07 #10 done9 line(s)

    shot 10·sharpness 1964.7

    1. Text-To-Speech1.00
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. A0.79
    6. Listen to the article0.99
    7. Talk to our agent1.00
    8. AlEngineer0.97
    9. EUROPE1.00
  • 2:21 #11 skipped

    shot 11·duplicate of #5

  • 2:40 #12 skipped

    shot 12·duplicate of #5

  • 3:22 #13 skipped

    shot 13·duplicate of #9

  • 3:31 #14 done17 line(s)

    shot 14·sharpness 2230.0

    1. Voxtral TTS and perspective for agent building1.00
    2. Local demo: [here]1.00
    3. Fallback videos1.00
    4. AIE1.00
    5. Vaxtral0.95
    6. 1.00
    7. 1.00
    8. 1.00
    9. Teet io ech0.82
    10. = Pad =.o0.63
    11. Voice Agent0.98
    12. ASSISTANT CONTEXT0.84
    13. ave knowledge of the Abbey Schedule tor April 8:0.89
    14. -12. MTak by Luke Hae (EveLan. Orowth0.53
    15. 1220 PM Reacfy mini ghing a boly e A by0.69
    16. AlEngineer0.95
    17. EUROPE1.00
  • 3:52 #15 done62 line(s)

    shot 15·sharpness 1816.6

    1. Chrome1.00
    2. File1.00
    3. Edit0.86
    4. History1.00
    5. Bookmarks1.00
    6. Profiles1.00
    7. Tab1.00
    8. Window Help0.97
    9. Thu Apr 9 11:450.97
    10. mistral-tts0.97
    11. 41.00
    12. External Headphones0.99
    13. localhost:51731.00
    14. a☆0.60
    15. Unear - Perso0.93
    16. Calendar0.98
    17. Deskare1.00
    18. Lucca0.99
    19. WaB0.80
    20. OMy PRs0.93
    21. M Malbox0.95
    22. Le Chat0.93
    23. Al Studio - Mistral0.95
    24. OGithub0.91
    25. Linkediln0.89
    26. Sheets1.00
    27. Docs0.98
    28. Paragraph Viewer1.00
    29. Chunk process0.91
    30. Voxtral1.00
    31. VOICE1.00
    32. Al Engineer London0.98
    33. Paul on_us0.95
    34. Neutral1.00
    35. Text to Speech1.00
    36. TEXT1.00
    37. Stop previene0.79
    38. Voice Agent0.98
    39. And so with the sunshine and the great bursts of leaves growing0.99
    40. **0.84
    41. 1.00
    42. u0.74
    43. Bitrate1.00
    44. on the trees, just as things grow in fast movies, I had that1.00
    45. familiar conviction that life was beginning over again with the0.99
    46. AIE1.00
    47. 1.00
    48. Checks1.00
    49. summer.1.00
    50. 1.00
    51. 1.00
    52. 1.00
    53. 1.00
    54. 1.00
    55. Output will appear here1.00
    56. 196 chars1.00
    57. PCM (fastest)1.00
    58. Generate0.94
    59. Powered by Mistral Al0.96
    60. Z0.89
    61. AlEngineer0.97
    62. EUROPE1.00
  • 4:26 #16 skipped

    shot 16·duplicate of #15

  • 4:52 #17 done70 line(s)

    shot 17·sharpness 1984.9

    1. Chrome1.00
    2. File1.00
    3. Edit0.98
    4. View1.00
    5. History1.00
    6. Blookmarks0.97
    7. Profiles1.00
    8. Tab1.00
    9. Window Help0.98
    10. q0.74
    11. Thu Apr 9 11:460.99
    12. mistral-ts0.98
    13. localhost:51731.00
    14. 0.98
    15. Work0.99
    16. New Chrome available0.99
    17. 0.56
    18. Linear - Perso0.93
    19. Calendar0.97
    20. Deskare0.99
    21. Lucca0.99
    22. waB0.82
    23. My PRs0.99
    24. M Malbox0.93
    25. Le Chat0.94
    26. Al Studio - Mistral0.96
    27. OGithub0.87
    28. Linkedln0.85
    29. Sheets0.97
    30. Docs0.97
    31. Paragraph Viewer1.00
    32. Chunk process1.00
    33. Voxtral1.00
    34. VOICE1.00
    35. Start session0.94
    36. Al Engineer London0.99
    37. Paul en_us0.93
    38. Neutral1.00
    39. Text to Speech0.98
    40. MODEL1.00
    41. Voice Agent1.00
    42. mistral-small-25060.98
    43. Start a session and speak — your words will appear here.0.99
    44. **0.87
    45. 1.00
    46. u0.58
    47. Bitrate1.00
    48. ASSISTANT CONTEXT1.00
    49. AIE1.00
    50. 1.00
    51. Checks1.00
    52. No lists, no markdown, no lengthy explanations.1.00
    53. You are a concise voice assistant. Reply in 1–2 short sentences.0.98
    54. 1.00
    55. 1.00
    56. 1.00
    57. 1.00
    58. 1.00
    59. You have knowledge of the Abbey Schedule for April 9:0.98
    60. - 11:15 AM — "Beyond Transcription: Building Voice Al That0.97
    61. Actually Understands Conversations* by Hervé Bredin0.98
    62. - 11:40 AM — Tak by Samuel Humeau (Mistral, Al Scientist)0.98
    63. - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.99
    64. - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.97
    65. Marafioti (Hugging Face, Multimodal Research)0.99
    66. - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.98
    67. OCR Pipelines to Multimodal...* (speaker not available)0.97
    68. Reset conversation1.00
    69. Z0.94
    70. Engineering the future of Al0.99
  • 5:24 #18 done78 line(s)

    shot 18·sharpness 2240.1

    1. Chrome1.00
    2. File0.99
    3. Edit0.99
    4. View1.00
    5. History1.00
    6. Bookmarks1.00
    7. Profiles0.99
    8. Tab1.00
    9. Window Help0.98
    10. 00.59
    11. Thu Apr 9 11:460.97
    12. mistral-ts0.99
    13. a0.67
    14. localhost:51730.99
    15. @☆0.75
    16. Work0.98
    17. New Chrome avallable0.96
    18. 0.89
    19. Unear - Perso0.98
    20. Calendar0.95
    21. Deskare1.00
    22. Lucca0.98
    23. waB0.62
    24. My PRs0.99
    25. M Malbox0.90
    26. Le Chat0.95
    27. AI Studio - Mistral0.97
    28. OGithub0.93
    29. Linkedln0.97
    30. Sheets0.92
    31. Docs0.96
    32. Paragraph Viewer0.97
    33. Chunk process0.99
    34. Voxtral1.00
    35. VOICE1.00
    36. Stop1.00
    37. Assistant speaking..0.97
    38. Al Engineer London1.00
    39. Paul an_us0.92
    40. Neutral1.00
    41. Text to Speech1.00
    42. MODEL1.00
    43. Hey, Paul, can you tell me what's the title of the session at 1220, please?0.97
    44. 1.00
    45. Voice Agent1.00
    46. mistral-small-25061.00
    47. The session at 12:20 PM is titled *Reachy mini: giving a body to Al* by Andrés Marafioti.0.98
    48. 1.00
    49. Bitrate1.00
    50. ASSISTANT CONTEXT1.00
    51. And what was the session at 11.15, please?1.00
    52. 1.00
    53. 1.00
    54. AIE1.00
    55. 1.00
    56. 1.00
    57. 1.00
    58. 1.00
    59. Checks1.00
    60. No lists, no markdown, no lengthy explanations.0.98
    61. You are a concise voice assistant. Reply in 1-2 short sentences.0.98
    62. You have knowledge of the Abbey Schedule for April 9:0.98
    63. - 11:15 AM — "Beyond Transcription: Building Voice Al That0.98
    64. The session at 11:15 AM is *Beyond Transcription: Building Voice Al That Actually Understands0.99
    65. Conversations" by Hervé Bredin.0.99
    66. Actually Understands Conversations* by Hervé Bredin0.98
    67. - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.98
    68. - 11:40 AM — Talk by Samuel Humeau (Mistral, Al Scientist)0.99
    69. 0:04/0:080.94
    70. - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.96
    71. Marafioti (Hugging Face, Multimodal Research)0.99
    72. - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.97
    73. OCR Pipelines to Multimodal...* (speaker not available)0.97
    74. Reset conversation0.97
    75. Powered by Mistral Al0.95
    76. Z0.91
    77. AlEngineer0.96
    78. EUROPE1.00
  • 5:44 #19 done76 line(s)

    shot 19·sharpness 2418.8

    1. Chrome1.00
    2. File0.98
    3. Edit0.96
    4. View1.00
    5. History1.00
    6. Bookmarks0.99
    7. Profiles1.00
    8. Tab1.00
    9. Window Help0.96
    10. Thu Apr 9 11:461.00
    11. mistral-tts0.89
    12. a0.86
    13. 0.88
    14. localhost:51730.99
    15. a☆0.74
    16. Work1.00
    17. New Chrome avalable0.99
    18. Unear - Perso0.96
    19. Calendar0.95
    20. Deskare1.00
    21. Lucca0.99
    22. W&B0.67
    23. My PRs0.99
    24. M Maibox0.87
    25. Le Chat0.90
    26. AIl Studio - Mistral.0.89
    27. GithubUinkedin0.95
    28. Sheets1.00
    29. Docs0.97
    30. Paragraph Viewer0.98
    31. Chunk process0.99
    32. Voxtral1.00
    33. VOICE1.00
    34. Start session0.97
    35. Al Engineer London1.00
    36. Paul on_us0.92
    37. Neutral1.00
    38. Text to Speech1.00
    39. MODEL1.00
    40. Hey, Paul, can you tell me what's the title of the session at 1220, please?0.97
    41. Voice Agent1.00
    42. mistral-small-25061.00
    43. The session at 12:20 PM is titled *Reachy mini: giving a body to Al* by Andrés Marafioti.0.97
    44. 1.00
    45. Bitrate0.98
    46. ASSISTANT CONTEXT0.99
    47. And what was the session at 11.15, please?1.00
    48. 1.00
    49. 1.00
    50. AIE1.00
    51. 1.00
    52. 1.00
    53. 1.00
    54. 1.00
    55. Checks1.00
    56. No lists, no markdown, no lengthy explanations.1.00
    57. Actually Understands Conversations* by Hervé Bredin0.99
    58. You are a concise voice assistant. Reply in 1–2 short sentences.0.98
    59. You have knowledge of the Abbey Schedule for April 9:0.98
    60. - 11:15 AM — "Beyond Transcription: Building Voice Al That0.98
    61. The session at 11:15 AM is *Beyond Transcription: Building Voice Al That Actually Understands0.98
    62. Conversations* by Hervé Bredin.0.99
    63. Did you enjoy it as much as I did?0.99
    64. - 11:40 AM — Talk by Samuel Humeau (Mistral, AI Scientist)0.98
    65. - 12:00 PM — Talk by Luke Harries (ElevenLabs, Growth Lead)0.99
    66. I don't have personal experiences or emotions, but I'm glad you enjoyed it!0.99
    67. - 12:20 PM — "Reachy mini: giving a body to Al" by Andrés0.98
    68. Marafioti (Hugging Face, Multimodal Research)0.99
    69. - 12:40 PM — "Rethinking Document Reading for Legal Al: From0.97
    70. OCR Pipelines to Multimodal...* (speaker not available)0.97
    71. 0:00/0:030.94
    72. Reset conversation1.00
    73. Powered by Mistral Al0.95
    74. Z0.92
    75. AlEngineer0.96
    76. EUROPE1.00
  • 5:58 #20 done12 line(s)

    shot 20·sharpness 2388.7

    1. Digression 1: vocal identity1.00
    2. Most large companies already have their vocal identity (advertisement,0.99
    3. the “voice of the company”)0.98
    4. AIE1.00
    5. 0.99
    6. 1.00
    7. 1.00
    8. 1.00
    9. 1.00
    10. Vocal identity will be part of branding of a larger number of companies.1.00
    11. AlEngineer0.97
    12. EUROPE1.00
  • 6:01 #21 skipped

    shot 21·duplicate of #19

  • 6:21 #22 skipped

    shot 22·duplicate of #15

  • 6:45 #23 skipped

    shot 23·duplicate of #15

Transcript

193 cues· 2,986 words· 15,884 chars

  1. 0:14 So I'm from Mistral AI, and we are going to talk about speech generation and text-to-speech.
  2. 0:22 There is an occasion.
  3. 0:23 We released last week our first text-to-speech model, and it's open source, so I really encourage you to check it out.
  4. 0:29 It's an extremely strong text-to-speech model.
  5. 0:32 We are very proud of it.
  6. 0:35 And for this occasion, I thought we could review some of the recent trend in text-to-speech architecture, since there is a dominant trend emerging these days, although this can change very quickly.
  7. 0:54 And so this talk is slightly academic and addressed to people who want to know a bit more about how you do text-to-speech.
  8. 1:01 This being said, we have a few years before the machines do all the science for us, so we might enjoy it today.
  9. 1:10 I'm Sam.
  10. 1:12 Yeah, I work at Mistral as an AI scientist.
  11. 1:14 Before, I was at Facebook Fair when it was called Facebook.
  12. 1:19 And Mistral, a few words about the company.
  13. 1:22 We are a frontier lab.
  14. 1:24 We've been founded a couple of years ago.
  15. 1:28 We produce frontier model, but we're also a B2B business.
  16. 1:32 We help organization in their AI transformation, which is kind of a buzzword, but literally every company is transforming with AI.
  17. 1:39 We help them by providing them tools, product, and dedicated people to help them in their custom needs.
  18. 1:49 Back to the text-to-speech.
  19. 1:50 So there are a few offline use case of speech generation, like the famous listen to the blog or listen to the article.
  20. 1:59 But nowadays, the king use case for text-to-speech is its usage within agents.
  21. 2:06 And in particular, it's used to interface with a chat agent, typically in a pipe like this, where you have a central chat agent that does text-to-text, but does it extremely well, and you want to talk to it, so you add a speech-to-text, and you want it to speak to you, so you add a text-to-speech.
  22. 2:26 As everybody in this conf will tell you the latency is key here So you can reduce the latency on the left by having the speech-to-text done in real time So that when you detect the end of term you already have the transcript.
  23. 2:39 It's already done And we are gonna focus a bit on the right side today it's also very important that as soon as you have the first audio packets you you start to to voice them out this way the perceived latency is lower and
  24. 2:57 In fact, since your LLM can stream some text to you, actually what you ultimately want is something like this, if you're going to interface a chat assistant, which is a real-time text input, text-to-speech, where as soon as you have the first token of the LLM, the machine starts to speak.
  25. 3:19 We're going to talk about it in the end of the talk.
  26. 3:23 I want to focus a bit at the beginning, at the output side, and what it means to stream audio.
  27. 3:30 So to illustrate this, I have this app that I vibe-coded for the occasion.
  28. 3:38 And so we are going to use this text-to-speech model that we released, that I mentioned.
  29. 3:44 And we are going to hear Paul.
  30. 3:45 So Paul is an actual human being that sounds like this.
  31. 3:50 The persistent anxiety that fills the rest of my life is calmed for as long as I have the flavor of something good in my mouth.
  32. 4:00 So this is like some actual recording on some actual person named Paul and we're copying his voice.
  33. 4:08 So with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies, I had that familiar conviction that life was beginning over again with the summer.
  34. 4:18 So let's focus first on what's happening here.
  35. 4:22 As you can see, we are copying the voice, and the first audio packet happens first, and we can start to emit audio, which greatly reduces the perceived latency, even though the full computation of the audio happens like a few seconds later.
  36. 4:40 So if you use it in an agent, so here I crafted a small agent using a speech-to-text, one of our LLM, and this very text-to-speech, so we can speak to Paul.
  37. 4:55 And hey, Paul, can you tell me what's the title of the session at 12.20, please?
  38. 5:04 The session at 12.20 PM is titled Ritchie Mini, Giving a Body to AI by Andres Marafioti.
  39. 5:11 And what was the session at 11.15, please?
  40. 5:19 The session at 11.15 AM is Beyond Transcription, Building Voice AI That Actually Understands Conversations by Hervé Bredin.
  41. 5:29 Did you enjoy it as much as I did?
  42. 5:33 I don't have personal experiences or emotions, but I'm glad you enjoyed it.
  43. 5:38 That's all I can do.
  44. 5:41 So the important thing here is that since the audio packet arrived first, you still have a decent latency, and you can enjoy the conversation with the agent, despite the fact that the audio is still not generated fully.
  45. 5:56 And so we're going to dig, oh, yeah, sorry.
  46. 5:59 I want to make one digression.
  47. 6:00 So I mentioned the voice cloning here.
  48. 6:03 This model can only need a few seconds to clone the voice of someone.
  49. 6:10 So again, this is how it sounded.
  50. 6:13 And this is how we generate text.

Chapters

  1. 0:00 Introduction and Mistral's new open-source TTS model
  2. 2:06 Text-to-speech in AI agents and latency
  3. 3:33 Live demo: Voice cloning with 'Paul'
  4. 6:00 Voice cloning capabilities and multilingual examples
  5. 8:01 Historical context of audio generation
  6. 8:55 Transformer-based architecture for TTS
  7. 10:00 Challenges of information density in audio
  8. 10:55 Comparison of bit rates: text vs. audio
  9. 11:39 Using neural audio codecs
  10. 13:10 Backbone transformer and frame-based generation
  11. 14:56 Text conditioning and model architecture
  12. 16:08 Latency performance metrics
  13. 16:22 Future outlook: Streaming text input
  14. 17:35 Q&A: Generating text and audio simultaneously
  15. 18:24 Q&A: Availability of voice cloning features
  16. 19:35 Q&A: Philosophical take on speech interfaces
  17. 20:44 Q&A: Next steps for streaming audio and text input

Open at this second