read-only demo

Videos a2muGkT4WD4

Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX — Adrien Grondin, Locally AI

index_state ready data_status ok

AI Engineer· published 2026-04-20· 0:10:50· en-US· indexed 2026-08-10 23:09

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:21, 1 of 1 keyframes kept
  5. Shot 4, 0:21 to 0:32, 1 of 1 keyframes kept
  6. Shot 5, 0:32 to 0:48, 1 of 1 keyframes kept
  7. Shot 6, 0:48 to 0:59, 1 of 1 keyframes kept
  8. Shot 7, 0:59 to 1:11, 1 of 1 keyframes kept
  9. Shot 8, 1:11 to 1:24, 1 of 1 keyframes kept
  10. Shot 9, 1:24 to 1:29, 1 of 1 keyframes kept
  11. Shot 10, 1:29 to 1:48, 1 of 1 keyframes kept
  12. Shot 11, 1:48 to 2:19, 1 of 1 keyframes kept
  13. Shot 12, 2:19 to 2:51, 0 of 1 keyframes kept
  14. Shot 13, 2:51 to 3:23, 0 of 1 keyframes kept
  15. Shot 14, 3:23 to 3:30, 1 of 1 keyframes kept
  16. Shot 15, 3:30 to 4:05, 1 of 1 keyframes kept
  17. Shot 16, 4:05 to 4:40, 1 of 1 keyframes kept
  18. Shot 17, 4:40 to 5:09, 0 of 1 keyframes kept
  19. Shot 18, 5:09 to 5:38, 0 of 1 keyframes kept
  20. Shot 19, 5:38 to 5:47, 1 of 1 keyframes kept
  21. Shot 20, 5:47 to 5:55, 0 of 1 keyframes kept
  22. Shot 21, 5:55 to 6:01, 1 of 1 keyframes kept
  23. Shot 22, 6:01 to 6:04, 1 of 1 keyframes kept
  24. Shot 23, 6:04 to 6:37, 1 of 1 keyframes kept
  25. Shot 24, 6:37 to 6:41, 0 of 1 keyframes kept
  26. Shot 25, 6:41 to 6:48, 0 of 1 keyframes kept
  27. Shot 26, 6:48 to 7:06, 0 of 1 keyframes kept
  28. Shot 27, 7:06 to 7:49, 1 of 1 keyframes kept
  29. Shot 28, 7:49 to 8:03, 1 of 1 keyframes kept
  30. Shot 29, 8:03 to 8:44, 1 of 1 keyframes kept
  31. Shot 30, 8:44 to 9:01, 1 of 1 keyframes kept
  32. Shot 31, 9:01 to 9:32, 1 of 1 keyframes kept
  33. Shot 32, 9:32 to 10:04, 1 of 1 keyframes kept
  34. Shot 33, 10:04 to 10:35, 0 of 1 keyframes kept
  35. Shot 34, 10:35 to 10:50, 1 of 1 keyframes kept

35 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
115
whisperx 115
chunks
19
from 115 cues
keyframes
26
kept of 35 captured
frames with text
26
580 lines read
chapters
0
from the source metadata
keyframe bytes
3.1 MB
word timings on 115 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 23:06 1m 12s
stt done 2026-08-10 23:07 11s
chunk done 2026-08-10 23:07 0s
text_embed done 2026-08-10 23:07 0s
keyframe done 2026-08-10 23:07 1m 02s
ocr done 2026-08-10 23:08 10s
frame_embed done 2026-08-10 23:08 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 668.7

    1. Al Engineer0.93
    2. EUROPE1.00
  • 0:06 #1 done2 line(s)

    shot 1·sharpness 819.1

    1. PRESENTING SPONSOR0.99
    2. Google DeepMind1.00
  • 0:11 #2 done3 line(s)

    shot 2·sharpness 939.2

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.95
    3. WorkOS OpenAI0.96
  • 0:18 #3 done5 line(s)

    shot 3·sharpness 509.3

    1. Running Gemma 4 on iPhone0.99
    2. with MLX0.96
    3. Adrien Grondin - Founder, Locally Al1.00
    4. AlEngineer0.99
    5. EUROPE1.00
  • 0:27 #4 done6 line(s)

    shot 4·sharpness 2430.3

    1. AIE1.00
    2. 1.00
    3. 1.00
    4. @adrgrondin0.96
    5. Google DeepMind1.00
    6. AlEngineer0.97
  • 0:40 #5 done4 line(s)

    shot 5·sharpness 1627.0

    1. AIE1.00
    2. 1.00
    3. Engineering the future of Al0.99
    4. AlEngineer0.96
  • 0:54 #6 done80 line(s)

    shot 6·sharpness 3117.4

    1. 09:411.00
    2. 09:411.00
    3. 09:411.00
    4. à0.66
    5. Apple Foundation>0.98
    6. Q0.74
    7. Apple Foundation>0.98
    8. Manage models1.00
    9. I want to plan a 7-day trip to Paris. Please1.00
    10. suggest must-see attractions, recommend1.00
    11. System1.00
    12. daily itinerary including restaurants and0.99
    13. good neighborhoods to stay in, and create a0.99
    14. Apple Foundation1.00
    15. AIE1.00
    16. 1.00
    17. 1.00
    18. including must-see attractions, neighborhoods0.99
    19. to stay in, and daily activities along with0.99
    20. Absolutely, Paris is a fantastic city with so0.99
    21. much to offer! Here's a suggested itinerary,0.98
    22. activities.1.00
    23. Meet Apple Foundation1.00
    24. understanding, refinement, generating creative0.99
    25. range of text generation tasks,like0.96
    26. content, and more.1.00
    27. On-device model by Apple. Excels at a diverse1.00
    28. summarization, entity extraction, text1.00
    29. 4 New0.75
    30. restaurant recommendations.0.99
    31. 1.00
    32. 1.00
    33. 1.00
    34. 1.00
    35. Must-See Attractions:1.00
    36. 1. Eiffel Tower - Iconic landmark and one of1.00
    37. Locally Al now supports Apple's on-device large0.98
    38. language model, the same model that powers0.99
    39. Apple Intelligence.1.00
    40. Featured1.00
    41. the most visited monuments in the world.0.97
    42. SmolLM 3 (3B)0.99
    43. Download1.00
    44. 3. Notre-Dame Cathedral - A masterpiece1.00
    45. 4. Sacré-Cour Basilica - Offers stunning0.97
    46. 2. Louvre Museum - Home to thousands of1.00
    47. works of art, including the Mona Lisa.0.99
    48. of French Gothic architecture.1.00
    49. views of Paris from Montmartre.1.00
    50. A model made by Hugging Face. Great for0.98
    51. complex reasoning, long conversations, and use0.99
    52. in English, French, Spanish, German, Italian, and1.00
    53. Portuguese. Recommended for iPhone 15 Pro or0.99
    54. 1,73 080.94
    55. newer.1.00
    56. 5. Champs-Élysées and Arc de Triomphe -0.99
    57. Reasoning0.92
    58. Famous avenue and grand monument.0.99
    59. 6. Montmartre and Sacré-Coeur - Artistic0.99
    60. 7. Musée d'Orsay - nowned for its0.95
    61. neighborhood with cobblestone streets.1.00
    62. collection of Imprconist masterpieces.0.97
    63. a trip to Paris0.98
    64. Plan1.00
    65. a complex topic simply0.98
    66. Explain0.96
    67. Desi1.00
    68. in everyday devices. Great for content creation,1.00
    69. A powerful model from Google, optimized for use1.00
    70. G Gemma 3n (E2B)0.97
    71. Download1.00
    72. +0.93
    73. Ask anything1.00
    74. Ask anything0.97
    75. Recommended for iPhone 15 Pro and newer.1.00
    76. text summarization, and conversational Al.0.99
    77. 2,51080.93
    78. Braintrust1.00
    79. WorkOS OpenAI0.96
    80. AlEngineer0.97
  • 1:03 #7 done5 line(s)

    shot 7·sharpness 1838.8

    1. Gemma 40.95
    2. AIE1.00
    3. # Braintrust0.98
    4. WorkOS OpenAI0.96
    5. AlEngineer0.98
  • 1:14 #8 done20 line(s)

    shot 8·sharpness 2088.0

    1. Adrien Grondin@adrgrondin - Apr 40.97
    2. Google's Gemma 4 E2B running on-device on iPhone 17 Pro0.97
    3. Gemma 4 is built from the same research as Gemini 3, has image0.99
    4. understanding capabilities and can reason if needed0.98
    5. Running at ~40tk/s with MLX optimized for Apple Silicon0.98
    6. AIE1.00
    7. Meet Gemma 41.00
    8. 1.00
    9. 1.00
    10. 1.00
    11. 0.96
    12. 0:121.00
    13. 2491.00
    14. t2 5040.93
    15. 5.9K1.00
    16. h 1M0.86
    17. 0.69
    18. AlEngineer0.97
    19. EUROPE1.00
    20. AlEngineer0.97
  • 1:25 #9 done7 line(s)

    shot 9·sharpness 1697.3

    1. AIE1.00
    2. MLX1.00
    3. 1.00
    4. 1.00
    5. AlEngineer0.97
    6. EUROPE1.00
    7. AlEngineer0.97
  • 1:42 #10 done9 line(s)

    shot 10·sharpness 878.7

    1. AIE1.00
    2. M51.00
    3. 1.00
    4. 1.00
    5. 0.98
    6. 1.00
    7. AlEngineer0.97
    8. EUROPE1.00
    9. AlEngineer0.97
  • 2:04 #11 done103 line(s)

    shot 11·sharpness 2221.4

    1. Private1.00
    2. github.com0.99
    3. ④ ☑ + D0.68
    4. PlatformSolutions0.95
    5. Resources1.00
    6. Open Source1.00
    7. Enterprise0.97
    8. Pricing1.00
    9. Search or jump to...0.96
    10. Sign in1.00
    11. Sign up1.00
    12. ml-explore/mlx-swift-Im Public0.97
    13. Notifications0.97
    14. ¥ Fork 1390.94
    15. ☆ Star 3660.94
    16. <> Code0.88
    17. Issues 330.90
    18. I7 Pull requests 250.96
    19. Discussions1.00
    20. Actions0.94
    21. ① Security and quality0.98
    22. Insights1.00
    23. P main0.84
    24. P9 Branches17 Tags0.94
    25. Q Go to file0.98
    26. (> Code0.90
    27. About1.00
    28. aleroot Fix tool calling for Llama 3 (#145)0.95
    29. ec9619b · 6 hours ago0.94
    30. 248 Commits1.00
    31. Readme1.00
    32. LLMs and VLMs with MLX Swift0.99
    33. 1.00
    34. 1.00
    35. .github0.97
    36. Fix doc comments and verify n CI (e176)0.94
    37. 3 days ago0.97
    38. MIT license0.99
    39. 1.00
    40. AIE1.00
    41. 1.00
    42. Libraries0.96
    43. Fix tool calling for Llama 3 (#145)0.99
    44. 6 hours ago0.99
    45. Aa Contributing0.93
    46. Code of conduct1.00
    47. 1.00
    48. 1.00
    49. Tests1.00
    50. Fix tool callng for Llama 3 (#145)0.97
    51. 6 hours ago0.99
    52. Ar Activity0.97
    53. 1.00
    54. 1.00
    55. 1.00
    56. scripts1.00
    57. Fix doc comments and verify in CI (#176)0.99
    58. 3 days ago0.96
    59. Custom properties0.99
    60. ☆ 366 stars0.95
    61. skills0.99
    62. Decouple from tokenizer and downloader packages (#118)1.00
    63. last week1.00
    64. ⑥14 watching0.94
    65. gitignore0.94
    66. feat: implement gemma3n text model in MLXLLM (#346)1.00
    67. 9 months ago1.00
    68. ¥ 139 forks0.92
    69. .pre-commit-config.yaml1.00
    70. improve swift-format checks -- same as mlx-swift (#399)1.00
    71. 7 months ago1.00
    72. Report repository0.98
    73. .spiyml0.87
    74. prepare mlx-swift-Im0.97
    75. 5 months ago1.00
    76. Releases 60.98
    77. .swift-format0.98
    78. initial commit0.98
    79. 2 years ago0.97
    80. 2.31.3 Latest0.98
    81. last week0.95
    82. ACKNOWLEDGMENTS.md1.00
    83. Add gemma 3 embedding model (#136)0.99
    84. 3 weeks ago1.00
    85. + 5 releases0.99
    86. CODE_OF_CONDUCT.md1.00
    87. add admin files0.97
    88. 2 years ago1.00
    89. CONTRIBUTING.md1.00
    90. Add MNIST Digit Prediction/Inference (#22)0.98
    91. 2 years ago1.00
    92. Packages1.00
    93. LICENSE0.99
    94. add admin files0.98
    95. 2 years ago0.96
    96. No packages published0.98
    97. Package.swift0.98
    98. Decouple from tokenizer and downloader packages (#118)1.00
    99. last week1.00
    100. Contributors 301.00
    101. AlEngineer0.96
    102. EUROPE1.00
    103. AlEngineer0.97
  • 2:38 #12 skipped

    shot 12·duplicate of #11

  • 2:55 #13 skipped

    shot 13·duplicate of #11

  • 3:28 #14 done5 line(s)

    shot 14·sharpness 1661.6

    1. AIE1.00
    2. 1.00
    3. AlEngineer0.96
    4. EUROPE1.00
    5. AlEngineer0.96
  • 3:41 #15 done68 line(s)

    shot 15·sharpness 2446.7

    1. Private1.00
    2. huggingface.co1.00
    3. 00.96
    4. ④ ⑤ + D0.74
    5. Hugging Face1.00
    6. Q. Search models, datasets, users...0.96
    7. Models Datasets SpacesBuckets Nw0.96
    8. Docs Pricing0.99
    9. Log In0.97
    10. Sign Up0.99
    11. MLX MLX Community0.97
    12. Activity Feed0.99
    13. Rx Request to join this org0.94
    14. Al & ML interests0.96
    15. mlx-community's models 20 gemma-4-e2b0.95
    16. o0.70
    17. †↓ Sort: Recently updated0.96
    18. 1.00
    19. 1.00
    20. None defined yet.0.99
    21. 1.00
    22. AIE1.00
    23. 1.00
    24. Recent Activity1.00
    25. - mlx-comnunity/gemma-4-e2b-nvfp40.96
    26. X Any-to-Any · 2B - Updated 6 days ago · à 1.62k - 10.87
    27. X Any-to-Any 2B · Updated 6 days ago · ±9690.84
    28. - mlx-community/gemma-4-e2b-mxfp80.97
    29. 1.00
    30. 1.00
    31. 1.00
    32. 1.00
    33. 1.00
    34. acuupdateda model d0.62
    35. aculß updated a model 2 days ago0.91
    36. mlx-community/VoxCPM2-4bit0.99
    37. - mlx-community/gemma-4-e2b-mxfp40.97
    38. X Any-to-Any 2B · Updated 6 days ago · ±5420.91
    39. - mlx-community/gemma-4-e2b-it-mxfp80.98
    40. X Any-to-Any · 2B - Updated 6 days ago ± 6900.91
    41. acou pblishd a model 2 days go0.82
    42. mlx-community/VoxCPM2-Bhit0.93
    43. mlx-community/VoxCPM2-4bit0.89
    44. - mlx-comnunity/gemma-4-e2b-it-mxfp40.95
    45. X Any-to-Any · 2B · Updated 6 days ago - ±1.32k0.89
    46. - mlx-community/gemma-4-e2b-it-bf160.98
    47. XAny-to-Any · 5B - Updated 6 days ago ·±8720.83
    48. View all activity0.98
    49. -mlx-community/gemma-4-e2b-it-8bit0.98
    50. - mlx-community/gemma-4-e2b-it-6bit0.98
    51. Team members4,3051.00
    52. X Any-to-Any · 2B · Updated 6 days ago - ± 1.17k - 10.87
    53. X Any-to-Any · 2B - Updated 6 days ago · ± 5090.92
    54. - mlx-community/gemma-4-e2b-it-5bit0.96
    55. - mlx-community/gemma-4-e2b-it-4bit0.98
    56. X Any-to-Any · 18 · Updated 6 days ago · ¿5530.84
    57. X Any-to-Any 1B · Updated 6 days ago · ± 53.9k -40.84
    58. - mlx-community/genma-4-e2b-bf160.98
    59. - mlx-community/gemma-4-e2b-8bit0.99
    60. X Any-to-Any · 5B · Updated 6 days ago · ± 6600.90
    61. X Any-to-Any · 2B · Updated 6 days ago · ± 1.41k - 10.89
    62. - mlx-community/genma-4-e2b-6bit0.97
    63. - mlx-community/gemma-4-e2b-5bit0.97
    64. X Any-to-Any · i 2B · Updated & days ago - ± 3890.92
    65. X Any-to-Any - i 1B - Updated 6 days ago · ± 3660.91
    66. AlEngineer0.97
    67. EUROPE1.00
    68. AlEngineer0.98
  • 4:36 #16 done64 line(s)

    shot 16·sharpness 2455.5

    1. Private1.00
    2. huggingface.co1.00
    3. ④ ☑ + D0.69
    4. Hugging Face1.00
    5. Q. Search models, datasets, users...0.97
    6. Models Datasets Spaces Buckets w0.95
    7. Docs Pricing0.99
    8. Log In0.97
    9. Sign Up0.99
    10. MLX MLX Community0.98
    11. Activity Feed0.98
    12. Rx Request to join this org0.95
    13. Follow 10,030.90
    14. Al & ML interests0.95
    15. ( mlx-community's models 20 gemma-4-e2b0.94
    16. o0.68
    17. †↓ Sort: Recently updated0.96
    18. 1.00
    19. None defined yet.0.99
    20. AIE1.00
    21. 1.00
    22. Recent Activity1.00
    23. - mlx-comnunity/gemma-4-e2b-nvfp40.96
    24. X Any-to-Any · 2B · Updated 6 days ago · ± 1.62k · 10.87
    25. X Anyto-Any · d 2B · Updated 6 days ago · ± 9690.83
    26. - mlx-community/gemma-4-e2b-mxfp80.97
    27. 1.00
    28. 1.00
    29. 1.00
    30. 1.00
    31. acul3 updated a model 2 das ago0.83
    32. acul3 updated a model 2 days ao0.89
    33. mlx-community/VoxCPM2-4bit0.99
    34. - mlx-community/gemma-4-e2b-mxfp40.97
    35. X Any-to-Any 2B · Updated 6 days ago · ± 5420.94
    36. - mlx-community/gemma-4-e2b-it-mxfp80.98
    37. X Any-to-Any · 2B - Updated 6 days ago · ± 6900.91
    38. mix-community/VoxCPM2-Bhit0.88
    39. acull published a model 2 day0.77
    40. mlx-community/VoxCPM2-4bit0.92
    41. - mlx-community/gemma-4-e2b-it-mxfp40.95
    42. X Any-to-Any 28· Updated6 days ago · 1.32k0.84
    43. - mlx-community/gemma-4-e2b-it-bf160.98
    44. XAny-to-Any ·i 5B - Updated 6 days ago · ±8720.84
    45. View all activity0.97
    46. -mlx-community/gemma-4-e2b-it-8bit0.99
    47. - mlx-community/gemma-4-e2b-it-6bit0.98
    48. Team members 4,3050.99
    49. X Any-to-Any · 2B · Updated 6 days ago - à 1.17k - 10.86
    50. X Any-to-Any · 2B - Updated 6 days ago · ± 5090.93
    51. - mlx-community/gemma-4-e2b-it-5bit0.95
    52. - mlx-community/gemma-4-e2b-it-4bit0.98
    53. X Any-to-Any · 18 · Updated 6 days ago - à 5530.85
    54. 2 Any-to-Any - i 18 · Updated 6 days ago · ± 53.9k -40.83
    55. - mlx-community/genma-4-e2b-bf160.97
    56. - mlx-community/gemma-4-e2b-8bit0.99
    57. X Any-to-Any · 5SB · Updated 6 days ago · ≤6600.87
    58. X Any-to-Any · t 2B - Updated 6 days ago · ± 1.41k 10.90
    59. - mlx-community/genma-4-e2b-6bit0.97
    60. - mlx-community/gemma-4-e2b-5bit0.97
    61. X Any-to-Any · i 2B · Updated & days ago - ± 3890.92
    62. X Any-to-Any - i 1B - Updated 6 days ago - ± 3660.92
    63. Engineering the future of Al1.00
    64. AlEngineer0.98
  • 4:44 #17 skipped

    shot 17·duplicate of #7

  • 5:13 #18 skipped

    shot 18·duplicate of #7

  • 5:45 #19 done8 line(s)

    shot 19·sharpness 1858.8

    1. AIE1.00
    2. 1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. AlEngineer0.97
    7. EUROPE1.00
    8. AlEngineer0.97
  • 5:50 #20 skipped

    shot 20·duplicate of #9

  • 5:56 #21 done25 line(s)

    shot 21·sharpness 1521.6

    1. aiDotEngineerLondon1.00
    2. 77%~0.90
    3. 0.62
    4. 0.56
    5. Slide0.95
    6. Appearance0.99
    7. Title0.92
    8. Body1.00
    9. Silde Number0.91
    10. Background0.98
    11. Standard1.00
    12. Dynamic1.00
    13. Current Fil0.95
    14. Color Fil0.98
    15. AIE1.00
    16. 1.00
    17. 1.00
    18. 1.00
    19. 1.00
    20. 40tk/s0.93
    21. D0.62
    22. Edil Slide Layout0.87
    23. AlEngineer0.97
    24. EUROPE1.00
    25. AlEngineer0.97
  • 6:02 #22 done26 line(s)

    shot 22·sharpness 1641.3

    1. aiDotEngineerLondon1.00
    2. 77%1.00
    3. 0.84
    4. 0.62
    5. Slide0.92
    6. Appearance1.00
    7. Title1.00
    8. Body1.00
    9. Siide Number0.96
    10. 21:100.75
    11. Gema 4 (2>0.69
    12. Beckground0.96
    13. Standard1.00
    14. Dynamic0.99
    15. Meet Gemma 40.99
    16. Current Fill0.97
    17. Color Fill0.99
    18. AIE1.00
    19. 1.00
    20. 1.00
    21. 1.00
    22. 1.00
    23. Edil Slide Layout0.85
    24. AlEngineer0.96
    25. EUROPE1.00
    26. AlEngineer0.98
  • 6:27 #23 done30 line(s)

    shot 23·sharpness 2200.2

    1. 21:110.85
    2. and at nighit)0.85
    3. Gemma 4 E2B0.96
    4. Louvre Museums (Book tickets in advancel0.92
    5. Focus on key wings like the Mona Lisa and0.93
    6. Greek antiqulties.)0.97
    7. Notre Dame Cathedral: Vew the exterior0.95
    8. and the ongoing restoration.)0.95
    9. • Are de Triomphe & Champs-Élysées: For0.94
    10. stunning views and grand scale.0.94
    11. Palace of Versalles (Day Trip): (Requres a0.92
    12. ful day, but woth the trip for the ryal0.88
    13. AIE1.00
    14. Musde drOrsay: (Home to incredible0.92
    15. grandeur.)1.00
    16. 1.00
    17. 1.00
    18. 1.00
    19. Montmartre & Sacrd-Coeur: (For0.91
    20. bohemian charm and breathtaking0.98
    21. Impressonist and Post-mpreslonist art0.88
    22. panoramic city vlews.)0.93
    23. 7-Day Itinerary Breakdown0.96
    24. This ltnerary blances maor sightseeing days0.89
    25. with deeper neighborhood exploration.0.98
    26. Day 1: Arrival & Iconic Introduction0.95
    27. (Eiffel Tower Focus0.96
    28. Ask anything1.00
    29. Engineering the future of Al1.00
    30. AlEngineer0.97

Transcript

115 cues· 1,630 words· 8,221 chars

  1. 0:15 okay hello everybody i'm going to show you today how to run gma4 on iphone with mlx so first let's introduce myself i'm adrian you can find my twitter if you want to learn more about all on device things i'm the developer of
  2. 0:32 Locally AI, so maybe you have already seen the app.
  3. 0:36 So Locally AI is a chatbot that allow you to run on-device models on your iPhone with MLX.
  4. 0:42 So I will just go through what is MLX in a few seconds.
  5. 0:48 Basically, as I said, it's a chatbot.
  6. 0:51 It's fully native.
  7. 0:51 You can also chat with Apple Foundation with it and many models that are compatible with MLX.
  8. 0:59 And one of these models is GMA4.
  9. 1:02 So basically, GMA4 by Gold in Mine have a lot of models, and some of them can run on iPhone, like the smaller ones, and they are pretty great.
  10. 1:13 Maybe you have seen on Twitter one of the posts I've made that I demo it running in the app on iPhone.
  11. 1:19 It's really fast.
  12. 1:20 It runs really well on MLX and behind the app.
  13. 1:23 So it's using MLX.
  14. 1:25 MLX is a framework made by Apple that is optimized for Apple Silicon.
  15. 1:30 So mainly the chip in iPhone, but also the chips on the Mac.
  16. 1:37 The app locally is available on iPad.
  17. 1:39 It works also very well on this device and Mac OS.
  18. 1:42 and everything is built to be as optimized as possible on these devices.
  19. 1:49 So if you want to run an iPhone language model, so Gemma works well, but you have a lot of model that you can run also.
  20. 1:57 The QWEN model and the small LM model from the Game Face.
  21. 2:02 The place you will want to go is GitHub and go to the repo MLX Swift LM.
  22. 2:07 I won't go into detail how you implement the repo.
  23. 2:10 I think I will let your agent implement that for you, but it's one repo that you need to install if you're developing an iOS or macOS or iPadOS app.
  24. 2:19 And you can use that to simply download the model and then run it.
  25. 2:24 The API is very straightforward, very simple to implement.
  26. 2:27 In less than 10 minutes, you can have an iOS app with a model that is running on your device.
  27. 2:32 That's very simple to do, as mentioned.
  28. 2:35 MLX.
  29. 2:36 So this is MLX Swift LM.
  30. 2:37 But if you're more into Python apps or macOS apps, you can also run MLX VLM.
  31. 2:42 from Prince, maybe you have seen him, like he's doing on device for audio with MLX audio and visual model with MLX VLM and also MLX video to run image generation model or video generation model.
  32. 2:59 MLX, the ecosystem is getting bigger, it's getting bigger, it's really great right now.
  33. 3:06 You can do pretty much everything, like omni-models, or as I said, text-to-speech, speech-to-speech.
  34. 3:12 There's a lot of things that you can do with a model.
  35. 3:16 And basically, let's say you integrate this model.
  36. 3:20 You can't just run any model on it.
  37. 3:23 You need to get some model, and there's a good place to get the model.
  38. 3:27 It's a GingFace.
  39. 3:28 I'm pretty sure everybody heard of a GingFace.
  40. 3:31 And on a GingFace, you will want to look for MLX community.
  41. 3:34 This is where all the weight of a model, the quantized weight, the full size, will be uploaded.
  42. 3:41 So you will just be able to go to this community, look for the models.
  43. 3:45 I think right now there's almost 4,000 or 5,000 models uploaded.
  44. 3:49 So the community is really active on it.
  45. 3:52 When the model is released by your lab, you will directly have it almost 30 minutes after release quantized in the 4-bit, 6-bit.
  46. 4:00 and everything that you can you can imagine here you have an example for jima for each ub which is the one i i run on on iphone there's like a lot of variant of it for from bf16 and make mxfp4 like 5-bit 6-bit there's everything so so you download mlx with tlm install it with your agent or anything then you go to mlx community and you just choose the model you want to run and
  47. 4:28 then with the ID you can just pass it to the framework and it will directly, it will be integrated with the GingFace to download the model with MLX Swift LM directly.
  48. 4:36 So you just need to grab the ID and pass it to the framework.
  49. 4:41 Usually when you're running the model on iPhone, what you want to do is selecting some quantized version of the model because the full size will be way too large.
  50. 4:54 What I recommend is trying, depending on the size, between 3-bit and 8-bit.

Open at this second