read-only demo

Videos OV56RddyFuU

Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning — Merve Noyan, Hugging Face

index_state ready data_status ok

AI Engineer· published 2026-05-13· 0:19:10· en-US· indexed 2026-08-10 19:52

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:23, 1 of 1 keyframes kept
  5. Shot 4, 0:23 to 0:29, 1 of 1 keyframes kept
  6. Shot 5, 0:29 to 0:47, 1 of 1 keyframes kept
  7. Shot 6, 0:47 to 0:56, 1 of 1 keyframes kept
  8. Shot 7, 0:56 to 1:23, 1 of 1 keyframes kept
  9. Shot 8, 1:23 to 1:58, 1 of 1 keyframes kept
  10. Shot 9, 1:58 to 2:00, 1 of 1 keyframes kept
  11. Shot 10, 2:00 to 2:33, 1 of 1 keyframes kept
  12. Shot 11, 2:33 to 2:47, 1 of 1 keyframes kept
  13. Shot 12, 2:47 to 3:02, 1 of 1 keyframes kept
  14. Shot 13, 3:02 to 3:11, 1 of 1 keyframes kept
  15. Shot 14, 3:11 to 3:46, 1 of 1 keyframes kept
  16. Shot 15, 3:46 to 4:20, 0 of 1 keyframes kept
  17. Shot 16, 4:20 to 4:48, 1 of 1 keyframes kept
  18. Shot 17, 4:48 to 5:13, 1 of 1 keyframes kept
  19. Shot 18, 5:13 to 5:32, 1 of 1 keyframes kept
  20. Shot 19, 5:32 to 5:48, 1 of 1 keyframes kept
  21. Shot 20, 5:48 to 6:17, 1 of 1 keyframes kept
  22. Shot 21, 6:17 to 6:47, 0 of 1 keyframes kept
  23. Shot 22, 6:47 to 6:56, 1 of 1 keyframes kept
  24. Shot 23, 6:56 to 7:36, 1 of 1 keyframes kept
  25. Shot 24, 7:36 to 8:09, 1 of 1 keyframes kept
  26. Shot 25, 8:09 to 8:56, 1 of 1 keyframes kept
  27. Shot 26, 8:56 to 9:16, 1 of 1 keyframes kept
  28. Shot 27, 9:16 to 9:32, 1 of 1 keyframes kept
  29. Shot 28, 9:32 to 9:46, 1 of 1 keyframes kept
  30. Shot 29, 9:46 to 10:09, 1 of 1 keyframes kept
  31. Shot 30, 10:09 to 10:13, 0 of 1 keyframes kept
  32. Shot 31, 10:13 to 10:21, 1 of 1 keyframes kept
  33. Shot 32, 10:21 to 10:51, 1 of 1 keyframes kept
  34. Shot 33, 10:51 to 10:55, 1 of 1 keyframes kept
  35. Shot 34, 10:55 to 11:33, 1 of 1 keyframes kept
  36. Shot 35, 11:33 to 11:48, 1 of 1 keyframes kept
  37. Shot 36, 11:48 to 12:02, 1 of 1 keyframes kept
  38. Shot 37, 12:02 to 12:09, 1 of 1 keyframes kept
  39. Shot 38, 12:09 to 12:41, 1 of 1 keyframes kept
  40. Shot 39, 12:41 to 12:45, 1 of 1 keyframes kept
  41. Shot 40, 12:45 to 12:53, 1 of 1 keyframes kept
  42. Shot 41, 12:53 to 12:56, 1 of 1 keyframes kept
  43. Shot 42, 12:56 to 13:38, 1 of 1 keyframes kept
  44. Shot 43, 13:38 to 14:16, 1 of 1 keyframes kept
  45. Shot 44, 14:16 to 14:26, 1 of 1 keyframes kept
  46. Shot 45, 14:26 to 14:36, 1 of 1 keyframes kept
  47. Shot 46, 14:36 to 14:59, 1 of 1 keyframes kept
  48. Shot 47, 14:59 to 15:32, 1 of 1 keyframes kept
  49. Shot 48, 15:32 to 15:37, 1 of 1 keyframes kept
  50. Shot 49, 15:37 to 15:43, 1 of 1 keyframes kept
  51. Shot 50, 15:43 to 16:03, 1 of 1 keyframes kept
  52. Shot 51, 16:03 to 16:05, 1 of 1 keyframes kept
  53. Shot 52, 16:05 to 16:10, 1 of 1 keyframes kept
  54. Shot 53, 16:10 to 16:20, 1 of 1 keyframes kept
  55. Shot 54, 16:20 to 16:28, 1 of 1 keyframes kept
  56. Shot 55, 16:28 to 16:36, 1 of 1 keyframes kept
  57. Shot 56, 16:36 to 17:04, 1 of 1 keyframes kept
  58. Shot 57, 17:04 to 17:30, 1 of 1 keyframes kept
  59. Shot 58, 17:30 to 17:37, 1 of 1 keyframes kept
  60. Shot 59, 17:37 to 17:44, 1 of 1 keyframes kept
  61. Shot 60, 17:44 to 17:50, 1 of 1 keyframes kept
  62. Shot 61, 17:50 to 17:52, 0 of 1 keyframes kept
  63. Shot 62, 17:52 to 17:55, 0 of 1 keyframes kept
  64. Shot 63, 17:55 to 17:56, 1 of 1 keyframes kept
  65. Shot 64, 17:56 to 18:05, 0 of 1 keyframes kept
  66. Shot 65, 18:05 to 18:15, 1 of 1 keyframes kept
  67. Shot 66, 18:15 to 18:46, 1 of 1 keyframes kept
  68. Shot 67, 18:46 to 18:53, 1 of 1 keyframes kept
  69. Shot 68, 18:53 to 18:56, 1 of 1 keyframes kept
  70. Shot 69, 18:56 to 19:09, 1 of 1 keyframes kept
  71. Shot 70, 19:09 to 19:10, 0 of 1 keyframes kept

71 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
201
whisperx 201
chunks
33
from 201 cues
keyframes
64
kept of 71 captured
frames with text
64
2,021 lines read
chapters
15
from the source metadata
keyframe bytes
8.2 MB
word timings on 201 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 16:56 2m 03s
stt done 2026-08-10 16:58 24s
chunk done 2026-08-10 16:58 0s
text_embed done 2026-08-10 19:52 1s
keyframe done 2026-08-10 16:59 1m 48s
ocr done 2026-08-10 17:00 38s
frame_embed done 2026-08-10 19:52 10s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:21 #3 done15 line(s)

    shot 3·sharpness 1494.4

    1. eer0.96
    2. raintrus1.00
    3. ze1.00
    4. jineer0.97
    5. Open Agent1.00
    6. OPE0.98
    7. eer1.00
    8. gger.de1.00
    9. Ecosystem1.00
    10. odal0.99
    11. Engineer1.00
    12. (and having an Intern Al Engineer at your fingertips)0.99
    13. EUROPE1.00
    14. AlEngineer0.98
    15. EUROPE1.00
  • 0:24 #4 done12 line(s)

    shot 4·sharpness 771.5

    1. eer0.97
    2. aintrus1.00
    3. ze1.00
    4. ineer1.00
    5. Open/source1.00
    6. eer1.00
    7. ger.de1.00
    8. odal0.97
    9. ngineer1.00
    10. EUROPE1.00
    11. AlEngineer0.98
    12. EUROPE1.00
  • 0:36 #5 done13 line(s)

    shot 5·sharpness 770.0

    1. er1.00
    2. raintrus1.00
    3. ze1.00
    4. jineer0.92
    5. PE0.99
    6. Open/source1.00
    7. eer0.96
    8. 1ger.de0.99
    9. dal1.00
    10. jineer1.00
    11. OPE1.00
    12. AlEngineer0.99
    13. EUROPE1.00
  • 0:54 #6 done36 line(s)

    shot 6·sharpness 1277.7

    1. B1.00
    2. trust1.00
    3. Open-source1.00
    4. deepseek-ai/DeepSeek-V3.1-Terminuslike291Follow DeepSeek97.8k0.98
    5. Text GenerationTransformersSafetensors deepseek_v30.99
    6. conversational0.98
    7. custom_code1.00
    8. text-generation-inference1.00
    9. License: mit0.98
    10. Model card1.00
    11. Files and versions o xet0.91
    12. Community1.00
    13. ¿ Edit model card0.97
    14. dev1.00
    15. DeepSeek-V3.1-Terminus1.00
    16. Downloads last mont0.99
    17. 11,0701.00
    18. deepseek1.00
    19. Safetensors0.97
    20. Model size 685B params Tensortype BF16·F8_E4M3 · F320.94
    21. 0 Chat template Files info0.93
    22. er1.00
    23. Inference Providers1.00
    24. Novita1.00
    25. Hugging Fac0.98
    26. X Twitter0.93
    27. deepseek al0.97
    28. Text Generation1.00
    29. Examples1.00
    30. Input a message to start chatting with deepseek-ai/DeepSeek-V3.1-Terminus.1.00
    31. Introduction1.00
    32. This update maintains the model's original capabilities while addressing issues1.00
    33. reported by users, including:1.00
    34. I anmuaae rnncictanru- Darlurina inetanree nf mivad Chinaca.Fnalich tovt and0.89
    35. AlEngineer0.97
    36. EUROPE1.00
  • 1:10 #7 done43 line(s)

    shot 7·sharpness 1485.0

    1. Open-source1.00
    2. deepseek-ai/DeepSeek-V3.1-Terminuslike2911.00
    3. Follow DeepSeek 97.8k0.98
    4. Text Generation1.00
    5. Transformers1.00
    6. Safetensors1.00
    7. deepseek_v31.00
    8. conversational1.00
    9. custom_code1.00
    10. text-generation-inference1.00
    11. License: mit1.00
    12. Model card1.00
    13. Files and versionsxet1.00
    14. Community 100.99
    15. Edit model card0.98
    16. 0.69
    17. DeepSeek-V3.1-Terminus1.00
    18. Downloads last month1.00
    19. 11,0701.00
    20. deepseek1.00
    21. Safetensors①0.96
    22. Model size 6858 params0.98
    23. Tensor type BF16 · F8_E4M3 · F320.96
    24. φ Chat template0.95
    25. Files info0.99
    26. Inference ProvidersNEw0.98
    27. Novita1.00
    28. DeepSeek Homepage1.00
    29. Chat DeepSeek V31.00
    30. Hugging Face1.00
    31. DeepSeek AI0.97
    32. Text Generation0.98
    33. Examples1.00
    34. Discord DeepSeek AI1.00
    35. WeChat DeepSeek AI0.96
    36. X Twitter0.96
    37. deepseek ai0.96
    38. License MIT0.99
    39. Input a message to start chatting with deepseek-ai/DeepSeek-V3.1-Terminus.1.00
    40. Introduction1.00
    41. This update maintains the model's original capabilities while addressing issues1.00
    42. reported by users, including:1.00
    43. I anauaae roncictancu Deducina inctancec nf mived Chinece-Fnalich tovt and0.89
  • 1:38 #8 done11 line(s)

    shot 8·sharpness 2987.7

    1. onAI0.83
    2. Why does open-source matter?1.00
    3. • Absolute control over models1.00
    4. • Cost reduction in multipliers1.00
    5. • Further customize, shrink models depending on your needs0.98
    6. • Guaranteed privacy for the end-user, enable on-0.99
    7. device/in-browser uses, data doesn't go to another server0.99
    8. NTRY1.00
    9. lind0.83
    10. AlEngineer0.98
    11. EUROPE1.00
  • 1:59 #9 done8 line(s)

    shot 9·sharpness 2896.1

    1. Why does open-source matter?1.00
    2. • Absolute control over models1.00
    3. • Cost reduction in multipliers1.00
    4. • Further customize, shrink models depending on your needs0.99
    5. • Guaranteed privacy for the end-user, enable on-0.99
    6. device/in-browser uses, data doesn't go to another server1.00
    7. AlEngineer0.98
    8. EUROPE1.00
  • 2:23 #10 done94 line(s)

    shot 10·sharpness 2730.2

    1. Artificial Analysis Index0.98
    2. Artificial Analysis Intelligence Index by Open Weights / Proprietary1.00
    3. 28 of 472 models1.00
    4. 00.67
    5. Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, τ2-Bench Telecom, Terminal-0.99
    6. + Add model from specific provider0.97
    7. Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt1.00
    8. Proprietary1.00
    9. Open Weights0.96
    10. A Artificial Analysis0.99
    11. 571.00
    12. 571.00
    13. 531.00
    14. 521.00
    15. 521.00
    16. 511.00
    17. 501.00
    18. 501.00
    19. 491.00
    20. 491.00
    21. 481.00
    22. 471.00
    23. 461.00
    24. 451.00
    25. 421.00
    26. 391.00
    27. 371.00
    28. 361.00
    29. 361.00
    30. 341.00
    31. 331.00
    32. 321.00
    33. 271.00
    34. 261.00
    35. 241.00
    36. 241.00
    37. 181.00
    38. G1.00
    39. AI0.91
    40. 80.64
    41. AI0.93
    42. Z0.87
    43. a0.51
    44. K0.93
    45. G1.00
    46. a0.67
    47. G1.00
    48. AI0.93
    49. G1.00
    50. M0.79
    51. 80.83
    52. C0.93
    53. C0.94
    54. C0.95
    55. C0.71
    56. Ca0.59
    57. C0.94
    58. C0.86
    59. C0.93
    60. C0.97
    61. C.0.61
    62. C0.95
    63. C0.96
    64. C0.97
    65. C0.88
    66. Cn0.58
    67. C0.95
    68. (max)0.84
    69. Gemini 3.11.00
    70. Pro Preview0.99
    71. GPT-5.40.99
    72. Claude Opus0.96
    73. 4.6 (max)0.98
    74. Claude1.00
    75. GLM-5.190.99
    76. MiniMax-M2.7 ?0.93
    77. Grok 4.200.97
    78. 0309 v20.84
    79. GPT-5.4 mini0.96
    80. (xhigh)0.93
    81. Kimi K2.5 90.92
    82. Gemini 30.98
    83. Flash0.99
    84. Qwen3.5 397B0.98
    85. Haiku1.00
    86. Nova 2.0 Pro1.00
    87. Preview1.00
    88. (medium)0.97
    89. Preview1.00
    90. gpt-oss-120B1.00
    91. K-EXAONE ?0.91
    92. gpt-oss-20B1.00
    93. K2 Think V2 90.89
    94. Llama 4 Maverick1.00
  • 2:35 #11 done57 line(s)

    shot 11·sharpness 2802.0

    1. Hugging Face Hub1.00
    2. Home for the open-source machine learning community: share & discover models,0.99
    3. datasets, apps, connect with the community and more!1.00
    4. Hugging Face1.00
    5. O Search models, datasets, users...0.96
    6. Models1.00
    7. Datasets1.00
    8. Spaces1.00
    9. Community0.97
    10. Docs Pricing0.99
    11. i0.56
    12. + New0.98
    13. Following3551.00
    14. New Post0.99
    15. Trending1.00
    16. last 7 days0.99
    17. Models Datasets0.98
    18. Spaces1.00
    19. Papers Collections Community1.00
    20. Posts1.00
    21. Models1.00
    22. Datasets1.00
    23. Spaces1.00
    24. merve1.00
    25. Upvotes Likes Articles0.99
    26. Profile1.00
    27. moonshotai/Kimi-K2-Instruct1.00
    28. Inbox (7113)1.00
    29. hysts updated a Space 2 minutes ago0.98
    30. Text GenerationUpdated...25.1k1.23k0.99
    31. Settings1.00
    32. $ Billing0.99
    33. SeqTex0.94
    34. DeepSite v20.99
    35. 10.2k1.00
    36. SeqTex generates texture based on textual conditions1.00
    37. Generate any application with DeepSeek0.99
    38. Organizations1.00
    39. Hugging Face0.99
    40. Wauplin updated a Space5 minutes ago1.00
    41. HuggingFaceTB/SmolLM3-3B1.00
    42. Text Generation3BUp...59.5k4850.97
    43. G Google0.98
    44. Responses.js1.00
    45. 090.83
    46. Deprem Yapay Zeka1.00
    47. Check out https://github.com/huggingface/responses.js0.98
    48. mistralai/Devstral-Small-25071.00
    49. Notebooks-explorers1.00
    50. Text Generation 24BU...9.58k2330.95
    51. - SODA0.89
    52. — PyTorch Image Models0.95
    53. Weyaxi updated a dataset 5 minutes ago0.98
    54. black-forest-labs/FLUX.1-Kontext-dev0.99
    55. Image-to-Image·Updated ...274k1.68k0.95
    56. Templates1.00
    57. Weyaxi/huggingface-leaderboard1.00
  • 2:50 #12 done48 line(s)

    shot 12·sharpness 1283.1

    1. Hugging Face Hub1.00
    2. Home for the machine learning community: share & discover1.00
    3. mo1.00
    4. 2.7M models of all libraries and tasks1.00
    5. more!1.00
    6. 925k+ datasets1.00
    7. + Ne0.94
    8. Spaces (apps)1.00
    9. New Post1.00
    10. Trending1.00
    11. last 7 days1.00
    12. Collections1.00
    13. Community1.00
    14. Posts1.00
    15. AIl0.53
    16. Models1.00
    17. Datasets1.00
    18. Spaces1.00
    19. merve0.98
    20. Prof1.00
    21. ecosystem: open-source libraries1.00
    22. cruct0.98
    23. Inbo0.98
    24. 25.1k1.00
    25. Setti0.99
    26. $ Billir0.92
    27. (transformers, diffusers & more!)1.00
    28. 10.2k1.00
    29. ith DeepSeek1.00
    30. Organiza1.00
    31. HuggingFaceTB/SmolLM3-3B1.00
    32. Hugging Face1.00
    33. Wauplin updated a Space5 minutes ago0.97
    34. Text Generation 3B·Up...59.5k0.96
    35. 4851.00
    36. G Google1.00
    37. Responses.js1.00
    38. Deprem Yapay Zeka0.97
    39. Check out https://github.com/huggingface/responses.js0.99
    40. mistralai/Devstral-Small-25071.00
    41. Notebooks-explorers1.00
    42. - SODA0.88
    43. —PyTorch Image Models0.97
    44. Weyaxi updated a dataset 5 minutes ago0.98
    45. black-forest-labs/FLUX.1-Kontext-dev0.99
    46. Templates1.00
    47. Weyaxi/huggingface-leaderboard1.00
    48. Vorse0.70
  • 3:07 #13 done53 line(s)

    shot 13·sharpness 1126.2

    1. /models1.00
    2. Main1.00
    3. Tasks1.00
    4. Libraries1.00
    5. Languages Licenses0.97
    6. Models2,665,3241.00
    7. Filter by name0.98
    8. Full-text search0.99
    9. Inference Available0.99
    10. ↑↓ Sort: Trending1.00
    11. Other1.00
    12. Tasks1.00
    13. Image-Text-to-Text·: 36B·Updated about 3 hours ago·259k·6040.94
    14. Qwen/Qwen3.5-35B-A3B1.00
    15. Text Generation0.99
    16. Any-to-Any0.96
    17. P0.61
    18. Image-Text-to-Text1.00
    19. Image-to-Text1.00
    20. Qwen/Qwen3.5-27B1.00
    21. Image-Text-to-Text28B·Updated 2 days ago108k3960.98
    22. 0.72
    23. Image-to-Image1.00
    24. Text-to-Image0.98
    25. Text-to-Video1.00
    26. Text-to-Speech1.00
    27. +441.00
    28. Qwen/Qwen3.5-397B-A17B1.00
    29. Image-Text-to-Text 403B·Updated 4 days ago·726k1.11k0.96
    30. Parameters1.00
    31. <1B0.99
    32. 6B0.96
    33. 12B0.99
    34. 32B1.00
    35. 128B0.90
    36. >500B0.99
    37. Qwen/Qwen3.5-122B-A10B1.00
    38. Image-Text-to-Text · 125B · Updated 3 days ago·108k· 3240.91
    39. Libraries1.00
    40. unsloth/Qwen3.5-35B-A3B-GGUF1.00
    41. PyTorch0.97
    42. TensorFlow1.00
    43. X JAX0.82
    44. Image-Text-to-Text 35B·Updated 3 days ago·265k·2750.94
    45. Transformers1.00
    46. Diffusers1.00
    47. zai-org/GLM-51.00
    48. sentence-transformers1.00
    49. Safetensors1.00
    50. Text Generation754B·Updated 14 days ago189k1.63k0.96
    51. ONNX1.00
    52. GGUF1.00
    53. Transformers.js0.98
  • 3:25 #14 done8 line(s)

    shot 14·sharpness 3036.3

    1. models → agents, serve locally1.00
    2. Agentic LLMs (thinking + tool calling): gpt-oss, Gemma-4,0.99
    3. Minimax M2.7, GLM-5, Nemotron3-Super0.99
    4. Agentic vision models (thinking + CUA): Qwen3.50.99
    5. (Alibaba), Kimi-K2.50.98
    6. mlx_lm.generate --prompt "How tall is Mt Everest?"0.99
    7. vllm serve Qwen/Qwen3-8B # then query with OpenAI Completion1.00
    8. llama-server -m model.gguf --port 80801.00
  • 3:53 #15 skipped

    shot 15·duplicate of #14

  • 4:34 #16 done67 line(s)

    shot 16·sharpness 1361.1

    1. compare open models1.00
    2. Main1.00
    3. Tasks1.00
    4. Libraries1.00
    5. Languages Licenses Other1.00
    6. Datasets 171.00
    7. Filter by name1.00
    8. Full-text search1.00
    9. ↑↓ Sort: Trending0.98
    10. Modalities0.99
    11. openai/gsm8k1.00
    12. allenai/olmOCR-bench1.00
    13. 3D1.00
    14. Audio1.00
    15. Document1.00
    16. Geospatial1.00
    17. Benchmark·Updated 17 days ago·17.6k·±765k1.24k0.96
    18. Benchmark · Updated Feb 19·± 4.14k · 1760.90
    19. Image1.00
    20. Tabular1.00
    21. Text1.00
    22. Time-series1.00
    23. ScaleAI/SWE-bench_Pro1.00
    24. cais/hle1.00
    25. Video1.00
    26. Benchmark·Updated Feb 23·731·±692k·780.91
    27. Benchmark · Updated Jan 20· 2.5k·±46.1k ·7610.91
    28. Size (rows)1.00
    29. SWE-bench/SWE-bench_Verified1.00
    30. collinear-ai/yc-bench1.00
    31. <1K0.93
    32. >1T0.95
    33. Benchmark·Updated Feb 27· 500·126k·280.92
    34. Benchmark - Updated 17 days ago · ± 104 · 150.91
    35. TIGER-Lab/MMLU-Pro1.00
    36. MathArena/aime_20261.00
    37. Format0.97
    38. Benchmark · Updated 29 days ago·12.1k ·± 113k·4640.92
    39. Benchmark · Updated Feb 16 · 30± 12.7k · 280.90
    40. 4 json0.87
    41. EE CSV0.78
    42. parquet0.94
    43. optimized-parquet1.00
    44. imagefolder1.00
    45. soundfolder1.00
    46. webdataset1.00
    47. harborframework/terminal-bench-2.01.00
    48. Benchmark · Updated Feb 17 · ± 2.98k - 190.89
    49. Idavidrein/gpqa1.00
    50. Benchmark · Updated Mar 5· 1.25k ·± 105k · 4080.87
    51. text0.94
    52. arrow0.98
    53. mteb/arguana1.00
    54. FutureMa/EvasionBench1.00
    55. Type1.00
    56. Benchmark·Updated Feb 22·11.5k-± 12.8k·50.91
    57. Benchmark · Updated Feb 19·16.7k·± 177·850.92
    58. Benchmark×1.00
    59. è Traces0.94
    60. mteb/BRIGHT1.00
    61. hf-audio/open-asr-leaderboard1.00
    62. Benchmark · Updated 7 days ago·1.35M - ±662-20.90
    63. Benchmark - Updated 6 days ago · 99.5k · ± 20.1k - 70.90
    64. likaixin/ScreenSpot-Pro1.00
    65. nvidia/compute-eval1.00
    66. Benchmark - Updated 22 days ago ·± 7.73k - 600.91
    67. Benchmark - Updated 20 days ago · 2.46k - ± 5.63k · 220.89
  • 4:53 #17 done45 line(s)

    shot 17·sharpness 1624.0

    1. compare open models1.00
    2. Datasets:ScaleAl/SWE-bench_Pro1.00
    3. like1.00
    4. 781.00
    5. FollowScale Al0.98
    6. 2841.00
    7. Benchmark1.00
    8. Modalities:1.00
    9. Text1.00
    10. Formats:1.00
    11. parquet0.95
    12. Size:1.00
    13. <1K1.00
    14. Libraries:1.00
    15. Datasets1.00
    16. pandas1.00
    17. Polars1.00
    18. +11.00
    19. Dataset card1.00
    20. 田Data Studio1.00
    21. Files and versions xet0.96
    22. Community1.00
    23. Leaderboard Official Benchmark0.97
    24. ① Learn more0.98
    25. Experimental1.00
    26. Task: SWE Bench Pro0.97
    27. #1.00
    28. MODEL1.00
    29. SCORE1.00
    30. zai-org/GLM-5.10.97
    31. 58.4*1.00
    32. 21.00
    33. MiniMaxAI/MiniMax-M2.51.00
    34. 55.41.00
    35. 31.00
    36. moonshotai/Kimi-K2.51.00
    37. 50.71.00
    38. Qwen/Qwen3-Coder-Next1.00
    39. source1.00
    40. 44.3*1.00
    41. 51.00
    42. Qwen/Qwen3-Coder-480B-A35B-Instruct0.99
    43. source1.00
    44. 38.71.00
    45. Show all 14 models0.99
  • 5:22 #18 done35 line(s)

    shot 18·sharpness 1258.2

    1. Inference Providers for Qwen/Qwen3-VL-235B-A22B-Thinking0.99
    2. ×0.80
    3. Novita1.00
    4. ★ Auto0.89
    5. mervenoyan1.00
    6. t routing0.98
    7. Python1.00
    8. JavaScript1.00
    9. CcURL0.95
    10. huggingface_hub0.98
    11. requests openai1.00
    12. Stream1.00
    13. import os0.98
    14. Copy1.00
    15. from huggingface_hub import InferenceClient1.00
    16. client = InferenceClient(0.99
    17. vibe-check1.00
    18. provider="novita",1.00
    19. api_key=os.environ["HF_TOKEN"],0.99
    20. completion = client.chat.completions.create(1.00
    21. model="Qwen/Qwen3-VL-235B-A22B-Thinking",1.00
    22. go serverless1.00
    23. messages=[1.00
    24. {0.85
    25. "role": "user",0.99
    26. "content": [0.96
    27. {0.95
    28. "type": "text",0.98
    29. "text": "Describe this image in one sentence."1.00
    30. 3,0.64
    31. {0.92
    32. "type": "image_url",0.99
    33. "image_url": {0.99
    34. "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Islan1.00
    35. print(completion.choices[0].message)1.00
  • 5:35 #19 done48 line(s)

    shot 19·sharpness 2413.3

    1. compare providers across models0.99
    2. Inference Providers ·Metrics for top trending models0.98
    3. Browse all0.96
    4. Filter by model or provider...1.00
    5. Model0.98
    6. Provider1.00
    7. Input $/1M0.98
    8. Output $/1M0.99
    9. Context1.00
    10. Latency(s)1.00
    11. Throughput(t/s)0.99
    12. Ggoogle/gemma-4-31B-it0.99
    13. novita1.00
    14. $0.141.00
    15. $0.401.00
    16. 262,1441.00
    17. 1.301.00
    18. 301.00
    19. G google/gemma-4-26B-A4B-it0.98
    20. novita1.00
    21. $0.131.00
    22. $0.401.00
    23. 262,1441.00
    24. 2.961.00
    25. 131.00
    26. Qwen/Qwen3.5-9B0.97
    27. 370.87
    28. together1.00
    29. $0.101.00
    30. $0.151.00
    31. 262,1441.00
    32. 0.791.00
    33. 621.00
    34. zai-org/GLM-51.00
    35. 70.73
    36. novita cheapest0.97
    37. $1.001.00
    38. $3.201.00
    39. 202,8001.00
    40. 2.411.00
    41. 340.97
    42. zai-org/GLM-51.00
    43. 70.68
    44. together1.00
    45. $1.001.00
    46. $3.201.00
    47. 202,7521.00
    48. 0.611.00
  • 6:14 #20 done

    shot 20·sharpness 2969.2

  • 6:21 #21 skipped

    shot 21·duplicate of #20

  • 6:55 #22 done

    shot 22·sharpness 948.8

  • 7:20 #23 done

    shot 23·sharpness 2543.9

The page's on-screen-text budget of 600 lines is spent, so the last cards in this grid list fewer lines than they hold. Narrow the page with ?frames= to read them.

Transcript

201 cues· 2,780 words· 14,580 chars

  1. 0:15 Hello everyone and welcome to this talk in OpenAgent ecosystem and I would like to call it having an AI engineer at your fingertips.
  2. 0:25 I'm Merve and I work in the open source team of HuggingFace.
  3. 0:28 How many of you are using HuggingFace on daily basis?
  4. 0:33 Oh, let's change that.
  5. 0:35 This is not OK.
  6. 0:38 But first, let's talk a bit about open source and what it is.
  7. 0:40 So when it comes to machine learning, open source is absolutely differential.
  8. 0:45 Basically, you have the open weight models that go in with non-commercial licenses.
  9. 0:52 We call them open weight.
  10. 0:53 And then we have open source models that have
  11. 0:56 commercially available licenses, such as this one from DeepSeek.
  12. 1:00 It's called the MIT license or Apache 2.0.
  13. 1:03 And then there is even more open models that have the code open.
  14. 1:09 If you have agents, the harness is open.
  15. 1:12 Everything is open.
  16. 1:13 And this matters even more by the fact that yesterday or the other day, it was revealed that the cloud performance was going down.
  17. 1:24 So if you have everything in the open, nothing changes without you knowing no performance degradation, without you knowing everything's great.
  18. 1:34 But on top of it, if you have access to the weights, you can shrink them.
  19. 1:39 You can quantize them.
  20. 1:41 You can fine tune them if you feel like it.
  21. 1:44 And it's absolute guaranteed privacy for your end user because you can deploy it to edge devices, browsers without the data going somewhere else.
  22. 1:54 This matters a lot in my opinion, even more these days with the security breaches and everything.
  23. 2:02 And there was this argument, maybe a few years ago, that open source models aren't as good as closed models.
  24. 2:08 No, this is not the case.
  25. 2:09 Like you see, for instance, the latest GLM 5.1 is absolutely crushing it.
  26. 2:14 And I'm actually using it in my coding setup.
  27. 2:18 This is the artificial analysis intelligence index.
  28. 2:22 And the green ones are open models.
  29. 2:24 Meanwhile, the black ones are the closed models.
  30. 2:28 And we just catched up.
  31. 2:30 And we will catch up even more with the upcoming models and stuff.
  32. 2:35 And let's go back to Hugging Face Hub.
  33. 2:37 So everything is facilitated through Hugging Face Hub, all of the open releases.
  34. 2:43 It's the infra layer for all of your open source workflows.
  35. 2:48 And as of now, it's hosting even more models.
  36. 2:50 I should have updated the number.
  37. 2:52 It's probably close to 3 million.
  38. 2:54 A lot of data sets, spaces, and everything.
  39. 2:57 But that's not all when it comes to the iGenetic ecosystem.
  40. 3:00 And this is what we are going to talk about today.
  41. 3:03 So when you go to the models, you can filter for agentic models.
  42. 3:09 They are mostly the trending ones.
  43. 3:12 And there is two types of models, in my opinion.
  44. 3:15 There is the vision LMs, and then there is the LLMs.
  45. 3:19 And the vision LMs can also act as a computer use agent over the screenshots.
  46. 3:24 They know where to click, et cetera, which is pretty cool.
  47. 3:27 And one trend I have recently noticed is the fact that
  48. 3:32 you have labs releasing their LLMs with vision capabilities, day zero.
  49. 3:39 Like, for instance, the GEMA-4 was an omni model, and still it's an agentic model.
  50. 3:45 There is Q1 3.5.

Chapters

  1. 0:00 Introduction to Open Agent Ecosystem
  2. 0:39 Importance of Open Source in Machine Learning
  3. 2:36 Hugging Face Hub overview
  4. 3:06 Agentic models and Vision-LMs
  5. 4:24 Benchmark datasets and model filtering
  6. 5:16 Inference providers and model routing
  7. 6:50 Local coding agents and tools
  8. 7:46 Hermes agents for memory management
  9. 9:20 Traces repository for agent sessions
  10. 10:22 Tips for finding and serving local models
  11. 12:07 Supercharging agents with Hugging Face skills
  12. 13:41 Live demonstration of agent-driven fine-tuning
  13. 14:41 Training vision models (object detection/segmentation)
  14. 15:00 Using Model Context Protocol (MCP) for agents
  15. 16:30 Case study: OCR processing for AI papers

Open at this second