Videos OV56RddyFuU
Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning — Merve Noyan, Hugging Face
Scene timeline
71 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 201
- whisperx 201
- chunks
- 33
- from 201 cues
- keyframes
- 64
- kept of 71 captured
- frames with text
- 64
- 2,021 lines read
- chapters
- 15
- from the source metadata
- keyframe bytes
- 8.2 MB
- word timings on 201 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 16:56 | 2m 03s |
stt |
done | — | 2026-08-10 16:58 | 24s |
chunk |
done | — | 2026-08-10 16:58 | 0s |
text_embed |
done | — | 2026-08-10 19:52 | 1s |
keyframe |
done | — | 2026-08-10 16:59 | 1m 48s |
ocr |
done | — | 2026-08-10 17:00 | 38s |
frame_embed |
done | — | 2026-08-10 19:52 | 10s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- eer0.96
- raintrus1.00
- ze1.00
- jineer0.97
- Open Agent1.00
- OPE0.98
- eer1.00
- gger.de1.00
- Ecosystem1.00
- odal0.99
- Engineer1.00
- (and having an Intern Al Engineer at your fingertips)0.99
- EUROPE1.00
- AlEngineer0.98
- EUROPE1.00
-
- eer0.97
- aintrus1.00
- ze1.00
- ineer1.00
- Open/source1.00
- eer1.00
- ger.de1.00
- odal0.97
- ngineer1.00
- EUROPE1.00
- AlEngineer0.98
- EUROPE1.00
-
- er1.00
- raintrus1.00
- ze1.00
- jineer0.92
- PE0.99
- Open/source1.00
- eer0.96
- 1ger.de0.99
- dal1.00
- jineer1.00
- OPE1.00
- AlEngineer0.99
- EUROPE1.00
-
- B1.00
- trust1.00
- Open-source1.00
- deepseek-ai/DeepSeek-V3.1-Terminuslike291Follow DeepSeek97.8k0.98
- Text GenerationTransformersSafetensors deepseek_v30.99
- conversational0.98
- custom_code1.00
- text-generation-inference1.00
- License: mit0.98
- Model card1.00
- Files and versions o xet0.91
- Community1.00
- ¿ Edit model card0.97
- dev1.00
- DeepSeek-V3.1-Terminus1.00
- Downloads last mont0.99
- 11,0701.00
- deepseek1.00
- Safetensors0.97
- Model size 685B params Tensortype BF16·F8_E4M3 · F320.94
- 0 Chat template Files info0.93
- er1.00
- Inference Providers1.00
- Novita1.00
- Hugging Fac0.98
- X Twitter0.93
- deepseek al0.97
- Text Generation1.00
- Examples1.00
- Input a message to start chatting with deepseek-ai/DeepSeek-V3.1-Terminus.1.00
- Introduction1.00
- This update maintains the model's original capabilities while addressing issues1.00
- reported by users, including:1.00
- I anmuaae rnncictanru- Darlurina inetanree nf mivad Chinaca.Fnalich tovt and0.89
- AlEngineer0.97
- EUROPE1.00
-
- Open-source1.00
- deepseek-ai/DeepSeek-V3.1-Terminuslike2911.00
- Follow DeepSeek 97.8k0.98
- Text Generation1.00
- Transformers1.00
- Safetensors1.00
- deepseek_v31.00
- conversational1.00
- custom_code1.00
- text-generation-inference1.00
- License: mit1.00
- Model card1.00
- Files and versionsxet1.00
- Community 100.99
- Edit model card0.98
- 三0.69
- DeepSeek-V3.1-Terminus1.00
- Downloads last month1.00
- 11,0701.00
- deepseek1.00
- Safetensors①0.96
- Model size 6858 params0.98
- Tensor type BF16 · F8_E4M3 · F320.96
- φ Chat template0.95
- Files info0.99
- Inference ProvidersNEw0.98
- Novita1.00
- DeepSeek Homepage1.00
- Chat DeepSeek V31.00
- Hugging Face1.00
- DeepSeek AI0.97
- Text Generation0.98
- Examples1.00
- Discord DeepSeek AI1.00
- WeChat DeepSeek AI0.96
- X Twitter0.96
- deepseek ai0.96
- License MIT0.99
- Input a message to start chatting with deepseek-ai/DeepSeek-V3.1-Terminus.1.00
- Introduction1.00
- This update maintains the model's original capabilities while addressing issues1.00
- reported by users, including:1.00
- I anauaae roncictancu Deducina inctancec nf mived Chinece-Fnalich tovt and0.89
-
- onAI0.83
- Why does open-source matter?1.00
- • Absolute control over models1.00
- • Cost reduction in multipliers1.00
- • Further customize, shrink models depending on your needs0.98
- • Guaranteed privacy for the end-user, enable on-0.99
- device/in-browser uses, data doesn't go to another server0.99
- NTRY1.00
- lind0.83
- AlEngineer0.98
- EUROPE1.00
-
- Why does open-source matter?1.00
- • Absolute control over models1.00
- • Cost reduction in multipliers1.00
- • Further customize, shrink models depending on your needs0.99
- • Guaranteed privacy for the end-user, enable on-0.99
- device/in-browser uses, data doesn't go to another server1.00
- AlEngineer0.98
- EUROPE1.00
-
- Artificial Analysis Index0.98
- Artificial Analysis Intelligence Index by Open Weights / Proprietary1.00
- 28 of 472 models1.00
- 00.67
- Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, τ2-Bench Telecom, Terminal-0.99
- + Add model from specific provider0.97
- Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt1.00
- Proprietary1.00
- Open Weights0.96
- A Artificial Analysis0.99
- 571.00
- 571.00
- 531.00
- 521.00
- 521.00
- 511.00
- 501.00
- 501.00
- 491.00
- 491.00
- 481.00
- 471.00
- 461.00
- 451.00
- 421.00
- 391.00
- 371.00
- 361.00
- 361.00
- 341.00
- 331.00
- 321.00
- 271.00
- 261.00
- 241.00
- 241.00
- 181.00
- G1.00
- AI0.91
- 80.64
- AI0.93
- Z0.87
- a0.51
- K0.93
- G1.00
- a0.67
- G1.00
- AI0.93
- G1.00
- M0.79
- 80.83
- C0.93
- C0.94
- C0.95
- C0.71
- Ca0.59
- C0.94
- C0.86
- C0.93
- C0.97
- C.0.61
- C0.95
- C0.96
- C0.97
- C0.88
- Cn0.58
- C0.95
- (max)0.84
- Gemini 3.11.00
- Pro Preview0.99
- GPT-5.40.99
- Claude Opus0.96
- 4.6 (max)0.98
- Claude1.00
- GLM-5.190.99
- MiniMax-M2.7 ?0.93
- Grok 4.200.97
- 0309 v20.84
- GPT-5.4 mini0.96
- (xhigh)0.93
- Kimi K2.5 90.92
- Gemini 30.98
- Flash0.99
- Qwen3.5 397B0.98
- Haiku1.00
- Nova 2.0 Pro1.00
- Preview1.00
- (medium)0.97
- Preview1.00
- gpt-oss-120B1.00
- K-EXAONE ?0.91
- gpt-oss-20B1.00
- K2 Think V2 90.89
- Llama 4 Maverick1.00
-
- Hugging Face Hub1.00
- Home for the open-source machine learning community: share & discover models,0.99
- datasets, apps, connect with the community and more!1.00
- Hugging Face1.00
- O Search models, datasets, users...0.96
- Models1.00
- Datasets1.00
- Spaces1.00
- Community0.97
- Docs Pricing0.99
- i0.56
- + New0.98
- Following3551.00
- New Post0.99
- Trending1.00
- last 7 days0.99
- Models Datasets0.98
- Spaces1.00
- Papers Collections Community1.00
- Posts1.00
- Models1.00
- Datasets1.00
- Spaces1.00
- merve1.00
- Upvotes Likes Articles0.99
- Profile1.00
- moonshotai/Kimi-K2-Instruct1.00
- Inbox (7113)1.00
- hysts updated a Space 2 minutes ago0.98
- Text GenerationUpdated...25.1k1.23k0.99
- Settings1.00
- $ Billing0.99
- SeqTex0.94
- DeepSite v20.99
- 10.2k1.00
- SeqTex generates texture based on textual conditions1.00
- Generate any application with DeepSeek0.99
- Organizations1.00
- Hugging Face0.99
- Wauplin updated a Space5 minutes ago1.00
- HuggingFaceTB/SmolLM3-3B1.00
- Text Generation3BUp...59.5k4850.97
- G Google0.98
- Responses.js1.00
- 090.83
- Deprem Yapay Zeka1.00
- Check out https://github.com/huggingface/responses.js0.98
- mistralai/Devstral-Small-25071.00
- Notebooks-explorers1.00
- Text Generation 24BU...9.58k2330.95
- - SODA0.89
- — PyTorch Image Models0.95
- Weyaxi updated a dataset 5 minutes ago0.98
- black-forest-labs/FLUX.1-Kontext-dev0.99
- Image-to-Image·Updated ...274k1.68k0.95
- Templates1.00
- Weyaxi/huggingface-leaderboard1.00
-
- Hugging Face Hub1.00
- Home for the machine learning community: share & discover1.00
- mo1.00
- 2.7M models of all libraries and tasks1.00
- more!1.00
- 925k+ datasets1.00
- + Ne0.94
- Spaces (apps)1.00
- New Post1.00
- Trending1.00
- last 7 days1.00
- Collections1.00
- Community1.00
- Posts1.00
- AIl0.53
- Models1.00
- Datasets1.00
- Spaces1.00
- merve0.98
- Prof1.00
- ecosystem: open-source libraries1.00
- cruct0.98
- Inbo0.98
- 25.1k1.00
- Setti0.99
- $ Billir0.92
- (transformers, diffusers & more!)1.00
- 10.2k1.00
- ith DeepSeek1.00
- Organiza1.00
- HuggingFaceTB/SmolLM3-3B1.00
- Hugging Face1.00
- Wauplin updated a Space5 minutes ago0.97
- Text Generation 3B·Up...59.5k0.96
- 4851.00
- G Google1.00
- Responses.js1.00
- Deprem Yapay Zeka0.97
- Check out https://github.com/huggingface/responses.js0.99
- mistralai/Devstral-Small-25071.00
- Notebooks-explorers1.00
- - SODA0.88
- —PyTorch Image Models0.97
- Weyaxi updated a dataset 5 minutes ago0.98
- black-forest-labs/FLUX.1-Kontext-dev0.99
- Templates1.00
- Weyaxi/huggingface-leaderboard1.00
- Vorse0.70
-
- /models1.00
- Main1.00
- Tasks1.00
- Libraries1.00
- Languages Licenses0.97
- Models2,665,3241.00
- Filter by name0.98
- Full-text search0.99
- Inference Available0.99
- ↑↓ Sort: Trending1.00
- Other1.00
- Tasks1.00
- Image-Text-to-Text·: 36B·Updated about 3 hours ago·259k·6040.94
- Qwen/Qwen3.5-35B-A3B1.00
- Text Generation0.99
- Any-to-Any0.96
- P0.61
- Image-Text-to-Text1.00
- Image-to-Text1.00
- Qwen/Qwen3.5-27B1.00
- Image-Text-to-Text28B·Updated 2 days ago108k3960.98
- 图0.72
- Image-to-Image1.00
- Text-to-Image0.98
- Text-to-Video1.00
- Text-to-Speech1.00
- +441.00
- Qwen/Qwen3.5-397B-A17B1.00
- Image-Text-to-Text 403B·Updated 4 days ago·726k1.11k0.96
- Parameters1.00
- <1B0.99
- 6B0.96
- 12B0.99
- 32B1.00
- 128B0.90
- >500B0.99
- Qwen/Qwen3.5-122B-A10B1.00
- Image-Text-to-Text · 125B · Updated 3 days ago·108k· 3240.91
- Libraries1.00
- unsloth/Qwen3.5-35B-A3B-GGUF1.00
- PyTorch0.97
- TensorFlow1.00
- X JAX0.82
- Image-Text-to-Text 35B·Updated 3 days ago·265k·2750.94
- Transformers1.00
- Diffusers1.00
- zai-org/GLM-51.00
- sentence-transformers1.00
- Safetensors1.00
- Text Generation754B·Updated 14 days ago189k1.63k0.96
- ONNX1.00
- GGUF1.00
- Transformers.js0.98
-
- models → agents, serve locally1.00
- Agentic LLMs (thinking + tool calling): gpt-oss, Gemma-4,0.99
- Minimax M2.7, GLM-5, Nemotron3-Super0.99
- Agentic vision models (thinking + CUA): Qwen3.50.99
- (Alibaba), Kimi-K2.50.98
- mlx_lm.generate --prompt "How tall is Mt Everest?"0.99
- vllm serve Qwen/Qwen3-8B # then query with OpenAI Completion1.00
- llama-server -m model.gguf --port 80801.00
-
- compare open models1.00
- Main1.00
- Tasks1.00
- Libraries1.00
- Languages Licenses Other1.00
- Datasets 171.00
- Filter by name1.00
- Full-text search1.00
- ↑↓ Sort: Trending0.98
- Modalities0.99
- openai/gsm8k1.00
- allenai/olmOCR-bench1.00
- 3D1.00
- Audio1.00
- Document1.00
- Geospatial1.00
- Benchmark·Updated 17 days ago·17.6k·±765k1.24k0.96
- Benchmark · Updated Feb 19·± 4.14k · 1760.90
- Image1.00
- Tabular1.00
- Text1.00
- Time-series1.00
- ScaleAI/SWE-bench_Pro1.00
- cais/hle1.00
- Video1.00
- Benchmark·Updated Feb 23·731·±692k·780.91
- Benchmark · Updated Jan 20· 2.5k·±46.1k ·7610.91
- Size (rows)1.00
- SWE-bench/SWE-bench_Verified1.00
- collinear-ai/yc-bench1.00
- <1K0.93
- >1T0.95
- Benchmark·Updated Feb 27· 500·126k·280.92
- Benchmark - Updated 17 days ago · ± 104 · 150.91
- TIGER-Lab/MMLU-Pro1.00
- MathArena/aime_20261.00
- Format0.97
- Benchmark · Updated 29 days ago·12.1k ·± 113k·4640.92
- Benchmark · Updated Feb 16 · 30± 12.7k · 280.90
- 4 json0.87
- EE CSV0.78
- parquet0.94
- optimized-parquet1.00
- imagefolder1.00
- soundfolder1.00
- webdataset1.00
- harborframework/terminal-bench-2.01.00
- Benchmark · Updated Feb 17 · ± 2.98k - 190.89
- Idavidrein/gpqa1.00
- Benchmark · Updated Mar 5· 1.25k ·± 105k · 4080.87
- text0.94
- arrow0.98
- mteb/arguana1.00
- FutureMa/EvasionBench1.00
- Type1.00
- Benchmark·Updated Feb 22·11.5k-± 12.8k·50.91
- Benchmark · Updated Feb 19·16.7k·± 177·850.92
- Benchmark×1.00
- è Traces0.94
- mteb/BRIGHT1.00
- hf-audio/open-asr-leaderboard1.00
- Benchmark · Updated 7 days ago·1.35M - ±662-20.90
- Benchmark - Updated 6 days ago · 99.5k · ± 20.1k - 70.90
- likaixin/ScreenSpot-Pro1.00
- nvidia/compute-eval1.00
- Benchmark - Updated 22 days ago ·± 7.73k - 600.91
- Benchmark - Updated 20 days ago · 2.46k - ± 5.63k · 220.89
-
- compare open models1.00
- Datasets:ScaleAl/SWE-bench_Pro1.00
- like1.00
- 781.00
- FollowScale Al0.98
- 2841.00
- Benchmark1.00
- Modalities:1.00
- Text1.00
- Formats:1.00
- parquet0.95
- Size:1.00
- <1K1.00
- Libraries:1.00
- Datasets1.00
- pandas1.00
- Polars1.00
- +11.00
- Dataset card1.00
- 田Data Studio1.00
- Files and versions xet0.96
- Community1.00
- Leaderboard Official Benchmark0.97
- ① Learn more0.98
- Experimental1.00
- Task: SWE Bench Pro0.97
- #1.00
- MODEL1.00
- SCORE1.00
- zai-org/GLM-5.10.97
- 58.4*1.00
- 21.00
- MiniMaxAI/MiniMax-M2.51.00
- 55.41.00
- 31.00
- moonshotai/Kimi-K2.51.00
- 50.71.00
- Qwen/Qwen3-Coder-Next1.00
- source1.00
- 44.3*1.00
- 51.00
- Qwen/Qwen3-Coder-480B-A35B-Instruct0.99
- source1.00
- 38.71.00
- Show all 14 models0.99
-
- Inference Providers for Qwen/Qwen3-VL-235B-A22B-Thinking0.99
- ×0.80
- Novita1.00
- ★ Auto0.89
- mervenoyan1.00
- t routing0.98
- Python1.00
- JavaScript1.00
- CcURL0.95
- huggingface_hub0.98
- requests openai1.00
- Stream1.00
- import os0.98
- Copy1.00
- from huggingface_hub import InferenceClient1.00
- client = InferenceClient(0.99
- vibe-check1.00
- provider="novita",1.00
- api_key=os.environ["HF_TOKEN"],0.99
- completion = client.chat.completions.create(1.00
- model="Qwen/Qwen3-VL-235B-A22B-Thinking",1.00
- go serverless1.00
- messages=[1.00
- {0.85
- "role": "user",0.99
- "content": [0.96
- {0.95
- "type": "text",0.98
- "text": "Describe this image in one sentence."1.00
- 3,0.64
- {0.92
- "type": "image_url",0.99
- "image_url": {0.99
- "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Islan1.00
- print(completion.choices[0].message)1.00
-
- compare providers across models0.99
- Inference Providers ·Metrics for top trending models0.98
- Browse all0.96
- Filter by model or provider...1.00
- Model0.98
- Provider1.00
- Input $/1M0.98
- Output $/1M0.99
- Context1.00
- Latency(s)1.00
- Throughput(t/s)0.99
- Ggoogle/gemma-4-31B-it0.99
- novita1.00
- $0.141.00
- $0.401.00
- 262,1441.00
- 1.301.00
- 301.00
- G google/gemma-4-26B-A4B-it0.98
- novita1.00
- $0.131.00
- $0.401.00
- 262,1441.00
- 2.961.00
- 131.00
- Qwen/Qwen3.5-9B0.97
- 370.87
- together1.00
- $0.101.00
- $0.151.00
- 262,1441.00
- 0.791.00
- 621.00
- zai-org/GLM-51.00
- 70.73
- novita cheapest0.97
- $1.001.00
- $3.201.00
- 202,8001.00
- 2.411.00
- 340.97
- zai-org/GLM-51.00
- 70.68
- together1.00
- $1.001.00
- $3.201.00
- 202,7521.00
- 0.611.00
The page's on-screen-text budget of 600
lines is spent, so the last cards in this grid list fewer lines than they
hold. Narrow the page with ?frames= to read them.
Transcript
201 cues· 2,780 words· 14,580 chars
- 0:15 Hello everyone and welcome to this talk in OpenAgent ecosystem and I would like to call it having an AI engineer at your fingertips.
- 0:25 I'm Merve and I work in the open source team of HuggingFace.
- 0:28 How many of you are using HuggingFace on daily basis?
- 0:33 Oh, let's change that.
- 0:35 This is not OK.
- 0:38 But first, let's talk a bit about open source and what it is.
- 0:40 So when it comes to machine learning, open source is absolutely differential.
- 0:45 Basically, you have the open weight models that go in with non-commercial licenses.
- 0:52 We call them open weight.
- 0:53 And then we have open source models that have
- 0:56 commercially available licenses, such as this one from DeepSeek.
- 1:00 It's called the MIT license or Apache 2.0.
- 1:03 And then there is even more open models that have the code open.
- 1:09 If you have agents, the harness is open.
- 1:12 Everything is open.
- 1:13 And this matters even more by the fact that yesterday or the other day, it was revealed that the cloud performance was going down.
- 1:24 So if you have everything in the open, nothing changes without you knowing no performance degradation, without you knowing everything's great.
- 1:34 But on top of it, if you have access to the weights, you can shrink them.
- 1:39 You can quantize them.
- 1:41 You can fine tune them if you feel like it.
- 1:44 And it's absolute guaranteed privacy for your end user because you can deploy it to edge devices, browsers without the data going somewhere else.
- 1:54 This matters a lot in my opinion, even more these days with the security breaches and everything.
- 2:02 And there was this argument, maybe a few years ago, that open source models aren't as good as closed models.
- 2:08 No, this is not the case.
- 2:09 Like you see, for instance, the latest GLM 5.1 is absolutely crushing it.
- 2:14 And I'm actually using it in my coding setup.
- 2:18 This is the artificial analysis intelligence index.
- 2:22 And the green ones are open models.
- 2:24 Meanwhile, the black ones are the closed models.
- 2:28 And we just catched up.
- 2:30 And we will catch up even more with the upcoming models and stuff.
- 2:35 And let's go back to Hugging Face Hub.
- 2:37 So everything is facilitated through Hugging Face Hub, all of the open releases.
- 2:43 It's the infra layer for all of your open source workflows.
- 2:48 And as of now, it's hosting even more models.
- 2:50 I should have updated the number.
- 2:52 It's probably close to 3 million.
- 2:54 A lot of data sets, spaces, and everything.
- 2:57 But that's not all when it comes to the iGenetic ecosystem.
- 3:00 And this is what we are going to talk about today.
- 3:03 So when you go to the models, you can filter for agentic models.
- 3:09 They are mostly the trending ones.
- 3:12 And there is two types of models, in my opinion.
- 3:15 There is the vision LMs, and then there is the LLMs.
- 3:19 And the vision LMs can also act as a computer use agent over the screenshots.
- 3:24 They know where to click, et cetera, which is pretty cool.
- 3:27 And one trend I have recently noticed is the fact that
- 3:32 you have labs releasing their LLMs with vision capabilities, day zero.
- 3:39 Like, for instance, the GEMA-4 was an omni model, and still it's an agentic model.
- 3:45 There is Q1 3.5.
loading
Chapters
- 0:00 Introduction to Open Agent Ecosystem
- 0:39 Importance of Open Source in Machine Learning
- 2:36 Hugging Face Hub overview
- 3:06 Agentic models and Vision-LMs
- 4:24 Benchmark datasets and model filtering
- 5:16 Inference providers and model routing
- 6:50 Local coding agents and tools
- 7:46 Hermes agents for memory management
- 9:20 Traces repository for agent sessions
- 10:22 Tips for finding and serving local models
- 12:07 Supercharging agents with Hugging Face skills
- 13:41 Live demonstration of agent-driven fine-tuning
- 14:41 Training vision models (object detection/segmentation)
- 15:00 Using Model Context Protocol (MCP) for agents
- 16:30 Case study: OCR processing for AI papers