Videos eAXxdtNlK04
Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse
Scene timeline
40 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 166
- whisperx 166
- chunks
- 31
- from 166 cues
- keyframes
- 22
- kept of 40 captured
- frames with text
- 22
- 504 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 4.5 MB
- word timings on 166 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:39 | 1m 14s |
stt |
done | — | 2026-08-11 00:40 | 19s |
chunk |
done | — | 2026-08-11 00:41 | 0s |
text_embed |
done | — | 2026-08-11 00:41 | 1s |
keyframe |
done | — | 2026-08-11 00:41 | 1m 30s |
ocr |
done | — | 2026-08-11 00:42 | 12s |
frame_embed |
done | — | 2026-08-11 00:42 | 4s |
Frames, and what the machine read
-
- ealangfuse0.84
- AI ENGINEER· SAN FRANCISCO0.96
- by ClickHouse0.99
- A MINIMAL SELF-OPTIMIZATION EXPERIMENT1.00
- Stop burning tokens1.00
- Self-improvement needs domain expertise first.0.99
- Annabell Schäfer - Growth Engineer0.99
- SELF-IMPROVING AGENTS0.98
-
- langfuse0.87
- pre.0.99
- karpathy / autoresearch0.98
- pops1.00
- Peter Steinberger0.97
- 0 ..0.59
- @steipete1.00
- :ode1.00
- Issues 541.00
- * Pull requests 130.96
- Vertaling tonen0.99
- Here's your monthly reminder that you shouldn't be prompting coding0.98
- agents anymore.1.00
- autoresea0.99
- Public1.00
- You should be designing loops that prompt your agents.0.99
- 11:58 a.m. - 7 jun. 2026 · 8,3 mln. Weergaven0.94
- master1.00
- 3Branches0 Ta0.96
- "I don't write prompts anymore. I1.00
- Stop prompting coding agents. Design the loops1.00
- autoresearch —autonomous research loops,0.97
- have loops."0.99
- that prompt them.1.00
- 87.2k★.1.00
-
- langfuse0.90
- code_compiles1.00
-
- langfuse0.99
- The target you give an agent is1.00
- always incomplete1.00
- optimal1.00
- destination1.00
- original1.00
- aim1.00
- START1.00
-
- langfuse0.92
- 02 - CONTEXT0.99
- Langfuse builds the infrastructure to run0.99
- improvement loops - for humans and agents1.00
- ONLINE1.00
- ONLINE1.00
- OFFLINE1.00
- OFFLINE1.00
- OFFLINE1.00
- Trace1.00
- Monitor1.00
- Build1.00
- Experiment1.00
- Evaluate1.00
- datasets0.98
- traces· sessions0.97
- dashboards · LLM-as-0.96
- datasets· features-as-0.96
- prompts · models - code0.91
- judges - custom evals0.96
- agents· prompts0.95
- judge . feedback0.95
- tests1.00
- variants1.00
- annotation1.00
- DEPLOY1.00
- Know what's1.00
- Find what's1.00
- Capture what you1.00
- Test hypothesis1.00
- Check if it1.00
- going on0.96
- interesting/1.00
- want to work1.00
- worked1.00
- wrong1.00
- STOP BURNING TOKENS0.99
-
- elangfuse0.98
- 03 - THE LOOPS FOR AGENT SYSTEMS1.00
- What is the clearest cut target function we can find?0.99
- What can we learn about the role of target functions from running1.00
- auto-optimization on it?1.00
- STOP BURNING TOKENS0.98
-
- elangfuse0.97
- 03 - THE CLEAR CUT TARGET1.00
- A single-label classification task has a clear cut target1.00
- function1.00
- Item to categorize0.98
- Available labels1.00
- Accuracy1.00
- Order1.00
- True label = predicted0.99
- label?1.00
- Complaint1.00
- yes/no1.00
- Inquiry1.00
- True label0.99
- Order1.00
- STOP BURNING TOKENS1.00
-
- langfuse0.93
- The 'Agent'0.92
- Single label classifier1.00
- The optimizer1.00
- Target function0.99
- Single label classification1.00
- Claude running in loop to get0.99
- ground truth1.00
- to 92% accuracy1.00
- STOP BURNING TOKENS1.00
-
- PART TWO - THE EXPERIMENT0.98
- 011.00
- A minimal self-optimization loop0.99
- RANDOM PETER1.00
- our shorthand for a prompt hand-seeded by domain experts. The human baseline.0.99
-
- langfuse0.96
- 05 - THE EXPERIMENT0.97
- The setup1.00
- The Target Function1.00
- Classified arXiv paper— primary label, from title +0.99
- TASK1.00
- abstract1.00
- Classification allows for a clean0.99
- accuracy signal0.99
- DATASET1.00
- 200 fit· 100 validate· 300 test0.98
- The'Agent'0.99
- CLASSIFIER1.00
- GPT-5-4 nano—the cheaper model0.97
- Simple on purpose, to constrain the1.00
- degrees of freedom for change1.00
- START PROMPT1.00
- A flat list of labels0.99
- OPTIMIZER1.00
- Claude Opus4.8 — proposes prompt updates0.96
- The Optimization process1.00
- Powerful model as optimizer, helpful1.00
- GPT-5-4 prompting guide1.00
- information as context1.00
- REFERENCES1.00
- task.md for overall loop process description0.99
- COMPARISON1.00
- GPT 5-5 baseline performed with 84.2%0.99
- accuracy1.00
- STOP BURNING TOKENS1.00
- Source: Data licensed under Apache License 2.00.98
- https://www.kaggle.com/datasets/sumitm004/arxiv-scientific-research-papers-dataset1.00
-
- langfuse0.94
- 05 - THE INPUTS0.94
- The target function maps inputs to their expected labels1.00
- Dataset arxiv-cs10-singlelabel-v1-small-fit0.99
- + New item0.99
- Upload CSV0.97
- 个0.95
- Items1.00
- Experiments1.00
- Q Search...0.95
- IDs1.00
- Filters1.00
- Columns 8/81.00
- 目1.00
- Fit set:1.00
- Item id1.00
- ①Status0.96
- C. Input0.99
- Expected Output1.00
- Metadata1.00
- Actions1.00
- Run, look at errors, form0.99
- f71f8d3d-aa6...0.96
- Active1.00
- 2021.00
- 06-1.00
- 05:1.00
- Team, Faster than the1.00
- title: "Faster than the1.00
- 1 Itens0.86
- Engineering"1.00
- label: "Software0.98
- 11 Items0.86
- seed: 202606151.00
- split: "fit"0.98
- hypothesis for prompt updates0.99
- Customer: Tool Integration,1.00
- Collaboration, and1.00
- }0.92
- source_id: "abs-1.00
- 2606.01772v1"1.00
- Validation set:1.00
- 460a43e7-56...1.00
- Active1.00
- 05:0.99
- 2021.00
- 06-1.00
- 3 Itens0.91
- title: "From Failed0.99
- Repairing Harness Flaws"1.00
- Trajectories to Reliable1.00
- LLM Agents: Diagnosing and1.00
- 1 Itens0.94
- }0.95
- Engineering"0.99
- label: "Software0.98
- 11 Items0.91
- seed: 202606151.00
- split: "fit"0.99
- source_id: "abs-0.98
- 2606.06324v1"1.00
- Run to decide if a change is1.00
- accepted1.00
- 1079da80-d9...1.00
- Active1.00
- 2021.00
- 06-1.00
- 05:0.96
- title: "Agentic Persona1.00
- Generation with Critique-0.99
- Engineering"1.00
- label: "Software0.98
- 11 Itens0.91
- seed: 202606151.00
- split: "fit"0.96
- ***0.87
- Test set:1.00
- Refinement: An Industrial1.00
- Evaluation"1.00
- }0.90
- source_id: "abs-1.00
- 2606.09637v1"1.00
- Run after loop converged to check0.99
- 2a4fd057-92...1.00
- Active1.00
- 06-1.00
- 05:0.99
- 2021.00
- 3 Itens0.97
- Code Smells Across Project1.00
- title: "Comparing ML-1.00
- Specific and General Python1.00
- }0.81
- Engineering"1.00
- label: "Software0.98
- 11 Items0.89
- seed: 202606151.00
- split: "fit"0.99
- source_id: "abs-1.00
- generalization1.00
- Characteristics"1.00
- 2606.01882v1"0.99
- STOP BURNING TOKENS1.00
-
- langfuse0.96
- 05 - THE INPUTS0.94
- The base prompt is a flat list of available labels0.99
- flat_list_categories1.00
- #11.00
- pc-flat-loop1.00
- Probe1.00
- Prompt Config Linked Generations Use Prompt0.99
- Text Prompt1.00
- Classify this paper with a label.1.00
- Available labels:1.00
- - Artificial Intelligence1.00
- - Computation and Language (Natural Language Processing)1.00
- - Computer Vision and Pattern Recognition0.99
- - Cryptography and Security0.98
- - Databases1.00
- - Distributed, Parallel, and Cluster Computing1.00
- - Information Retrieval0.99
- - Networking and Internet Architecture1.00
- - Robotics1.00
- - Software Engineering0.99
- Title: {{title}}0.99
- Authors: {{authors}}0.98
- Summary: {{summary}}0.98
- Return only the label name, nothing else.0.99
- STOP BURNING TOKENS0.99
-
- 11.00
- # Prompt Auto-Improvement Loop1.00
- 21.00
- 06 - THE Optimization process0.98
- 31.00
- ## Context1.00
- 41.00
- You are improving a classification prompt through an automated1.00
- The loop is step by step1.00
- fit/validate/test loop.1.00
- 51.00
- The baseline prompt lives in Langfuse prompt management - fetch0.99
- it via the CLI at the start.1.00
- instructed via task.md0.98
- 61.00
- You interact with Langfuse exclusively via the Langfuse CLI.1.00
- 71.00
- Do not touch the test dataset until a stopping condition is met.0.99
- 81.00
- 90.99
- ## Datasets0.99
- 101.00
- - Fit: https://cloud.langfuse.com/project/0.99
- Run the base prompt on fit + validate1.00
- cmq74a8mb01yqad0j43ke3ax7/datasets/cmqfems63091rad0cmfg8q8nd/1.00
- items1.00
- Score accuracy — per item and overall0.97
- 111.00
- - Validate: https://cloud.langfuse.com/project/0.99
- cmq74a8mb01yqad0j43ke3ax7/datasets/cmqfenrra093mad0ctht6cntj/1.00
- Error analysis on the fit set — find patterns0.99
- items1.00
- 121.00
- - Test: untouched until a stopping condition is met1.00
- 131.00
- Propose an update for the biggest error category1.00
- 141.00
- ## Stopping conditions1.00
- 151.00
- Stop the loop when either of these is met:0.97
- Re-run — accept only if validate improves0.98
- 161.00
- 1. 92% accuracy on t0.97
- validate dataset0.99
- 171.00
- 2. 15 improvement loops completed1.00
- STOP IF 15 runs completed OR 92%1.00
- 181.00
- 191.00
- When a stopping condition is met, run the test dataset once and1.00
- accuracy reached1.00
- report final accuracy across all three splits.1.00
- 201.00
- Perform final run on test set1.00
- 211.00
- ## Loop0.99
- 221.00
- 231.00
- ### Step 0 - Fetch the prompt0.96
- 241.00
- Pull the latest production version0.99
- via the CLI. This is your baseline0.99
- STOP BURNING TOKENS0.99
- 251.00
- 261.00
- ### Step 1 - Baseline0.96
- 271.00
- 1. Run the current prompt against t0.98
-
- langfuse0.89
- 07 - RESULTS·QUALITY0.97
- The loop improves accuracy of the classification1.00
- Self-improvement undeniably lifts performance, and, by holding it on a cheaper model, saves cost.0.99
- Validation accuracy per prompt iteration — keep/discard decision boundary0.99
- Kept1.00
- Discarded1.00
- 1.01.00
- EXacth act-rac0.57
- 0.81.00
- 0.7801.00
- 0.7701.00
- 0.8301.00
- 0.7801.00
- 0.7701.00
- 0.790.8300.96
- Overall level remains0.97
- aroun80%0.94
- 0.6801.00
- -0.6800.99
- 0.61.00
- +15% from0.98
- 0.41.00
- baseline1.00
- Test set1.00
- 0.21.00
- accuracy:1.00
- 80.2%1.00
- 0.01.00
- 21.00
- 31.00
- 41.00
- 51.00
- 61.00
- 71.00
- Prompt version0.99
- STOP BURNING TOKENS1.00
-
- langfuse1.00
- 07 - RESULTS · QUALITY0.96
- The loop clarified decision boundaries and added1.00
- examples1.00
- Changes v1 → v40.95
- Classify this paper with a label.0.99
- Classify this paper with a label.0.99
- Added general classification approach0.99
- transfor1.00
- Reserve Artificial Intelligence for a0.96
- Added information on how to decide1.00
- "Uses a neural network / LLM / learning" is not by itself a reason0.94
- generating, translating.0.96
- sunnarizing, dialogue, QA,0.87
- between two similar classes1.00
- - A faster approximate-mearest-neighbour / vector index for search -» Databases (the0.87
- Added examples on specific frequently0.99
- mendations for a user's query -> Infarnation0.91
- confused labels0.99
- Scheduling or running jobs across a Gu cluster - Bistributed, Parallel, and Cluster0.83
- sage, even if it retrieves or ranks passages as a step.0.99
- Available labels:0.95
- - Artificial Intelligence0.98
- - Computation and Language (Natural Language Pracessing)0.95
- Computer Wision and Pattern Recognition0.93
- Available labels:0.96
- Artificial Intelligence0.97
- - Seftware Engineering0.97
- Sunnary: ((suary))0.76
- - Distributed, Paratlel, and Cluster Computing0.99
- Enformation Retrieval0.90
- - Software Engineering0.97
- Titte: ((tite))0.81
- Authors: €authors?}0.88
- Sunnary: (<summary))0.75
- - Distributed, Parallel, and Cluster Conputing0.97
- - Information Retrieval0.94
- - Robotics0.95
- Networking and Baternet Architecture0.90
- Return only the label nase,0.93
- Return only the label name, nothing else.0.97
- Improvement loop added rules and examples1.00
- STOP BURNING TOKENS0.99
-
- ealangfuse0.92
- 07 - RESULTS· QUALITY0.97
- The classes and the clear right/wrong give high signal0.99
- feedback for the optimizer to improve the setup1.00
- Both baselines complete and logged: Fit 68.0%· Validate 68.0%.0.97
- Error analysis (fit, 64 errors): The dominant pattern is unambiguous - "Artificial Intelligence" is0.99
- the #1 wrong-answer sink (22/64), "Computation and Language (NLP)" #2 (11/64). The biggest0.99
- specific confusions are Information Retrieval → AI (8×) and Networking → AI (5×). The bare baseline0.98
- has zero guidance, so papers that merely use ML get dumped into the two method-magnet labels0.99
- instead of their true application domain.1.00
- Loop-1 will address exactly this pattern: add a concise "classify by primary domain/contribution,1.00
- not the technique; AI and NLP are over-chosen method labels" rule. Let me update tracking and write0.99
- the loop-1 build script.0.99
- Clearly quantifiable1.00
- Data backed'update'1.00
- Enough examples to draw0.98
- and reliable failure0.97
- +0.90
- hypothesis1.00
- +0.89
- general conclusions from0.98
- =0.87
- modes1.00
- STOP BURNING TOKENS0.98
-
- ight/wrong give high signal0.99
- to improve the setup1.00
- date68.0%.1.00
- n is unambiguous - "Artificial Intelligence" is0.99
- anguage (NLP)"#2(11/64). The biggest0.98
- (8×) and Networking → AI(5×). The bare baseline0.97
- dumped into the two method-magnet labels1.00
- cise "classify by primary domain/contribution,0.99
- dlabels" rule. Let me update tracking and write0.98
-
- idate 68.0%.0.89
- n is unambiguous - "Artificial Intelligence" is0.98
- anguage (NLP)"#2(11/64). The biggest0.98
- (8×) and Networking → AI(5×). The bare baseline0.98
- dumped into the two method-magnet labels1.00
- cise "classify by primary domain/contribution,0.98
- labels" rule. Let me update tracking and write1.00
- ough examples to draw0.99
- eneral conclusions from1.00
Transcript
166 cues· 3,151 words· 17,090 chars
- 0:00 Hi everyone, my name is Annabelle, I'm a growth engineer at Lengfuse and we're the largest open source observability and evaluation platform for your AI system.
- 0:08 And today I'm going to share about how you should stop burning your tokens and why you should start building in domain expertise early in into your loop design but also into your overall application design to make sure you're continuously improving and updating your application.
- 0:24 It is June 2026, and the whole internet is about loops right now.
- 0:29 So we have Boris Czerny saying he doesn't write any prompts anymore.
- 0:32 He has loops.
- 0:33 And Peter Steinberger being like, you should be designing loops and not prompt your agents.
- 0:37 And I mean, the whole capacity auto research and auto improvement topic blew up earlier this year already.
- 0:43 But all of them are coming a little bit from the developer perspective.
- 0:46 And coding is one for a good reason, one of the cases where this whole
- 0:53 automatically reaching a goal work quite well, because they've always had at least one target function that was quite clear, and this was, does the code compile or not?
- 1:02 And of course, just because code compiles doesn't mean it's great and doesn't mean that all the features are exactly how you wanted them, but you at least know that you shipped something that worked.
- 1:11 And of course, this was then expanded over time to expand this target function, but essentially it's a, does it work or does it not?
- 1:20 In all other fields, and especially if you're building AI applications for some kind of domains like medicine compliance or healthcare or all kinds of chatbots, these target functions are not nearly as clear.
- 1:34 And a target that you give an agent is actually also always incomplete.
- 1:38 So you might initially think that you're trying to head down here, but actually your optimal destination is up there.
- 1:45 And to get there and to understand this, it takes quite some time to figure this out.
- 1:50 And it's just inherently a difficult problem, especially if you're in a field where a clear yes or no, like does the code compile, doesn't really cut it.
- 2:00 So it is very difficult, but at the same time, we at LengQ see that the teams who are investing heavily here in the middle, so making sure they capture what they actually want to work, so the target function, and build this out and make sure they have good evaluators that are evaluating this, are the ones who manage to continuously upgrade and improve their application over time, and also ship with confidence, because if you know it's working as you intended to, then you can also sleep well at night in case you push a code change.
- 2:28 Okay, so now we know it's super important and you should be doing it, but at the same time it's almost impossible to actually get it right.
- 2:37 So we were setting out and wondering what's the clearest cut target function we can find for an agent to use?
- 2:43 And also if you run an auto optimization on it, what can we learn about the role of target functions from running it on it?
- 2:51 And the clearest cut target function we could find was a single label classification task that has a very clear cut yes or no.
- 2:58 So, for example, let's say we have an item to categorize.
- 3:01 There's a true label, for example, oh, it's an order.
- 3:03 And then we have a set of available labels, order, complaint, inquiry.
- 3:06 And our classifier then assigns one label.
- 3:09 And you can very clearly say, is the true label equal to the predicted label?
- 3:13 And if yes, then this one is right and the next one might be wrong.
- 3:17 And overall, you can calculate an accuracy value.
- 3:20 and get a very clear signal from here.
- 3:23 So we put this target function together with an agent and also with an optimizer and how this minimal loop looked like I'm going to share with you now.
- 3:32 So here's our minimal self-optimization loop.
- 3:35 Up there in the first two rows, we can see our target function.
- 3:39 In our case, we opted for classified archive papers.
- 3:42 So it's a set of papers that, based on their title and abstract, got a primary label from the author to categorize it for the other people in research.
- 3:55 And we have the ground truth here, and we have it in 200 items in a fit dataset, 100 in a validate dataset, and 300 in a test dataset, just to also make sure we're not overfitting.
- 4:07 Then we have our agent, which is more or less just a simple prompt based on GPT-5 for nano, because we wanted to see how a very small and cheap model performs on the auto-improvement, because the good ones actually got really, really good, but also very expensive.
- 4:22 over time, and we have this flat list of labels, and then we have our optimization process that runs through Cloud Code, leveraging Cloud Opus 4.8, so one of the frontier models, and it proposes the prompt updates and has this context reference, the GPT-5-4 prompting guide, as well as a task MD that is describing the loop.
- 4:40 Looking a little bit closer on how this looks like, here we can see our target function.
- 4:44 In this case, especially our fit data set within length view.
- 4:48 So we see an input column where the title and the abstracts are inside, as well as an expected output column.
- 4:53 So what kind of label should be applied?
- 4:56 And on those three different data sets, the idea is on the fit data set, you run it, you look for the errors and especially error clusters.
- 5:05 And Opus should then formulate hypothesis for prompt updates.
- 5:09 In the validation set, we then check if those prompt updates actually also generalize to unseen data.
- 5:15 And then finally, when we're reaching a plateau or our stopping criteria, then we're running it on a test set to see how well we actually generalize for untouched data that was not part of the training process.
- 5:28 Our base prompt is a very flat list of labels and just a simple task, classify this paper with a label.
- 5:34 Of course, if we would run the prompt, we would probably add some pros, how to think about it and all of this.
- 5:39 But we just wanted to see what happens if we use the very base version of this prompt and how the system is dealing with it.
- 5:48 And our loop is a step-by-step instructed loop done through a task markdown file.
loading
Chapters
- 0:00 Introduction and the goal of avoiding token waste
- 0:24 The current trend of "designing loops" vs. prompt engineering
- 0:47 Why coding (with its clear "compile" target) set a false precedent
- 1:20 Challenges of defining target functions in non-coding domains
- 2:30 Case study: The arXiv paper classification experiment
- 3:30 Setup of the minimal self-optimization loop
- 5:48 The step-by-step iteration process and stopping criteria
- 7:00 Results: Achieving a 15% improvement in accuracy
- 9:53 Deep dive into the reasoning behind the improvement
- 11:26 Translating binary "high signal" feedback to other applications
- 13:12 Defining what "good" looks like for your specific domain
- 15:05 The importance of human-agent collaboration and data review
- 15:57 Don't review it only with your coding agents, but review it as a human.