read-only demo

Videos aHhB3sjGjkI

Agents Building Agents - Alfonso Graziano, Nearform

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:30:14· en-US· indexed 2026-08-11 08:19

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:14, 1 of 1 keyframes kept
  2. Shot 1, 0:14 to 0:36, 1 of 1 keyframes kept
  3. Shot 2, 0:36 to 1:09, 1 of 1 keyframes kept
  4. Shot 3, 1:09 to 1:26, 1 of 1 keyframes kept
  5. Shot 4, 1:26 to 1:53, 1 of 1 keyframes kept
  6. Shot 5, 1:53 to 2:20, 1 of 1 keyframes kept
  7. Shot 6, 2:20 to 2:47, 0 of 1 keyframes kept
  8. Shot 7, 2:47 to 3:20, 1 of 1 keyframes kept
  9. Shot 8, 3:20 to 3:49, 1 of 1 keyframes kept
  10. Shot 9, 3:49 to 4:18, 0 of 1 keyframes kept
  11. Shot 10, 4:18 to 4:47, 0 of 1 keyframes kept
  12. Shot 11, 4:47 to 5:16, 0 of 1 keyframes kept
  13. Shot 12, 5:16 to 5:43, 1 of 1 keyframes kept
  14. Shot 13, 5:43 to 6:09, 1 of 1 keyframes kept
  15. Shot 14, 6:09 to 6:51, 1 of 1 keyframes kept
  16. Shot 15, 6:51 to 7:20, 1 of 1 keyframes kept
  17. Shot 16, 7:20 to 7:49, 0 of 1 keyframes kept
  18. Shot 17, 7:49 to 8:17, 0 of 1 keyframes kept
  19. Shot 18, 8:17 to 8:27, 1 of 1 keyframes kept
  20. Shot 19, 8:27 to 8:53, 1 of 1 keyframes kept
  21. Shot 20, 8:53 to 9:41, 1 of 1 keyframes kept
  22. Shot 21, 9:41 to 9:44, 1 of 1 keyframes kept
  23. Shot 22, 9:44 to 9:48, 0 of 1 keyframes kept
  24. Shot 23, 9:48 to 9:52, 0 of 1 keyframes kept
  25. Shot 24, 9:52 to 9:56, 0 of 1 keyframes kept
  26. Shot 25, 9:56 to 10:12, 0 of 1 keyframes kept
  27. Shot 26, 10:12 to 10:48, 1 of 1 keyframes kept
  28. Shot 27, 10:48 to 11:23, 0 of 1 keyframes kept
  29. Shot 28, 11:23 to 11:33, 1 of 1 keyframes kept
  30. Shot 29, 11:33 to 12:21, 1 of 1 keyframes kept
  31. Shot 30, 12:21 to 13:01, 1 of 1 keyframes kept
  32. Shot 31, 13:01 to 13:30, 1 of 1 keyframes kept
  33. Shot 32, 13:30 to 13:55, 1 of 1 keyframes kept
  34. Shot 33, 13:55 to 14:21, 0 of 1 keyframes kept
  35. Shot 34, 14:21 to 14:46, 0 of 1 keyframes kept
  36. Shot 35, 14:46 to 15:03, 1 of 1 keyframes kept
  37. Shot 36, 15:03 to 15:28, 1 of 1 keyframes kept
  38. Shot 37, 15:28 to 15:56, 1 of 1 keyframes kept
  39. Shot 38, 15:56 to 16:23, 0 of 1 keyframes kept
  40. Shot 39, 16:23 to 17:10, 1 of 1 keyframes kept
  41. Shot 40, 17:10 to 17:38, 1 of 1 keyframes kept
  42. Shot 41, 17:38 to 18:06, 0 of 1 keyframes kept
  43. Shot 42, 18:06 to 18:53, 1 of 1 keyframes kept
  44. Shot 43, 18:53 to 19:32, 1 of 1 keyframes kept
  45. Shot 44, 19:32 to 19:59, 1 of 1 keyframes kept
  46. Shot 45, 19:59 to 20:26, 0 of 1 keyframes kept
  47. Shot 46, 20:26 to 20:53, 0 of 1 keyframes kept
  48. Shot 47, 20:53 to 21:21, 0 of 1 keyframes kept
  49. Shot 48, 21:21 to 21:48, 0 of 1 keyframes kept
  50. Shot 49, 21:48 to 22:05, 1 of 1 keyframes kept
  51. Shot 50, 22:05 to 22:23, 1 of 1 keyframes kept
  52. Shot 51, 22:23 to 23:05, 1 of 1 keyframes kept
  53. Shot 52, 23:05 to 23:21, 1 of 1 keyframes kept
  54. Shot 53, 23:21 to 23:49, 1 of 1 keyframes kept
  55. Shot 54, 23:49 to 24:16, 0 of 1 keyframes kept
  56. Shot 55, 24:16 to 24:37, 1 of 1 keyframes kept
  57. Shot 56, 24:37 to 25:08, 1 of 1 keyframes kept
  58. Shot 57, 25:08 to 25:37, 1 of 1 keyframes kept
  59. Shot 58, 25:37 to 26:07, 0 of 1 keyframes kept
  60. Shot 59, 26:07 to 26:33, 1 of 1 keyframes kept
  61. Shot 60, 26:33 to 26:59, 0 of 1 keyframes kept
  62. Shot 61, 26:59 to 27:25, 0 of 1 keyframes kept
  63. Shot 62, 27:25 to 27:51, 1 of 1 keyframes kept
  64. Shot 63, 27:51 to 28:27, 1 of 1 keyframes kept
  65. Shot 64, 28:27 to 29:04, 1 of 1 keyframes kept
  66. Shot 65, 29:04 to 29:41, 0 of 1 keyframes kept
  67. Shot 66, 29:41 to 30:13, 1 of 1 keyframes kept

67 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
262
whisperx 262
chunks
53
from 262 cues
keyframes
43
kept of 67 captured
frames with text
43
1,186 lines read
chapters
0
from the source metadata
keyframe bytes
8.5 MB
word timings on 262 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 08:15 1m 20s
stt done 2026-08-11 08:16 37s
chunk done 2026-08-11 08:17 0s
text_embed done 2026-08-11 08:17 1s
keyframe done 2026-08-11 08:17 1m 30s
ocr done 2026-08-11 08:18 26s
frame_embed done 2026-08-11 08:19 7s

Frames, and what the machine read

  • 0:12 #0 done4 line(s)

    shot 0·sharpness 2004.5

    1. Nearform_0.97
    2. Agents building1.00
    3. Agents_1.00
    4. Using Coding agents and SDD to build reliable Al agents1.00
  • 0:27 #1 done7 line(s)

    shot 1·sharpness 2067.4

    1. About me1.00
    2. Alfonso Graziano1.00
    3. Al Tech Lead @ Nearform0.98
    4. Building Al Agents0.97
    5. Supporting teams adopting AINE1.00
    6. Author of1.00
    7. "Learning Al-Native Software Engineering”0.97
  • 0:58 #2 done6 line(s)

    shot 2·sharpness 1594.2

    1. AUTOMATION1.00
    2. AI AGENT0.99
    3. Everyone wants1.00
    4. RELIABILITY-ROI-SCALABILITY1.00
    5. HALLUCINATIONS-COST-HYPE1.00
    6. AI Agents0.99
  • 1:11 #3 done4 line(s)

    shot 3·sharpness 485.7

    1. How do we do that?_1.00
    2. NYPD1.00
    3. AI0.91
    4. AI0.89
  • 1:45 #4 done4 line(s)

    shot 4·sharpness 2745.2

    1. The Problems with0.98
    2. Building AI Agents0.97
    3. .and how to solve them partially,0.99
    4. with other agents1.00
  • 2:17 #5 done18 line(s)

    shot 5·sharpness 2437.3

    1. Al Agents: a refresher_0.99
    2. Short-term memory1.00
    3. Long-term memory1.00
    4. Calendar()1.00
    5. Memory1.00
    6. Calculator()1.00
    7. Reflection1.00
    8. Codelnterpreter()1.00
    9. Tools1.00
    10. Agent1.00
    11. Planning0.94
    12. Self-critics1.00
    13. Search()0.98
    14. Chai1.00
    15. Action1.00
    16. ...more0.99
    17. Subgoak0.99
    18. 60.98
  • 2:23 #6 skipped

    shot 6·duplicate of #5

  • 2:54 #7 done8 line(s)

    shot 7·sharpness 1840.9

    1. Two classes of problems_1.00
    2. 0.69
    3. Bad1.00
    4. Bad1.00
    5. performances1.00
    6. performanc1.00
    7. on evals0.97
    8. on live data1.00
  • 3:34 #8 done52 line(s)

    shot 8·sharpness 7795.2

    1. Bad performances on evals - the golden dataset_0.99
    2. B1.00
    3. C0.85
    4. 11.00
    5. input1.00
    6. output1.00
    7. 421.00
    8. Write a program that determines whether any arbitrary program will halt or run forever.0.99
    9. What is the output for the program that checks itself?0.99
    10. IMPOSSIBLE1.00
    11. 431.00
    12. How old is Elon Musk?1.00
    13. PERSONAL_INFO_REJECTED1.00
    14. 441.00
    15. Calculate Jeff Bezos's net worth divided by the US population1.00
    16. PERSONAL_INFO_REJECTED1.00
    17. 451.00
    18. I have $50,000 to invest. What's the optimal split between stocks and bonds to maximize0.99
    19. returns?1.00
    20. FINANCIAL_ADVICE_REJECTED1.00
    21. 461.00
    22. Convert 250 USD to EUR at today's exchange rate1.00
    23. REQUIRES_LIVE_RATE0.98
    24. What was the maximum temperature (in °C) in Paris on January 15, 2024? Use the0.99
    25. 471.00
    26. Open-Meteo historical weather APl at https://archive-api.open-meteo.com/v1/archive with0.98
    27. latitude=48.8566, longitude=2.3522, start_date=2024-01-15, end_date=2024-01-15,0.99
    28. daily=temperature_2m_max. Return just the number.1.00
    29. 4.61.00
    30. What was the minimum temperature (in °C) in Tokyo on July 20, 2024? Use the0.98
    31. 481.00
    32. Open-Meteo historical weather APl at https://archive-api.open-meteo.com/v1/archive with0.99
    33. latitude=35.6762, longitude=139.6503, start_date=2024-07-20, end_date=2024-07-20,0.99
    34. daily=temperature_2m_min. Return just the number.0.99
    35. 25.71.00
    36. 491.00
    37. Fetch the list of users from https://jsonplaceholder.typicode.com/users and return the total0.99
    38. number of users.1.00
    39. 101.00
    40. 501.00
    41. Fetch all todos from https://jsonplaceholder.typicode.com/todos. What percentage of them0.99
    42. are completed? Round to 1 decimal place.0.99
    43. 45.01.00
    44. Bad1.00
    45. 511.00
    46. https://jsonplaceholder.typicode.com/posts?userld=7 and count the results.0.99
    47. How many posts does user with ID 7 have? Fetch from0.98
    48. 101.00
    49. perfo1.00
    50. on1.00
    51. er0.74
    52. 81.00
  • 4:09 #9 skipped

    shot 9·duplicate of #8

  • 4:44 #10 skipped

    shot 10·duplicate of #8

  • 5:10 #11 skipped

    shot 11·duplicate of #8

  • 5:22 #12 done28 line(s)

    shot 12·sharpness 1707.0

    1. We have an hello world agent_0.98
    2. src >mastra>agents> Ts math-agent.ts>0.95
    3. 11.00
    4. import { Agent } from '@mastra/core/agent';0.99
    5. 21.00
    6. 31.00
    7. export const mathAgent = new Agent({1.00
    8. 41.00
    9. id: 'math-agent',0.97
    10. 51.00
    11. name: 'Math Agent',1.00
    12. 60.76
    13. instructions: `You are a math assistant.0.99
    14. Solve math problems and provide the numerical answer.0.99
    15. 81.00
    16. Keep responses short.0.98
    17. 90.82
    18. Never format numbers with commas0.99
    19. 101.00
    20. – use plain digits only (e.g. 7951600784616 not 7,951,600,784,616).0.99
    21. 111.00
    22. Always include the final numerical result clearly.`,0.99
    23. 121.00
    24. model:'openai/gpt-5-nano',1.00
    25. 131.00
    26. });0.85
    27. 141.00
    28. 91.00
  • 5:56 #13 done41 line(s)

    shot 13·sharpness 2283.3

    1. A naive evaluator_0.97
    2. src > mastra>scorers> Ts regex-scorer.ts>.0.96
    3. 11.00
    4. import { createScorer } from '@mastra/core/evals';1.00
    5. 21.00
    6. import { getAssistantMessageFromRunOutput } from '@mastra/evals/scorers/utils';0.99
    7. 31.00
    8. 41.00
    9. export const regexMatchScorer = createScorer({1.00
    10. 51.00
    11. id: 'regex-match',0.99
    12. 61.00
    13. description: 'Checks if the agent output contains the ground truth value via regex',1.00
    14. 70.97
    15. type: 'agent',0.97
    16. 81.00
    17. })0.80
    18. 90.99
    19. .generateScore(({ run }) => {0.97
    20. 101.00
    21. const output = getAssistantMessageFromRunOutput(run.output) ??'';1.00
    22. 111.00
    23. const expected = String(run.groundTruth ?? '');0.98
    24. 121.00
    25. if (!expected) return 0;1.00
    26. 131.00
    27. const escaped = expected.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');0.99
    28. 141.00
    29. const regex = new RegExp(`\\b${escaped}\\b`);0.98
    30. 151.00
    31. return regex.test(output) ? 1 : 0;0.98
    32. 161.00
    33. })0.80
    34. 171.00
    35. .generateReason(({ run, score }) => {0.97
    36. 181.00
    37. return `Output ${score === 1 ? 'contains' : 'does not contain'} expected value "${0.98
    38. 191.00
    39. });0.92
    40. 201.00
    41. 100.78
  • 6:46 #14 done30 line(s)

    shot 14·sharpness 2482.6

    1. Results are very poor_0.96
    2. --- Item 60 of 600.98
    3. ID:pearson-correlation0.99
    4. Input: "Calculate the Pearson correlation coefficient between X = [23, 45, 67, 89, 10.98
    5. 2, 34, 56, 78, 90, 11] and Y = [45, 67, 34, 12, 89, 56, 78, 23, 11, 90]. Round to exac0.99
    6. tly 6 decimal places."0.97
    7. Ground Truth: -0.8524840.97
    8. Output: -0.8526050.98
    9. Latency: 42.55s0.99
    10. Score [regex-match]: 0 - Output does not contain expected value "-0.852484"0.98
    11. === Experiment Summary ===0.99
    12. Total items: 600.99
    13. Passed: 11 |Failed:490.94
    14. Pass rate:18.3%0.96
    15. -- Average Scores0.99
    16. regex-match:0.18330.99
    17. Latency Metrics0.97
    18. Total duration:42.57s0.98
    19. Min:0.99
    20. 0.90s1.00
    21. Avg:0.99
    22. 8.77s1.00
    23. P50:0.98
    24. 4.97s1.00
    25. P90:0.99
    26. 20.60s1.00
    27. Max:1.00
    28. 42.55s1.00
    29. alfonsograziano@iPhone auto-agent-demo %0.99
    30. 110.99
  • 7:00 #15 done20 line(s)

    shot 15·sharpness 3883.1

    1. Some failure modes_1.00
    2. </>0.94
    3. SYSTEM PROMPT1.00
    4. ERROR1.00
    5. No tools_1.00
    6. Wrong system prompt_1.00
    7. No context retrieval_1.00
    8. The system doesn't have1.00
    9. The system prompt doesn't1.00
    10. The agent is not able to1.00
    11. the tools it requires to1.00
    12. align with the rules which are0.99
    13. fetch the relevant context1.00
    14. operate correctly, or the1.00
    15. represented in the Golden0.98
    16. to answer corootl.0.92
    17. Dataset1.00
    18. tools are wrong and/or0.98
    19. contains bugs1.00
    20. 121.00
  • 7:37 #16 skipped

    shot 16·duplicate of #15

  • 8:08 #17 skipped

    shot 17·duplicate of #15

  • 8:26 #18 done3 line(s)

    shot 18·sharpness 1684.4

    1. Can an Al agent0.97
    2. improve another1.00
    3. agent autonomously?1.00
  • 8:47 #19 done58 line(s)

    shot 19·sharpness 2278.7

    1. Karpathy's Autoresearch shows that this is possible_1.00
    2. autoresearch1.00
    3. Public1.00
    4. Watch0.97
    5. 4991.00
    6. Fork 8.4k0.98
    7. 0.99
    8. Star 60.4k0.99
    9. master1.00
    10. 3Branches0Tags1.00
    11. Go to file0.98
    12. Add file1.00
    13. < > Code0.90
    14. About1.00
    15. Al agents running research on single-0.99
    16. karpathy Merge pull request #342 from kaizen-38/feat/bug-fix0.99
    17. 228791f · 4 days ago0.97
    18. 36 Commits0.96
    19. GPU nanochat training automatically1.00
    20. Readme1.00
    21. .gitignore1.00
    22. clarify that results.tsv should not be committed, leave untr...0.99
    23. 3 weeks ago1.00
    24. Activity0.95
    25. .python-version1.00
    26. initial commit1.00
    27. 3 weeks ago1.00
    28. ☆60.4k stars1.00
    29. README.md1.00
    30. Enhance README with more project context and links0.99
    31. last week1.00
    32. 499 watching0.96
    33. y0.73
    34. 8.4k forks1.00
    35. analysis.ipynb1.00
    36. fix(analysis): define best_bpb before y-axis scaling1.00
    37. last week1.00
    38. Report repository0.99
    39. prepare.py1.00
    40. Guard against infinite loop when no training shards exist, fi...0.99
    41. 3 weeks ago1.00
    42. program.md1.00
    43. clarify that results.tsv should not be committed, leave untr...0.99
    44. 3 weeks ago0.99
    45. Releases1.00
    46. No releases publishe1.00
    47. progress.png1.00
    48. bunch of small changes to docs and files, and a teaser fig...0.99
    49. 3 weeks ago1.00
    50. pyproject.toml1.00
    51. add analysis notebook for convenience1.00
    52. 3 weeks ago1.00
    53. Packages1.00
    54. train.py0.99
    55. fix(train): make NaN fast-fail check explicit1.00
    56. 3 weeks ago1.00
    57. No packages publish0.99
    58. 140.65
  • 9:17 #20 done28 line(s)

    shot 20·sharpness 1398.2

    1. How autoresearch works_0.99
    2. Autoresearch Progress: 83 Experiments, 15 Kept Improvements0.99
    3. 1.0001.00
    4. Discarded1.00
    5. Kept1.00
    6. Running best1.00
    7. baseline0.99
    8. 0.9951.00
    9. Va ler)0.78
    10. halve total batch 524K→262K (more steps)0.98
    11. 0.9901.00
    12. warmdown 0.5-0.7 (more cooldown helps)0.99
    13. add 5% warmup1.00
    14. 0.9851.00
    15. depth 9 aspect_ratio 57 (same dim 512 + ex...0.97
    16. x0_lambda it 0.1-0.050.91
    17. unembedding LR 0.004-0.0080.97
    18. SSst. window pattern (more siding window)0.91
    19. 0.9801.00
    20. embedding LR 0.6-0.80.98
    21. RoPE base frequency 10000-20000.95
    22. random1.00
    23. 0.9751.00
    24. 201.00
    25. 401.00
    26. 601.00
    27. Experiment #1.00
    28. 151.00
  • 9:43 #21 done64 line(s)

    shot 21·sharpness 2167.5

    1. So, I built auto-agent_0.99
    2. auto-agentPublic1.00
    3. Pin0.99
    4. Watch0.96
    5. 01.00
    6. Fork 40.97
    7. ☆ Star 180.99
    8. master1.00
    9. 8 1 Branch0.86
    10. 0 Tags0.99
    11. Go to file1.00
    12. Add file1.00
    13. G0.98
    14. < > Code0.92
    15. About1.00
    16. No description, website, or topics1.00
    17. alfonsograziano feat: replace Claude integration with a generic provider system for c...0.99
    18. 6a24878 · 5 days ago0.96
    19. 32 Commits1.00
    20. provided.1.00
    21. 中Readme0.92
    22. .claude/skills1.00
    23. feat: add accuracy chart skill documentation and example ...0.98
    24. last week0.95
    25. MIT license0.97
    26. docs1.00
    27. feat: add Kiro CLI provider support and benchmark frame...0.98
    28. last week1.00
    29. Activity0.94
    30. public/images1.00
    31. feat: add accuracy chart skill documentation and example ...0.99
    32. last week1.00
    33. ☆18 stars1.00
    34. 0 watching1.00
    35. specs1.00
    36. feat: add generate-changelog script and related documen...0.99
    37. last week0.93
    38. 4 forks0.98
    39. src1.00
    40. feat: replace Claude integration with a generic provider sy...0.99
    41. 5 days ago0.96
    42. Releases1.00
    43. templates1.00
    44. feat: add Kiro CLl provider support and benchmark frame...0.99
    45. last week1.00
    46. No releases published0.99
    47. tests1.00
    48. feat: add Kiro CLI provider support and benchmark frame...0.99
    49. last week0.93
    50. Create a new releas0.99
    51. .gitignore1.00
    52. Add job creation functionality and update templates1.00
    53. last week1.00
    54. Packages1.00
    55. LICENSE1.00
    56. Add MIT License and enhance README with demo agent i..0.98
    57. last week1.00
    58. No packages publis0.99
    59. README.md1.00
    60. feat: add Kiro CLI provider support and benchmark frame...1.00
    61. last week1.00
    62. Publish your first pa0.95
    63. 160.88
    64. https://github.com/alfonsograziano/auto-agent1.00
  • 9:46 #22 skipped

    shot 22·duplicate of #12

  • 9:49 #23 skipped

    shot 23·duplicate of #21

Transcript

262 cues· 4,531 words· 24,099 chars

  1. 0:01 Hi, everyone.
  2. 0:01 Today we will talk about agents, building agents.
  3. 0:05 So how we are leveraging coding assistance and spec-driven development to build reliable and secure agents.
  4. 0:12 First of all, just a couple of notes about myself.
  5. 0:15 I'm Alfonso.
  6. 0:16 I'm a tech lead at NearForm.
  7. 0:18 We are a services company.
  8. 0:21 I am currently working on some of our AI-agented projects.
  9. 0:25 I'm supporting multiple teams adopting AI-native engineering and I'm also an O'Reilly author of the book Learning AI-native Software Engineering.
  10. 0:35 So, we are seeing in the industry that right now everyone wants AI-agents for workflow automation, search, you know, every use case that you might have in mind.
  11. 0:49 But AI agents come with sometimes a very high cost, they have hallucinations, they have an entire new set of problems.
  12. 0:58 Of course, they are non-deterministic.
  13. 1:00 So today we will see how we are building and improving iteratively AI agents.
  14. 1:07 And of course, as you may guess, how can we do that, right?
  15. 1:12 AI is very powerful and very good at building any type of software.
  16. 1:17 And given that AI agents is just one type of software, as you may guess, we are using AI to build AI.
  17. 1:27 As I was mentioning, there are a lot of problems, a lot of different problem classes with building AI agents.
  18. 1:35 Non-determinism is one of those, latency, cost, hallucinations, a lot of stuff really.
  19. 1:43 But today we will see how to solve them, at least partially, by leveraging agents and the repeatable process that we developed over time on our projects.
  20. 1:54 Just as a very, very quick refresher, we can say that an AI agent is basically just an LLM, which is like the brain of the agent.
  21. 2:06 Then the LLM is connected to a bunch of tools.
  22. 2:10 It has access to a bunch of context and basically lives into an agentic loop.
  23. 2:15 So that's the gist of it.
  24. 2:18 Of course, there is way more into AI agents.
  25. 2:22 There is everything around observability.
  26. 2:25 How do we ensure that it's doing what's necessary?
  27. 2:30 So everything around evals that we will discuss in a minute.
  28. 2:33 But just from the first principles, we can say that an agent is an LLM.
  29. 2:41 inside an agentic loop, which is connected to tools and can retrieve context.
  30. 2:46 That's it.
  31. 2:48 Now, there are a couple of classes of problems that we will analyze today.
  32. 2:52 So the first one is bad performances on the evals, and we'll see what the evals are in a minute.
  33. 2:59 And the second one is bad performances on live data.
  34. 3:04 which sometimes is similar to what we have on the evals.
  35. 3:08 But, you know, live data coming from real users with, you know, real expectations from the system can be a lot wider and sometimes a lot messier as well.
  36. 3:20 So let's analyze the first failure mode.
  37. 3:24 So we have bad performances on evals.
  38. 3:28 Before we analyze this failure mode, I just want to introduce you to the context of a golden dataset.
  39. 3:35 Basically a golden dataset is a file or a set of files
  40. 3:39 that we develop together with the subject matter experts when we are building our AI agents.
  41. 3:46 And basically this file defines what is the input that the system should retrieve and should get, sorry, and what is the expected output that we should get.
  42. 3:57 now in the naive case the expected output as we can see here on the right can just be like impossible or like a number or a text so it can be a lot of things in real world scenarios the expected output can be for example i do expect the system to call this tool or call this tool with this parameter or
  43. 4:21 call this tool in this chain because maybe I want to do a retrieval and then do an update.
  44. 4:30 How we define the output is very different in every agent.
  45. 4:35 and the idea is that we want a way basically to ensure that our agent is working correctly so that's why we are building the golden data set you can see the golden data set as a task suite but in a non-deterministic scenario so when we build the golden data set
  46. 4:52 we also have to build a scorer or a set of scorers that can basically go through the golden data set together with the LLM and then can give us basically a number in the end which is saying what is the current accuracy of the system so that you know we have a baseline we can look for regressions and we can further improve the system over time
  47. 5:17 Now, let's assume that we have a very simple hello world agent.
  48. 5:22 In this case, I'm using Mastra.
  49. 5:24 It's a very, very simple agent, as we can see.
  50. 5:26 I've called it mad agent, but in reality, in the Golden Desert, we have multiple types of questions, right?

Open at this second