read-only demo

Videos EUsPvBeIx70

Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra

index_state ready data_status ok

AI Engineer· published 2026-07-12· 0:16:24· en-US· indexed 2026-08-11 03:22

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:25, 1 of 1 keyframes kept
  2. Shot 1, 0:25 to 0:51, 1 of 1 keyframes kept
  3. Shot 2, 0:51 to 1:16, 1 of 1 keyframes kept
  4. Shot 3, 1:16 to 1:41, 0 of 1 keyframes kept
  5. Shot 4, 1:41 to 2:06, 1 of 1 keyframes kept
  6. Shot 5, 2:06 to 2:33, 1 of 1 keyframes kept
  7. Shot 6, 2:33 to 3:00, 0 of 1 keyframes kept
  8. Shot 7, 3:00 to 3:27, 0 of 1 keyframes kept
  9. Shot 8, 3:27 to 3:58, 1 of 1 keyframes kept
  10. Shot 9, 3:58 to 4:30, 0 of 1 keyframes kept
  11. Shot 10, 4:30 to 5:16, 0 of 1 keyframes kept
  12. Shot 11, 5:16 to 5:41, 1 of 1 keyframes kept
  13. Shot 12, 5:41 to 6:06, 0 of 1 keyframes kept
  14. Shot 13, 6:06 to 6:31, 0 of 1 keyframes kept
  15. Shot 14, 6:31 to 6:56, 0 of 1 keyframes kept
  16. Shot 15, 6:56 to 7:21, 1 of 1 keyframes kept
  17. Shot 16, 7:21 to 7:46, 0 of 1 keyframes kept
  18. Shot 17, 7:46 to 8:17, 1 of 1 keyframes kept
  19. Shot 18, 8:17 to 8:48, 0 of 1 keyframes kept
  20. Shot 19, 8:48 to 9:20, 1 of 1 keyframes kept
  21. Shot 20, 9:20 to 9:53, 0 of 1 keyframes kept
  22. Shot 21, 9:53 to 10:39, 0 of 1 keyframes kept
  23. Shot 22, 10:39 to 11:16, 1 of 1 keyframes kept
  24. Shot 23, 11:16 to 11:53, 0 of 1 keyframes kept
  25. Shot 24, 11:53 to 12:19, 1 of 1 keyframes kept
  26. Shot 25, 12:19 to 12:45, 0 of 1 keyframes kept
  27. Shot 26, 12:45 to 13:11, 1 of 1 keyframes kept
  28. Shot 27, 13:11 to 13:37, 0 of 1 keyframes kept
  29. Shot 28, 13:37 to 14:03, 1 of 1 keyframes kept
  30. Shot 29, 14:03 to 14:29, 0 of 1 keyframes kept
  31. Shot 30, 14:29 to 14:55, 0 of 1 keyframes kept
  32. Shot 31, 14:55 to 15:20, 1 of 1 keyframes kept
  33. Shot 32, 15:20 to 15:46, 0 of 1 keyframes kept
  34. Shot 33, 15:46 to 16:24, 1 of 1 keyframes kept

34 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
170
whisperx 170
chunks
30
from 170 cues
keyframes
16
kept of 34 captured
frames with text
16
251 lines read
chapters
0
from the source metadata
keyframe bytes
2.1 MB
word timings on 170 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 03:20 1m 23s
stt done 2026-08-11 03:21 18s
chunk done 2026-08-11 03:21 0s
text_embed done 2026-08-11 03:21 0s
keyframe done 2026-08-11 03:21 46s
ocr done 2026-08-11 03:22 7s
frame_embed done 2026-08-11 03:22 2s

Frames, and what the machine read

  • 0:07 #0 done9 line(s)

    shot 0·sharpness 983.0

    1. phaidra1.00
    2. PATENT PENDING0.99
    3. AI ENGINEER WORLD'S FAIR 20260.99
    4. Semantic Blindness: We gave an LLM1.00
    5. 500,000 sensors — and it got confused.0.98
    6. What broke, and the architecture that fixed it.1.00
    7. Vanč Levstik0.99
    8. Raahul Singh · Staff AI Research Engineer0.96
    9. Vanč Levstik · Senior Engineering Manager0.99
  • 0:35 #1 done9 line(s)

    shot 1·sharpness 983.2

    1. phaidra1.00
    2. PATENT PENDING0.99
    3. AI ENGINEER WORLD'S FAIR 20261.00
    4. Semantic Blindness: We gave an LLM1.00
    5. 500,000 sensors — and it got confused.0.98
    6. What broke, and the architecture that fixed it.1.00
    7. Raahul Singh1.00
    8. Raahul Singh · Staff Al Research Engineer0.98
    9. Vanč Levstik · Senior Engineering Manager0.99
  • 1:10 #2 done40 line(s)

    shot 2·sharpness 906.3

    1. THE PROBLEM1.00
    2. The naive build0.98
    3. A site is tens of thousands of heterogeneous components -0.99
    4. GPUs, chillers, pumps, valves, PDUs, switches, sensors.1.00
    5. EVERYEQUIPMENT NAME0.96
    6. rack-12 / gpu-030.99
    7. chiller-a / supply-temp1.00
    8. pump-07 / speed0.99
    9. valve-22 / position1.00
    10. switch-09 / port-40.97
    11. meter-44 / kw0.97
    12. cdu-2 / flow0.98
    13. tower-3 / fan0.97
    14. pdu-5 / leg-b0.91
    15. ahu-1 / return-temp0.99
    16. sensor-88 / humidity1.00
    17. busbar-4 / current1.00
    18. rack-12 / gpu-84 coil-2 / delta-p0.95
    19. damper-7 / pos0.96
    20. vfd-09 / hz0.86
    21. tank-1 / level0.97
    22. rack-31 / gpu-111.00
    23. filter-6 / delta-p0.98
    24. breaker-15 / state0.99
    25. zone-3 / temp0.99
    26. rack-31 / gpu-120.99
    27. pump-08 / speed0.96
    28. manifold-4 / flow0.98
    29. node-22 / inlet-tenp0.97
    30. pdu-6 / leg-a0.95
    31. tower-4 / fan0.99
    32. meter-45 / km0.95
    33. rack-77 / gpu-010.99
    34. exch-3 / approach0.97
    35. valve-23 / position0.99
    36. crac-2 / supply1.00
    37. Raahul Singh0.98
    38. LLM0.93
    39. → "find the matches." At demo scale it works beautifully. So we shipped it.0.97
    40. Phaidra ai0.95
  • 1:33 #3 skipped

    shot 3·duplicate of #2

  • 1:53 #4 done11 line(s)

    shot 4·sharpness 810.6

    1. THE PROBLEM0.99
    2. works.any() vs works.all()0.99
    3. Demo=works.any()1.00
    4. Product=works.all()1.00
    5. One situation has to work. Easy to reach.1.00
    6. Every situation, every time. The naive build0.99
    7. Looks finished.1.00
    8. fails quietly and inconsistently.1.00
    9. Raahul Singh1.00
    10. After Andrej Karpathy - "Software in the era of Al"0.96
    11. Phaidra al0.89
  • 2:19 #5 done12 line(s)

    shot 5·sharpness 1087.8

    1. DOO0.50
    2. THE PROBLEM0.99
    3. works.any() vs works.all()0.98
    4. Demo=works.any()1.00
    5. Product=works.all()1.00
    6. One situation has to work. Easy to reach.1.00
    7. Every situation, every time. The naive build0.99
    8. Looks finished.1.00
    9. fails quietly and inconsistently.1.00
    10. Raahul Singh0.97
    11. After Andrej Karpathy - "Software in the era of Al*0.96
    12. Phaidra.ai1.00
  • 2:52 #6 skipped

    shot 6·duplicate of #5

  • 3:03 #7 skipped

    shot 7·duplicate of #5

  • 3:49 #8 done13 line(s)

    shot 8·sharpness 759.3

    1. THEPROBLEM1.00
    2. The fan-out wall0.99
    3. shard5,800 names0.95
    4. ~460K1.00
    5. names1.00
    6. 92 parallel LLM calls0.98
    7. flat list · no structure0.96
    8. +86more·92 shardstotal0.97
    9. What broke: the model returned phantom equipment that doesn't exist - and silently dropped0.99
    10. Raahul Singh1.00
    11. real components that were right there.0.99
    12. Each shard is semantically blind — it loses the forest for the trees. Cost scales O(instances) → it breaks.0.98
    13. Phaidraa0.87
  • 4:14 #9 skipped

    shot 9·duplicate of #8

  • 4:44 #10 skipped

    shot 10·duplicate of #4

  • 5:38 #11 done20 line(s)

    shot 11·sharpness 1105.9

    1. INSIGHT 1·THE LINEARISER0.95
    2. A million nodes → ~50 lines0.96
    3. data_center1.00
    4. 1,089,0111.00
    5. tree nodes1.00
    6. →~1000:1→1.00
    7. building (10)0.95
    8. data_hall (200)0.99
    9. rack (6,400)0.99
    10. gpu (468,800)0.92
    11. A 64-GPU site and a 460K-GPU0.99
    12. gpu:1.00
    13. data_center > building > data_hall > rack > gpu0.98
    14. chiller: data_center > building > chiller_plant > chiller1.00
    15. site produce nearly the same summary.1.00
    16. switch: data_center > building > data_hall > leaf_switch1.00
    17. The LLM sees every path and the full0.99
    18. Raahul Singh0.98
    19. containment at once.0.98
    20. Phaidra ai0.93
  • 5:44 #12 skipped

    shot 12·duplicate of #11

  • 6:12 #13 skipped

    shot 13·duplicate of #11

  • 6:41 #14 skipped

    shot 14·duplicate of #11

  • 7:09 #15 done12 line(s)

    shot 15·sharpness 826.7

    1. INSIGHT 2·THE PLANNER0.98
    2. From searching to planning1.00
    3. "Show me the GPUs in data hall 11 that1.00
    4. "scope_under": "<data_hall subtree>".0.97
    5. "collect": "gpu",0.96
    6. are running hot.1.00
    7. "filter": "running_hot"0.99
    8. The LLM reads the 50-line summary - never a0.98
    9. list of names.0.99
    10. # a plan - a recipe, not a search0.99
    11. Raahul Singh1.00
    12. Phaidra ai0.94
  • 7:34 #16 skipped

    shot 16·duplicate of #15

  • 8:10 #17 done14 line(s)

    shot 17·sharpness 1003.1

    1. INSIGHT 3-EXECUTION0.96
    2. DeO0.56
    3. Pre-indexed trees·instant1.00
    4. nodes = index.by_type["gpu"]0.99
    5. # 0(1) lookup0.97
    6. scoped = subtree(data_hall_11)1.00
    7. # pointer walk0.98
    8. result = scoped filter(running_hot)0.99
    9. # set intersection1.00
    10. sub-ms1.00
    11. Pure set operations on in-memory trees –0.96
    12. no disk I/O, no further LLM calls.0.97
    13. Raahul Singh1.00
    14. Phaidra.al0.91
  • 8:41 #18 skipped

    shot 18·duplicate of #17

  • 9:01 #19 done16 line(s)

    shot 19·sharpness 1019.4

    1. INSIGHT4·SENTINELS0.98
    2. Messy naming, constant cost1.00
    3. The LLM sees the facility's naming rule - <convention>0.98
    4. if the LLM read the names0.99
    5. - never the names.0.99
    6. It writes a filter pattern - <filter>.0.97
    7. sentinel—flat0.98
    8. Deterministic code applies it to every candidate name.0.99
    9. SK0.92
    10. 100K1.00
    11. SM names0.95
    12. Raahul Singh0.98
    13. constant1.00
    14. 5K names or 5M, it never sees them.0.99
    15. The LLM cost doesn't move as the system grows -0.98
    16. Phaidra.ai0.88
  • 9:40 #20 skipped

    shot 20·duplicate of #19

  • 10:33 #21 skipped

    shot 21·duplicate of #19

  • 11:12 #22 done18 line(s)

    shot 22·sharpness 936.7

    1. THE IMPACT0.99
    2. Accuracy that holds at scale1.00
    3. Before vs after the change - same model, same data, same ground truth, 3 runs per case.0.99
    4. 100%1.00
    5. 100%1.00
    6. 100%1.00
    7. 80%1.00
    8. 100%0.99
    9. 25%1.00
    10. 31%1.00
    11. accuracy1.00
    12. Vanč Levstik1.00
    13. 64 GPUs0.96
    14. 9,216 GPUs1.00
    15. ~460K GPUs0.98
    16. Before - naive (dump names into context)0.96
    17. After - structured (plan + deterministic resolve)0.98
    18. Phaidra.ai0.91
  • 11:45 #23 skipped

    shot 23·duplicate of #22

Transcript

170 cues· 2,690 words· 14,576 chars

  1. 0:00 Welcome, everyone.
  2. 0:01 My name is Rahul Singh.
  3. 0:02 I'm a staff AI research engineer at Phaedra.
  4. 0:05 And I'm Vansh Listic.
  5. 0:06 I'm a senior engineering manager also at Phaedra.
  6. 0:09 And today, we want to talk about the time when we gave an LLM 500,000 sensor names and it got confused.
  7. 0:15 We call this problem semantic blindness.
  8. 0:18 At Phaedra, we build AI agents for AI factories.
  9. 0:22 This includes agents which allow our customers to talk about their data centers and explain to themselves and understand how the data centers are working and what problems they are facing on a day-to-day basis.
  10. 0:34 User queries can be anything from what chiller is running hot to analyze the distribution of temperatures across my data halls to is any of my GPUs facing any problems.
  11. 0:47 Now, from this variety of queries, you can see that these include talking about specific equipments, is Chiller 6 all right, to talking about groups of equipments, GPUs and Data Hall 1.
  12. 1:00 The industry has not really figured out a common name pattern yet, and every single customer can have their own things, from simple names like RACs, with GPUs, with data holes, which give you an exact idea of where the printings are, to something that is more difficult to comprehend, like CH3 something something six.
  13. 1:21 We've seen all kinds of names in the industry going forward.
  14. 1:24 Now, when you're building a demo system, this works because you're working at a small scale.
  15. 1:30 A simple LLM can look at all the names of your equipment and figure out what the user is talking about.
  16. 1:36 But this problem really becomes intractable as you go for scale.
  17. 1:41 For example, at one gigawatt scale factories, you will see 400,000 plus GPUs.
  18. 1:48 And to support those GPUs, you have power meters, you have chillers, you have other equipments.
  19. 1:53 LLM context windows are finite, and you will very quickly saturate them, and it just becomes a problem.
  20. 1:59 Like we say, a product is something that works for all scenarios and does not fail silently.
  21. 2:03 A demo just has to work for one.
  22. 2:07 In addition to having the LLM figure out these names, we could also have embedded them in a RAC database, a vector embedding approach.
  23. 2:16 But the problem is, oftentimes, the names are so similar that semantic search just fails.
  24. 2:21 There's very little difference between a string name of 20 characters long, which is differed by, let's say, one character, chiller 6 instead of chiller 7.
  25. 2:31 or cdu something versus well something else it's very small so you get a lot of problems with getting accurate recall also llms suffer from what we call a frequency penalty if you keep on outputting very similar names over and over again or very similar tokens more accurately over and over again
  26. 2:53 there are internal penalties in the lm which just shut off their output so if the user says list all the names of uh let's say gpus in ism and there are let's say a hundred just by listing through them the lms uh gadgets would see uh think that well it's spiraling
  27. 3:10 and it would just shut the system up.
  28. 3:13 So we can't really have these two approaches.
  29. 3:15 RAG would not work.
  30. 3:16 Just naive LLMs would not work.
  31. 3:19 As we move into production systems, we need something that can scale as the system scales.
  32. 3:24 Now, there are naive solutions which we shall discuss going forward.
  33. 3:27 A naive approach to solve this problem would be just to, well, divide and conquer.
  34. 3:32 Take your different equipments, branch them in different shards, and pass them through LLMs.
  35. 3:38 Parallel calls should work, right?
  36. 3:40 Well, that's what we thought.
  37. 3:42 The problem is you get horrible recall and hallucinations.
  38. 3:46 You will see LLMs invent phantom equipment that do not exist and also silently drop things that do exist.
  39. 3:53 Now, this creates a problem for machine critical systems where the users need to know exactly what is happening with their systems.
  40. 4:00 Any problems, they will quickly erode user trust.
  41. 4:02 At the same time, you will miss on specific problems, which can cascade into bigger and bigger problems going forward.
  42. 4:10 Anywho, something that we realized here is that as the size of these physical infrastructure grows, the LLM-based solutions cannot grow with the size of individual components or instances or nodes.
  43. 4:25 We have to find something that grows sub-linearly with increasing equipment count.
  44. 4:30 And this is what we figured out.
  45. 4:32 So we should not grow with instances.
  46. 4:34 We should grow with tree depth.
  47. 4:36 Now, what do we mean by tree depth here?
  48. 4:38 We realized that there's a hierarchical structure in which an AI factory is arranged.
  49. 4:42 You will have data centers.
  50. 4:43 Then each data centers will have different data halls.

Open at this second