read-only demo

Videos CDqzWpwkSls

Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo

index_state ready data_status ok

AI Engineer· published 2026-07-07· 0:25:53· en-US· indexed 2026-08-11 05:00

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:15, 1 of 1 keyframes kept
  2. Shot 1, 0:15 to 0:16, 1 of 1 keyframes kept
  3. Shot 2, 0:16 to 0:51, 0 of 1 keyframes kept
  4. Shot 3, 0:51 to 1:16, 1 of 1 keyframes kept
  5. Shot 4, 1:16 to 1:41, 0 of 1 keyframes kept
  6. Shot 5, 1:41 to 2:06, 0 of 1 keyframes kept
  7. Shot 6, 2:06 to 2:34, 1 of 1 keyframes kept
  8. Shot 7, 2:34 to 3:02, 0 of 1 keyframes kept
  9. Shot 8, 3:02 to 3:30, 0 of 1 keyframes kept
  10. Shot 9, 3:30 to 3:44, 0 of 1 keyframes kept
  11. Shot 10, 3:44 to 4:11, 1 of 1 keyframes kept
  12. Shot 11, 4:11 to 4:39, 0 of 1 keyframes kept
  13. Shot 12, 4:39 to 5:05, 1 of 1 keyframes kept
  14. Shot 13, 5:05 to 5:32, 0 of 1 keyframes kept
  15. Shot 14, 5:32 to 5:57, 1 of 1 keyframes kept
  16. Shot 15, 5:57 to 6:23, 0 of 1 keyframes kept
  17. Shot 16, 6:23 to 6:49, 1 of 1 keyframes kept
  18. Shot 17, 6:49 to 7:16, 0 of 1 keyframes kept
  19. Shot 18, 7:16 to 7:43, 0 of 1 keyframes kept
  20. Shot 19, 7:43 to 8:17, 1 of 1 keyframes kept
  21. Shot 20, 8:17 to 8:50, 0 of 1 keyframes kept
  22. Shot 21, 8:50 to 8:53, 1 of 1 keyframes kept
  23. Shot 22, 8:53 to 9:12, 1 of 1 keyframes kept
  24. Shot 23, 9:12 to 9:26, 1 of 1 keyframes kept
  25. Shot 24, 9:26 to 9:59, 1 of 1 keyframes kept
  26. Shot 25, 9:59 to 10:27, 1 of 1 keyframes kept
  27. Shot 26, 10:27 to 10:56, 0 of 1 keyframes kept
  28. Shot 27, 10:56 to 11:26, 1 of 1 keyframes kept
  29. Shot 28, 11:26 to 11:55, 0 of 1 keyframes kept
  30. Shot 29, 11:55 to 11:58, 1 of 1 keyframes kept
  31. Shot 30, 11:58 to 12:26, 1 of 1 keyframes kept
  32. Shot 31, 12:26 to 12:54, 0 of 1 keyframes kept
  33. Shot 32, 12:54 to 13:21, 0 of 1 keyframes kept
  34. Shot 33, 13:21 to 13:31, 1 of 1 keyframes kept
  35. Shot 34, 13:31 to 14:03, 1 of 1 keyframes kept
  36. Shot 35, 14:03 to 14:36, 0 of 1 keyframes kept
  37. Shot 36, 14:36 to 15:08, 1 of 1 keyframes kept
  38. Shot 37, 15:08 to 15:40, 0 of 1 keyframes kept
  39. Shot 38, 15:40 to 16:06, 1 of 1 keyframes kept
  40. Shot 39, 16:06 to 16:32, 0 of 1 keyframes kept
  41. Shot 40, 16:32 to 16:58, 0 of 1 keyframes kept
  42. Shot 41, 16:58 to 17:23, 0 of 1 keyframes kept
  43. Shot 42, 17:23 to 17:49, 0 of 1 keyframes kept
  44. Shot 43, 17:49 to 18:19, 1 of 1 keyframes kept
  45. Shot 44, 18:19 to 18:50, 0 of 1 keyframes kept
  46. Shot 45, 18:50 to 18:53, 1 of 1 keyframes kept
  47. Shot 46, 18:53 to 19:20, 0 of 1 keyframes kept
  48. Shot 47, 19:20 to 19:47, 0 of 1 keyframes kept
  49. Shot 48, 19:47 to 20:14, 0 of 1 keyframes kept
  50. Shot 49, 20:14 to 20:52, 0 of 1 keyframes kept
  51. Shot 50, 20:52 to 21:29, 0 of 1 keyframes kept
  52. Shot 51, 21:29 to 21:58, 1 of 1 keyframes kept
  53. Shot 52, 21:58 to 22:27, 0 of 1 keyframes kept
  54. Shot 53, 22:27 to 22:56, 0 of 1 keyframes kept
  55. Shot 54, 22:56 to 23:38, 1 of 1 keyframes kept
  56. Shot 55, 23:38 to 24:09, 0 of 1 keyframes kept
  57. Shot 56, 24:09 to 24:39, 1 of 1 keyframes kept
  58. Shot 57, 24:39 to 25:10, 0 of 1 keyframes kept
  59. Shot 58, 25:10 to 25:40, 0 of 1 keyframes kept
  60. Shot 59, 25:40 to 25:50, 1 of 1 keyframes kept
  61. Shot 60, 25:50 to 25:52, 1 of 1 keyframes kept

61 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
223
whisperx 223
chunks
46
from 223 cues
keyframes
28
kept of 61 captured
frames with text
28
423 lines read
chapters
0
from the source metadata
keyframe bytes
5.8 MB
word timings on 223 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 04:57 1m 11s
stt done 2026-08-11 04:58 28s
chunk done 2026-08-11 04:59 0s
text_embed done 2026-08-11 04:59 1s
keyframe done 2026-08-11 04:59 1m 17s
ocr done 2026-08-11 05:00 12s
frame_embed done 2026-08-11 05:00 5s

Frames, and what the machine read

  • 0:09 #0 done6 line(s)

    shot 0·sharpness 598.2

    1. duolingo0.99
    2. AIE 20260.98
    3. the human in your loop0.98
    4. isn't thinking0.99
    5. build Al systems for discernment, not approval0.99
    6. angel ortmann lee1.00
  • 0:16 #1 done2 line(s)

    shot 1·sharpness 500.0

    1. duolingo1.00
    2. background1.00
  • 0:30 #2 skipped

    shot 2·duplicate of #1

  • 0:54 #3 done39 line(s)

    shot 3·sharpness 2119.2

    1. BACKGROUND1.00
    2. trusting technology1.00
    3. Al Overview0.99
    4. COVID-19 symptoms range from mild to severe and typically appear 20.99
    5. exposure. The most common signs include fever or chills, cough, fatigue, .0.98
    6. Many things that used to be manual1.00
    7. taste or smell. Centers for Disease Control and Preven..10.95
    8. muscle or body aches, sore throat, congestion or runny nose, and a new loss of0.99
    9. cognitive tasks are already offloaded to0.99
    10. technology daily0.99
    11. 5:050.99
    12. Add List0.98
    13. Edit0.88
    14. As Al becomes more integrated into daily0.99
    15. Lists1.00
    16. 1210.94
    17. 888 A8 Contacts0.80
    18. decrease1.00
    19. life, trust will increase and caution will0.99
    20. iCloud0.90
    21. ANiCioud0.53
    22. Famly0.93
    23. Golden Gate1.00
    24. Recreation1.00
    25. National1.00
    26. Area1.00
    27. Emeryville0.96
    28. Berkeley0.99
    29. Gmail0.91
    30. Oakland1.00
    31. At Cmal0.55
    32. Moscone West1.00
    33. 23 min0.99
    34. 22.8 km0.97
    35. South San0.94
    36. Francisco1.00
    37. Pacifica1.00
    38. San Francisco0.99
    39. International Airport1.00
  • 1:33 #4 skipped

    shot 4·duplicate of #3

  • 2:03 #5 skipped

    shot 5·duplicate of #3

  • 2:31 #6 done15 line(s)

    shot 6·sharpness 2334.4

    1. BACKGROUND1.00
    2. cognitivesurrender1.00
    3. when humans forgo deliberation and adopt Al output as0.99
    4. their own with minimal scrutiny(Shaw & Nave, Wharton, 2026)0.99
    5. Al is reshaping human reasoning, often overriding human instinct1.00
    6. 80%0.87
    7. and deliberate reasoning0.98
    8. Al can supplement or supplant a human's thinking without the0.99
    9. accepted Al answers1.00
    10. human ever realizing it1.00
    11. even when Al was0.99
    12. wrong1.00
    13. Given Al access on an exam, participants scored0.99
    14. +25% when Al was right0.98
    15. -15% when Al was wrong1.00
  • 2:48 #7 skipped

    shot 7·duplicate of #6

  • 3:24 #8 skipped

    shot 8·duplicate of #6

  • 3:32 #9 skipped

    shot 9·duplicate of #0

  • 4:05 #10 done10 line(s)

    shot 10·sharpness 2135.5

    1. EXPERIMENT1.00
    2. whatis thedet?0.98
    3. The Duolingo English Test (DET) is a high-stakes fully-online English proficiency1.00
    4. exam accepted by over 6,000 programs worldwide.0.99
    5. To ensure score legitimacy, integrity, and trust for institutions, we have1.00
    6. many layers of security, including:1.00
    7. identity verification0.99
    8. locked-down testing environment1.00
    9. Al-assisted monitoring1.00
    10. human proctor review1.00
  • 4:31 #11 skipped

    shot 11·duplicate of #10

  • 5:02 #12 done11 line(s)

    shot 12·sharpness 1187.4

    1. duolingo english test0.98
    2. Copy-typing: reproducing text from another source1.00
    3. detecting1.00
    4. instead of composing it independently.1.00
    5. copytyping1.00
    6. We detect copy-typing using a CNN-Transformer0.99
    7. model that analyzes keystroke patterns to0.99
    8. distinguish transcription from composition1.00
    9. When flagged for copy-typing, proctors must verify1.00
    10. whether the test taker is likely copy-typing or not1.00
    11. Conservative model threshold (~1% FPR)1.00
  • 5:19 #13 skipped

    shot 13·duplicate of #12

  • 5:52 #14 done24 line(s)

    shot 14·sharpness 1794.9

    1. EXPERIMENT1.00
    2. experiment1.00
    3. THE QUESTION1.00
    4. Would skilled reviewers catch a false alarm or rubber-stamp it?1.00
    5. THE SETUP1.00
    6. Add fake Al signals and surface them to human reviewers as part of their normal1.00
    7. workflow1.00
    8. Info1.00
    9. Alerts1.00
    10. History1.00
    11. Flags & Summary0.99
    12. Behaviors1.00
    13. Decertification1.00
    14. Cheating Ring1.00
    15. ×0.68
    16. General1.00
    17. Flag1.00
    18. Time Description0.99
    19. Decision1.00
    20. Signals0.99
    21. Timeline1.00
    22. S4131.00
    23. 64:22Check if copy-typing behaviors are present1.00
    24. OYONO?0.98
  • 6:05 #15 skipped

    shot 15·duplicate of #14

  • 6:26 #16 done13 line(s)

    shot 16·sharpness 1939.7

    1. EXPERIMENT1.00
    2. initial findings0.98
    3. Despite a conservative model and human review, reviewers1.00
    4. 50%0.84
    5. endorsed fake alerts at near coin-flip rates suggesting evidence1.00
    6. of automation bias0.99
    7. accepted fake signals1.00
    8. But this unacceptable when college admissions and visas are1.00
    9. ≈coin-fliprate1.00
    10. on the line1.00
    11. The problem...0.99
    12. del1.00
    13. interface1.00
  • 7:03 #17 skipped

    shot 17·duplicate of #16

  • 7:30 #18 skipped

    shot 18·duplicate of #16

  • 8:03 #19 done12 line(s)

    shot 19·sharpness 1862.7

    1. EXPERIMENT1.00
    2. solution1.00
    3. To address these signs of automation bias, we0.99
    4. targeted the human-Al interaction loop.1.00
    5. 71%0.99
    6. Updated proctoring guidelines to emphasize:0.99
    7. 50%0.97
    8. Al signal = preliminary alert0.99
    9. Must find independent evidence before0.99
    10. upholding a flag1.00
    11. Old Instructions1.00
    12. New Instructions0.99
  • 8:30 #20 skipped

    shot 20·duplicate of #19

  • 8:51 #21 done2 line(s)

    shot 21·sharpness 444.5

    1. duolingo1.00
    2. implications1.00
  • 9:05 #22 done3 line(s)

    shot 22·sharpness 1064.0

    1. IMPLICATIONS1.00
    2. what does this mean for Al devs?0.97
    3. Your interaction loop determines how effective your Al is and what you learn from it.1.00
  • 9:22 #23 done9 line(s)

    shot 23·sharpness 1446.0

    1. IMPLICATIONS1.00
    2. make the decision data a param0.96
    3. Model1.00
    4. 0.96
    5. Human1.00
    6. 0.99
    7. Decision1.00
    8. Earlier, we defined human-in-the-loop Al as a system where a human makes a0.99
    9. decision. But the process is not linear, it is cyclical.1.00

Transcript

223 cues· 4,269 words· 24,237 chars

  1. 0:00 Hello, my name is Angel Ortman-Lee and I'm a software engineer at Duolingo.
  2. 0:05 I work on security for the Duolingo English test.
  3. 0:08 Today's talk is the human in your loop isn't thinking.
  4. 0:11 Build AI systems for discernment, not approval.
  5. 0:15 So let's get started with some background.
  6. 0:18 What is Human-in-the-Loop AI?
  7. 0:20 Human-in-the-Loop AI is a framework where a system or process is actively involving a human who is participating in the operation, supervision, or decision-making for an automated system.
  8. 0:30 Humans are involved because they're able to ensure the accuracy, safety, or ethical decision-making for that system.
  9. 0:38 They are a piece of that decision-making where AI normally falls short.
  10. 0:43 You can kind of think of this as a linear process where a model provides some sort of output and a human sees that and makes the decision.
  11. 0:52 Now, in the age of AI, trust is a big part of the conversation.
  12. 0:56 Recently, as technology has developed, a lot of things that used to be manual cognitive tasks are now part of technology.
  13. 1:03 So no one really memorizes phone numbers anymore.
  14. 1:05 They're just sitting in our contacts list on our smartphones.
  15. 1:09 When you're driving over somewhere, you just put the directions into your GPS and you don't really think about the details of that route.
  16. 1:15 Instead, you just trust the GPS to come up with the optimal way for you.
  17. 1:19 When you're searching something in the search engine, you used to look at the top results.
  18. 1:23 And now, as AI is becoming more increasingly integrated into our day-to-day life, you might be seeing some AI summaries at the top instead.
  19. 1:32 Here's an example of Googling COVID-19 symptoms, and you see that Gemini summary at the top.
  20. 1:38 And instead of looking at the CDC website or the WHO website, you might actually look to that as your final summary.
  21. 1:46 answer to the question that you are asking to the search engine.
  22. 1:49 As these pieces of AI are becoming more integrated into your day-to-day life, and these atomic little things like searching for an answer or getting a result from an app, your trust of AI systems will increase and caution will decrease.
  23. 2:04 This is something that's happening across all of society.
  24. 2:07 A study at Wharton was looking exactly that.
  25. 2:10 How is AI reshaping human reasoning, and what does it mean for us as we continue to use it day to day?
  26. 2:17 They've observed an interesting phenomenon called cognitive surrender, when a human forgoes deliberation and adopts AI output as their own with minimal scrutiny.
  27. 2:27 They saw this in a study where humans were asked to look at reasoning exams, and they were given AI resources to answer those questions.
  28. 2:35 They realized that AI can either supplement or supplant a human's thinking, and the human might not even know it's happening.
  29. 2:42 For example, for questions where the AI was right, the human performance increased by 25 percentage points, whereas when the AI was wrong, it decreased by 15.
  30. 2:53 This suggests that the humans were taking in that AI cognition and not really thinking critically about what that answer is, and instead using those answers as part of their own reasoning together with their human instinct and deliberate reasoning, amplifying those results.
  31. 3:13 Most interestingly, they saw that 80% of participants were accepting those AI answers even when it was wrong.
  32. 3:20 So they were lowering that barrier to entry and just trusting the AI even if they weren't really critically examining the correctness of that result.
  33. 3:31 So at the Duolingo English test, we wanted to look at the same thing, and we published some research titled When Machines Mislead.
  34. 3:39 It was a case study on the English exam, specifically in AI human interaction.
  35. 3:45 So for context, what is the DET?
  36. 3:48 The DET is a high-stakes exam that measures your English proficiency, and it's fully online and remotely proctored.
  37. 3:55 You can take it on your own machine in the comfort of your home, and it still provides high quality results that 6,000 programs worldwide are trusting day to day.
  38. 4:06 As you can imagine, with a fully online assessment, there are some interesting things that from a technological standpoint that have to happen to ensure good quality security.
  39. 4:16 This includes identity verification, a lockdown testing environment, as well as a variety of ways in which AI-assisted monitoring is happening to predict different types of cheating.
  40. 4:27 Lastly, we have a final round of scrutiny where we have human proctors review all of the video footage of the exam, as well as all of those AI results.
  41. 4:40 Specifically for the study, we wanted to target one of our AI cheating detection systems, which is copy typing.
  42. 4:48 Copy typing is a form of cheating where you're writing down information that you're reading currently as opposed to something that you're thinking right now at this moment.
  43. 4:57 So as you can imagine, your typing patterns differ when you're transcribing versus composing.
  44. 5:03 And this is something that our custom model is measuring.
  45. 5:06 It's looking at anomalies and keystroke patterns, and it's flagging them for sessions that are unusual.
  46. 5:13 So our model is highly conservative, and so we're prioritizing fairness for our test takers.
  47. 5:20 So this is not a very common flag.
  48. 5:23 However, when human proctors are taking a look at these, they're well trained to properly examine those segments and see if this is a flag.
  49. 5:34 So for our experiment, we wanted to answer the question, would a skilled reviewer catch a false alarm or would they just rubber stamp it?
  50. 5:41 So we know that our proctors are highly accurate at detecting various forms of cheating.

Open at this second