Videos CDqzWpwkSls
Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo
Scene timeline
61 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 223
- whisperx 223
- chunks
- 46
- from 223 cues
- keyframes
- 28
- kept of 61 captured
- frames with text
- 28
- 423 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.8 MB
- word timings on 223 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 04:57 | 1m 11s |
stt |
done | — | 2026-08-11 04:58 | 28s |
chunk |
done | — | 2026-08-11 04:59 | 0s |
text_embed |
done | — | 2026-08-11 04:59 | 1s |
keyframe |
done | — | 2026-08-11 04:59 | 1m 17s |
ocr |
done | — | 2026-08-11 05:00 | 12s |
frame_embed |
done | — | 2026-08-11 05:00 | 5s |
Frames, and what the machine read
-
- duolingo0.99
- AIE 20260.98
- the human in your loop0.98
- isn't thinking0.99
- build Al systems for discernment, not approval0.99
- angel ortmann lee1.00
-
- duolingo1.00
- background1.00
-
- BACKGROUND1.00
- trusting technology1.00
- Al Overview0.99
- COVID-19 symptoms range from mild to severe and typically appear 20.99
- exposure. The most common signs include fever or chills, cough, fatigue, .0.98
- Many things that used to be manual1.00
- taste or smell. Centers for Disease Control and Preven..10.95
- muscle or body aches, sore throat, congestion or runny nose, and a new loss of0.99
- cognitive tasks are already offloaded to0.99
- technology daily0.99
- 5:050.99
- Add List0.98
- Edit0.88
- As Al becomes more integrated into daily0.99
- Lists1.00
- 1210.94
- 888 A8 Contacts0.80
- decrease1.00
- life, trust will increase and caution will0.99
- iCloud0.90
- ANiCioud0.53
- Famly0.93
- Golden Gate1.00
- Recreation1.00
- National1.00
- Area1.00
- Emeryville0.96
- Berkeley0.99
- Gmail0.91
- Oakland1.00
- At Cmal0.55
- Moscone West1.00
- 23 min0.99
- 22.8 km0.97
- South San0.94
- Francisco1.00
- Pacifica1.00
- San Francisco0.99
- International Airport1.00
-
- BACKGROUND1.00
- cognitivesurrender1.00
- when humans forgo deliberation and adopt Al output as0.99
- their own with minimal scrutiny(Shaw & Nave, Wharton, 2026)0.99
- Al is reshaping human reasoning, often overriding human instinct1.00
- 80%0.87
- and deliberate reasoning0.98
- Al can supplement or supplant a human's thinking without the0.99
- accepted Al answers1.00
- human ever realizing it1.00
- even when Al was0.99
- wrong1.00
- Given Al access on an exam, participants scored0.99
- +25% when Al was right0.98
- -15% when Al was wrong1.00
-
- EXPERIMENT1.00
- whatis thedet?0.98
- The Duolingo English Test (DET) is a high-stakes fully-online English proficiency1.00
- exam accepted by over 6,000 programs worldwide.0.99
- To ensure score legitimacy, integrity, and trust for institutions, we have1.00
- many layers of security, including:1.00
- identity verification0.99
- locked-down testing environment1.00
- Al-assisted monitoring1.00
- human proctor review1.00
-
- duolingo english test0.98
- Copy-typing: reproducing text from another source1.00
- detecting1.00
- instead of composing it independently.1.00
- copytyping1.00
- We detect copy-typing using a CNN-Transformer0.99
- model that analyzes keystroke patterns to0.99
- distinguish transcription from composition1.00
- When flagged for copy-typing, proctors must verify1.00
- whether the test taker is likely copy-typing or not1.00
- Conservative model threshold (~1% FPR)1.00
-
- EXPERIMENT1.00
- experiment1.00
- THE QUESTION1.00
- Would skilled reviewers catch a false alarm or rubber-stamp it?1.00
- THE SETUP1.00
- Add fake Al signals and surface them to human reviewers as part of their normal1.00
- workflow1.00
- Info1.00
- Alerts1.00
- History1.00
- Flags & Summary0.99
- Behaviors1.00
- Decertification1.00
- Cheating Ring1.00
- ×0.68
- General1.00
- Flag1.00
- Time Description0.99
- Decision1.00
- Signals0.99
- Timeline1.00
- S4131.00
- 64:22Check if copy-typing behaviors are present1.00
- OYONO?0.98
-
- EXPERIMENT1.00
- initial findings0.98
- Despite a conservative model and human review, reviewers1.00
- 50%0.84
- endorsed fake alerts at near coin-flip rates suggesting evidence1.00
- of automation bias0.99
- accepted fake signals1.00
- But this unacceptable when college admissions and visas are1.00
- ≈coin-fliprate1.00
- on the line1.00
- The problem...0.99
- del1.00
- interface1.00
-
- EXPERIMENT1.00
- solution1.00
- To address these signs of automation bias, we0.99
- targeted the human-Al interaction loop.1.00
- 71%0.99
- Updated proctoring guidelines to emphasize:0.99
- 50%0.97
- Al signal = preliminary alert0.99
- Must find independent evidence before0.99
- upholding a flag1.00
- Old Instructions1.00
- New Instructions0.99
-
- duolingo1.00
- implications1.00
-
- IMPLICATIONS1.00
- what does this mean for Al devs?0.97
- Your interaction loop determines how effective your Al is and what you learn from it.1.00
-
- IMPLICATIONS1.00
- make the decision data a param0.96
- Model1.00
- →0.96
- Human1.00
- →0.99
- Decision1.00
- Earlier, we defined human-in-the-loop Al as a system where a human makes a0.99
- decision. But the process is not linear, it is cyclical.1.00
Transcript
223 cues· 4,269 words· 24,237 chars
- 0:00 Hello, my name is Angel Ortman-Lee and I'm a software engineer at Duolingo.
- 0:05 I work on security for the Duolingo English test.
- 0:08 Today's talk is the human in your loop isn't thinking.
- 0:11 Build AI systems for discernment, not approval.
- 0:15 So let's get started with some background.
- 0:18 What is Human-in-the-Loop AI?
- 0:20 Human-in-the-Loop AI is a framework where a system or process is actively involving a human who is participating in the operation, supervision, or decision-making for an automated system.
- 0:30 Humans are involved because they're able to ensure the accuracy, safety, or ethical decision-making for that system.
- 0:38 They are a piece of that decision-making where AI normally falls short.
- 0:43 You can kind of think of this as a linear process where a model provides some sort of output and a human sees that and makes the decision.
- 0:52 Now, in the age of AI, trust is a big part of the conversation.
- 0:56 Recently, as technology has developed, a lot of things that used to be manual cognitive tasks are now part of technology.
- 1:03 So no one really memorizes phone numbers anymore.
- 1:05 They're just sitting in our contacts list on our smartphones.
- 1:09 When you're driving over somewhere, you just put the directions into your GPS and you don't really think about the details of that route.
- 1:15 Instead, you just trust the GPS to come up with the optimal way for you.
- 1:19 When you're searching something in the search engine, you used to look at the top results.
- 1:23 And now, as AI is becoming more increasingly integrated into our day-to-day life, you might be seeing some AI summaries at the top instead.
- 1:32 Here's an example of Googling COVID-19 symptoms, and you see that Gemini summary at the top.
- 1:38 And instead of looking at the CDC website or the WHO website, you might actually look to that as your final summary.
- 1:46 answer to the question that you are asking to the search engine.
- 1:49 As these pieces of AI are becoming more integrated into your day-to-day life, and these atomic little things like searching for an answer or getting a result from an app, your trust of AI systems will increase and caution will decrease.
- 2:04 This is something that's happening across all of society.
- 2:07 A study at Wharton was looking exactly that.
- 2:10 How is AI reshaping human reasoning, and what does it mean for us as we continue to use it day to day?
- 2:17 They've observed an interesting phenomenon called cognitive surrender, when a human forgoes deliberation and adopts AI output as their own with minimal scrutiny.
- 2:27 They saw this in a study where humans were asked to look at reasoning exams, and they were given AI resources to answer those questions.
- 2:35 They realized that AI can either supplement or supplant a human's thinking, and the human might not even know it's happening.
- 2:42 For example, for questions where the AI was right, the human performance increased by 25 percentage points, whereas when the AI was wrong, it decreased by 15.
- 2:53 This suggests that the humans were taking in that AI cognition and not really thinking critically about what that answer is, and instead using those answers as part of their own reasoning together with their human instinct and deliberate reasoning, amplifying those results.
- 3:13 Most interestingly, they saw that 80% of participants were accepting those AI answers even when it was wrong.
- 3:20 So they were lowering that barrier to entry and just trusting the AI even if they weren't really critically examining the correctness of that result.
- 3:31 So at the Duolingo English test, we wanted to look at the same thing, and we published some research titled When Machines Mislead.
- 3:39 It was a case study on the English exam, specifically in AI human interaction.
- 3:45 So for context, what is the DET?
- 3:48 The DET is a high-stakes exam that measures your English proficiency, and it's fully online and remotely proctored.
- 3:55 You can take it on your own machine in the comfort of your home, and it still provides high quality results that 6,000 programs worldwide are trusting day to day.
- 4:06 As you can imagine, with a fully online assessment, there are some interesting things that from a technological standpoint that have to happen to ensure good quality security.
- 4:16 This includes identity verification, a lockdown testing environment, as well as a variety of ways in which AI-assisted monitoring is happening to predict different types of cheating.
- 4:27 Lastly, we have a final round of scrutiny where we have human proctors review all of the video footage of the exam, as well as all of those AI results.
- 4:40 Specifically for the study, we wanted to target one of our AI cheating detection systems, which is copy typing.
- 4:48 Copy typing is a form of cheating where you're writing down information that you're reading currently as opposed to something that you're thinking right now at this moment.
- 4:57 So as you can imagine, your typing patterns differ when you're transcribing versus composing.
- 5:03 And this is something that our custom model is measuring.
- 5:06 It's looking at anomalies and keystroke patterns, and it's flagging them for sessions that are unusual.
- 5:13 So our model is highly conservative, and so we're prioritizing fairness for our test takers.
- 5:20 So this is not a very common flag.
- 5:23 However, when human proctors are taking a look at these, they're well trained to properly examine those segments and see if this is a flag.
- 5:34 So for our experiment, we wanted to answer the question, would a skilled reviewer catch a false alarm or would they just rubber stamp it?
- 5:41 So we know that our proctors are highly accurate at detecting various forms of cheating.
loading