Videos ZSQb5fzRFPw
Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
Scene timeline
49 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 154
- whisperx 154
- chunks
- 29
- from 154 cues
- keyframes
- 39
- kept of 49 captured
- frames with text
- 38
- 862 lines read
- chapters
- 6
- from the source metadata
- keyframe bytes
- 5.3 MB
- word timings on 154 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 02:48 | 1m 38s |
stt |
done | — | 2026-08-11 02:49 | 16s |
chunk |
done | — | 2026-08-11 02:50 | 0s |
text_embed |
done | — | 2026-08-11 02:50 | 1s |
keyframe |
done | — | 2026-08-11 02:50 | 1m 45s |
ocr |
done | — | 2026-08-11 02:51 | 17s |
frame_embed |
done | — | 2026-08-11 02:52 | 7s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.91
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.90
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.99
-
- Cua1.00
- AlEngineer0.98
- World's Fair0.99
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.99
- World'sFair1.00
- Cua1.00
- THE OLD LOOP0.99
- 1.0 = the Human Loop0.99
- The agent watched the screen, chose the next move, acted like a user, and repeated.1.00
- PRESENTED BY1.00
- SCREENSHOT1.00
- REASON1.00
- CLICK TARGET1.00
- Safari1.00
- Edit0.98
- View1.00
- Window1.00
- Help1.00
- "The user asked for a0.97
- Safari1.00
- Edit1.00
- View1.00
- Go1.00
- Window1.00
- Help1.00
- Microsoft1.00
- PDF. The export button0.97
- is visible in the lower1.00
- Reports1.00
- Quarterly report0.96
- Export PDF1.00
- right."0.99
- Reports1.00
- Quarterly report1.00
- Export0.89
- Export1.00
- Choose one action, then observe again.0.99
- 1.0 METHODS0.99
- screenshots1.00
- OCR / SoM0.99
- mouse1.00
- keyboard1.00
- scroll1.00
- type1.00
- wait1.00
- pyautogui1.00
- screen first - action next · foreground repeat0.97
- cua.ai1.00
- D0.62
- Cua0.80
- Engineering the future of Al0.98
- World's Fair0.99
-
- AlEngineer0.99
- World'sFair1.00
- Cua1.00
- THE OLD LOOP0.96
- 1.0 = the Human Loop0.98
- The agent watched the screen, chose the next move, acted like a user, and repeated.1.00
- PRESENTED BY1.00
- SCREENSHOT1.00
- REASON1.00
- CLICK TARGET1.00
- Safari1.00
- View1.00
- Window1.00
- Help1.00
- "The user asked for a1.00
- Safari1.00
- Edit1.00
- View1.00
- Go1.00
- Window1.00
- Help1.00
- Microsoft1.00
- PDF. The export button0.99
- is visible in the lower0.97
- Reports1.00
- Quarterly report0.99
- Export PDF1.00
- right."0.99
- Reports1.00
- Quarterly report1.00
- Choose one action, then observe again.0.99
- 1.0 METHODS0.99
- screenshots1.00
- OCR / SoM0.99
- mouse1.00
- keyboard1.00
- scroll1.00
- type1.00
- wait1.00
- pyautogui1.00
- screen first · action next · foreground repeat0.97
- cua.ai0.99
- 00.57
- 000.64
- TRACK 7·JULY 1,20260.95
- ComputerUse1.00
- World's Fair0.97
-
- AlEngineer0.98
- World's Fair0.97
- Cua1.00
- COMPUTER-USE 1.0 EXAMPLE0.98
- A GUI-first operator0.99
- The old loop wired the model straight to the GUl: prompt and screenshot in, action out.0.99
- computer_use_setup.py1.00
- MODEL1.00
- response = client.responses.create(1.00
- computer-use-preview1.00
- model="computer-use-preview",1.00
- The loop is native to the model.1.00
- tools=[{1.00
- "type": "computer_use_preview",0.99
- "display_width": 1024,0.99
- "display_height": 768,1.00
- TOOL1.00
- },0.52
- "environment": "mac",0.98
- provider-native tools1.00
- input=[{0.99
- "role": "user",0.98
- OpenAI Responses API or Anthropic beta.0.97
- "content": [0.99
- {"type": "input_text", "text": task},0.99
- {"type": "input_image", "image_url": screenshot_url},0.98
- INPUT1.00
- H],0.50
- text + screenshot1.00
- truncation="auto",1.00
- The next action depends on pixels.1.00
- Computer-Use model setup · prompt + screenshot loop0.98
- cua.ai1.00
- OUTPUT1.00
- next: computer_call1.00
- 40.95
- TRACK 7·JULY 1,20260.96
- World's Fair1.00
- AlEngineet0.89
- Computer Use0.98
-
- AlEngineer1.00
- World's Fair0.96
- Cua1.00
- BACKGROUND CONTROL1.00
- À la SkyLight Recipe0.96
- Focus without raise, post to pid, prime Chromium, keep AX live0.99
- yabai pattern1.00
- 011.00
- Flip AppKit active without raising the window or switching0.99
- Spaces.1.00
- SLEventPostToPid1.00
- 820.87
- SkyLight sends mouse events to the target pid. outside the0.99
- HID stream.1.00
- primer click1.00
- 031.00
- An off-screen click at -1,-1 warms Chromium's activation gate.0.99
- live AX1.00
- 840.87
- Private AX SPI keeps occluded Electron trees fresh.0.99
- TRACK 7· JULY 1,20260.96
- Computer Use1.00
- World's Fair0.97
-
- AlEngineer0.99
- World's Fair0.99
- Cua1.00
- BACKGROUND CONTROL1.00
- À la SkyLight Recipe1.00
- Focus without raise, post to pid, prime Chromium, keep AX live1.00
- yabai pattern1.00
- 011.00
- 58°1.00
- Flip AppKit active without raising the window or switching0.99
- Spaces.1.00
- SLEventPostToPid1.00
- 020.86
- SkyLight sends mouse events to the farget pid. outside the0.99
- HID stream.1.00
- primer click1.00
- 030.99
- An off-screen click at -1,-1 warms Chromium's activation gate.0.99
- live AX1.00
- 840.94
- Private AX SPI keeps occluded Electron trees fresh.1.00
- TRACK 7· JULY 1,20260.95
- ComputerUse1.00
- World's Fair0.99
-
- World's Fair0.99
- AlEngineer0.99
- Cua1.00
- CROSS-PLATFORM1.00
- One interface to rule them all.0.98
- Different operating systems, different display servers, one agent-facing contract.1.00
- MAY 28260.96
- JUNE 20260.98
- Windows1.00
- Linux1.00
- UIA + messages1.00
- AT-SPI first0.99
- State comes from UIA/MSAA: pixels from PrintWindow/WGC.1.00
- Xll and Wayland expose opposite input models, so the driver0.99
- PostMessage handles many apps: SendInput is the explicit0.99
- keeps the agent contract semantic.0.97
- foreground rung.1.00
- X11 / XWAYLAND0.95
- background pixels, XTEST keys, recording0.98
- UIA / MSAA0.99
- PrintWindow / WGC0.98
- PostMessage1.00
- NATIVE WAYLAND1.00
- AT-SPI actions; raw keys need fallback0.98
- Windows: UIA + messages - Linux: X11 routes pixels, Wayland routes semantics0.98
- cua.ai1.00
- 000.66
- TRACK 7· JULY 1,20260.93
- ComputerUse1.00
- World's Fair0.96
-
- AlEngineer0.98
- World'sFair1.00
- Cua1.00
- REGRESSION MATRIX1.00
- Every action path gets tested0.99
- 500+1.00
- Tests run against real toolkits and app-state changes.1.00
- SCORED ACTION CHECKS0.99
- effect plus background-held contract1.00
- 5 MODALITY RUNS0.99
- ACTION PATH X DISPATCH X SCOPE0.98
- TOOLKIT LANES1.00
- MODE0.99
- TARGET1.00
- DISPATCH1.00
- VERDICT0.98
- WINDOWS1.00
- WPF1.00
- WinUI31.00
- WebView21.00
- Electron1.00
- ax-bg1.00
- element_index1.00
- background1.00
- effect + held0.97
- px-bg1.00
- x,y0.91
- background1.00
- effect + held0.97
- QLINUX0.98
- GTK3/41.00
- Qt5/61.00
- Electron1.00
- ax-fg1.00
- element_index1.00
- foreground1.00
- effect1.00
- MACOS0.99
- AppKit1.00
- SwiftUI1.00
- WKWebView1.00
- Electron1.00
- px-fg1.00
- x,y0.93
- foreground1.00
- effect1.00
- px-desktop1.00
- screen pixels1.00
- desktop1.00
- effect1.00
- EFFECT1.00
- CONTRACT1.00
- Did the app state actually1.00
- Did the target stay0.99
- change?1.00
- backgrounded?1.00
- 6 PARITY CONTROLS1.00
- checkbox1.00
- click target1.00
- slider1.00
- scroll1.00
- text input1.00
- 8 actions x 6 controls x 10+ lanes = 500+ scored checks0.99
- cua.ai0.98
- Oun0.73
- TRACK 7· JULY 1,20260.95
- Computer Use0.98
- World's Fair0.96
-
- AlEngineer0.99
- World'sFair1.00
- Cua1.00
- USED TODAY1.00
- Our early adopters1.00
- These products and agent labs use the same MCP and CLI driver for desktop apps.1.00
- PRESENTED BY1.00
- Microsoft1.00
- 1OH0.55
- Clicky1.00
- Hermes1.00
- Qwen Code0.98
- H Company1.00
- Droid Factory1.00
- CUA PRODUCT0.96
- AGENT RUNTIME1.00
- CODING AGENT0.99
- AGENT PLATFORM0.99
- FACTORY1.00
- ENTRY POINTS0.99
- TARGETS1.00
- LOOP1.00
- MCP, CLI, and SDK0.96
- real desktop apps0.98
- observe, act, verify0.97
- MCP and CLI across real apps1.00
- cua.ai0.99
- n0.73
- TRACK 7· JULY 1,20260.95
- ComputerUse1.00
- World's Fair0.98
-
- AlEngineer0.99
- World's Fair0.98
- Cua1.00
- INTELLIGENCE1.00
- Now: can the agent use the computer?1.00
- PRESENTED BY1.00
- The driver gives agents hands. The next question is whether those hands1.00
- produce reliable work.1.00
- Microsoft1.00
- SURFACE1.00
- TEST1.00
- SPEAKER1.00
- Cua Driver1.00
- Cua Bench0.95
- Dillon DuPont0.97
- observe, act, verify0.97
- measure real task progress0.99
- intelligence layer1.00
- Cua Bench · intelligence section0.96
- cua.ai1.00
- TRACK 7· JULY 1, 20260.93
- ComputerUse1.00
- World's Fair0.99
-
- AlEngineer0.98
- World's Fair0.99
- Cua1.00
- ACTIONS → CAN WE TRUST THEM0.98
- Cua Driver gives agents actions.1.00
- CuaBench measures whether those actions solve the task.1.00
- PRESENTED BY1.00
- Microsoft1.00
- 011.00
- 020.98
- 030.77
- setup1.00
- act1.00
- score1.00
- known state1.00
- the agent works1.00
- checks pass or fail0.99
- verifiable environments · every result reproducible0.98
- cua.ai1.00
- TRACK 7· JULY 1,20260.95
- World's Fair0.99
- Computer Use1.00
-
- AlEngineer0.98
- World'sFair1.00
- Cua1.00
- A TASK = THREE PIECES0.99
- Every task is three pieces.1.00
- task.py1.00
- PRESENTED BY1.00
- from cua_bench import setup_task, solve_task, evaluate_task0.99
- Microsoft1.00
- @setup_task1.00
- def setup(env): env.open("invoice.xlsx")0.99
- @solve_task1.00
- def solve(agent): agent.run("Put the total in B2")1.00
- @evaluate_task1.00
- def check(env): return env.cell("B2") = "42"0.99
- setup · solve · evaluate · five platforms0.93
- cua.ai1.00
- TRACK 7· JULY 1, 20260.96
- World's Fair0.97
- AlEngineer0.99
- Computer Use0.97
-
- AlEngineer0.98
- World'sFair1.00
- Cua1.00
- ONE FILE, ANY PLATFORM1.00
- Environments take scale and expertise.1.00
- env.py1.00
- env window1.00
- from bench_ui import launch_window, get_element_rect1.00
- # inline HTML, no browser, no server1.00
- rect + agent0.94
- pid = launch_window(html="<h1>Invoice</h1>")0.99
- Invoice0.99
- # real screen-space coords for the agent0.99
- rect = get_element_rect(pid, "h1")0.99
- THE AGENT SEES EXACTLY WHAT YOU DREW0.98
- author · launch· inspect · in one file0.92
- cua.ai1.00
- TRACK 7· JULY 1, 20260.96
- AlEngineer1.00
- Computer Use0.97
- World'sFair1.00
-
- AlEngineer0.99
- World'sFair1.00
- Cua1.00
- 130 TASKS· CUA-BENCH-KICAD0.96
- One open catalog.1.00
- 1301.00
- 420.86
- 51.00
- VERIFIABLE TASKS1.00
- ENVIRONMENTS1.00
- PLATFORMS1.00
- $ cb run dataset cua-bench-kicad -m sonnet-4-51.00
- linux · windows · android · macos · web0.93
- cua.ai1.00
- TRACK 7· JULY 1,20260.96
- AlEngineer0.99
- Computer Use0.96
- World'sFair1.00
Transcript
154 cues· 2,311 words· 12,442 chars
- 0:12 Thank you for taking the time for coming over here.
- 0:15 I'm Francesco, I'm the CEO of the company.
- 0:19 Alongside me, a couple of other folks, my CTO, Dylan, and my chief of infra, Rob, they're gonna work on the stage in a while.
- 0:27 But before we do that, who's excited for some computer-using agent talk happening now?
- 0:32 Are you guys excited?
- 0:33 Lovely.
- 0:35 If I were to ask like what was a computer using agent like one year ago, probably half the crowd would say I don't have any idea what really computer use mean.
- 0:45 So today I'm gonna take you to a journey.
- 0:50 basically from our vision where we come from so far on computer user.
- 0:57 This new shape of agents that are talking and up to model intelligence.
- 1:08 So we're gonna start with the vision of Quadriver, where we're coming from.
- 1:13 And...
- 1:15 If you, how many of you guys have been working with computer use for one year?
- 1:22 How about like three years?
- 1:25 Lovely, okay, so our team has plenty of experience, like we go all the way back our time on Microsoft, we were working on this type of GUI agents, we were calling them back in the days.
- 1:37 And there is a,
- 1:42 There is an example of like old-fashioned human agent loop.
- 1:47 We basically refer to this as human loop where you will have like an agent loop, you will take a screenshot that the agents will have to reason and plan through, and then you will basically work with an action space in terms of like clicking, typing, scrolling around.
- 2:07 This is what we would refer as the old-fashioned computer use 1.0, just to set the tone for this talk.
- 2:17 And we come along a long way since this type of computer using agents.
- 2:27 This again, I'm gonna skim over these slides, but that's the old-fashioned way of representing these agents loop, as a human would do.
- 2:35 We...
- 2:40 Here we go.
- 2:42 Over like two months ago, we released a project in the open source.
- 2:45 It's called Quadriver.
- 2:47 And we made it working like in the background.
- 2:55 That means that your computer user will not take over your screen as like the computer used 1.0.
- 3:04 kind of like Agent Loop was doing back in the days.
- 3:07 And it all started from Codex releasing their computer user code.
- 3:18 model two months ago.
- 3:20 So we cannot take the challenge because we were already working with this type of background computer user.
- 3:26 So over one weekend, we hacked something together.
- 3:30 And the trick here is really not having your agents take over your screen.
- 3:36 So there is a lot of dark magic happening behind the wall just
- 3:41 give you some context.
- 3:42 There are like some undocumented API living in the Apple framework and basically ships with your laptop.
- 3:50 And as you can see here, like it's in the demo, you have like an agent that is not taking over control over your laptop.
- 3:59 We made it working not only for MacOS, but also spanning across Windows and Linux.
- 4:04 This is the very first driver that is living on your laptop, and it lets really any AI agents connect to the underlying operating system, either using accessibility trees or a screenshot-level approach.
- 4:22 We kind of take on both.
- 4:26 This is what really the agents see for what it concerns.
- 4:30 You will have to install Quadriver, the agents will take a snapshot of the Windows data, and you will have to observe, and we really like take
- 4:44 take one different action path to really make the ground computer use happening.
- 4:50 So you really have to observe the space.
- 4:54 In this case, just by calling getWindowState, you get an accessibility tree representation plus a screenshot.
- 5:02 And then you will go and
- 5:07 try a background execution using accessibility tree, and if that doesn't work, we go all the way and make the heavy lifting for you and just try a pixel background click.
- 5:17 This is like, kind of like best step for background at this stage.
- 5:22 It's not like behaving the same way on MacOS, Windows, and Linux.
- 5:26 So we do like some of the lifting for you so that your AI agent can run on this on your MacBook.
loading