AI that operates the Mac
Computer Use is how an AI agent sees the screen the way macOS does, acts in real apps one verified step at a time, and gets measured on real Mac tasks. It is the research behind computer use in Eney - published, benchmarked, and built in Swift.
Four stages, repeated until the task is done. Each one is a research line of its own, and each one has a number behind it.
Capture the screen and its structure
A screenshot plus the accessibility tree of the focused window. When an app ships no usable tree, Screen2AX rebuilds one from the pixels.
Pick the next step
A vision-language model reads the task, the history, and the screen, and answers with one action - in a grammar the agent can parse and check before anything happens.
Click, type, scroll, press
The action is mapped from model space to screen points and performed with native system events in the real app. One step per turn, never a batch.
Look again before going on
The next screenshot shows what the step did. Risky actions wait for the user, and finished runs are scored by evaluators that inspect files and app state - not by another model's opinion.
macOS describes every window as a tree of elements - the structure screen readers and agents rely on. Only about a third of Mac apps ship it completely. Screen2AX rebuilds that tree from a screenshot.
A detector finds UI elements and groups, a vision-language model describes them, and the result is assembled into a tree that mirrors macOS's own accessibility structure - in real time, from a single screenshot.
When GUIrilla was built, macOS barely existed in public agent data - 0.06% of OS-Atlas, 2.45% of automatically collected desktop UI. GUIrilla crawls real apps and turns every screen into grounded tasks - no hand labeling.
Feed it an application bundle. It installs the app, walks every reachable state, and leaves behind a structured graph of the entire interface - then cleans up after itself.
GUIrilla takes an application bundle, installs it, and uninstalls it once the crawl is done.
Exploration runs through the accessibility tree, with handlers for pop-ups, menu unrolling, element and graph ordering, and empty or invisible elements. Together with the agents they raised task discovery 5× in Stocks and 3× in Maps.
An input agent fills fields with plausible data, an order-and-login agent queues irreversible actions last and pings a human when a login is required, and a task agent turns each action and its UI context into a task.
A structured interface graph of the whole app, and one grounded task per action: a screenshot, an accessibility state, and a single click or type.
GUIrilla-Task: 27,171 tasks across 1,108 apps and 23 genres, about 4,200 unique full-desktop screens, each a screenshot, an accessibility state, and one grounded click or type. MacApp Trees: 561 GB of structured interaction graphs, released with the dataset.
GUIrilla-Gold: 1,283 tasks verified by five accessibility-expert annotators, reaching 90.26% task alignment. Task wording matters, too: training on tasks rewritten around functional role gave 53.55% accuracy against 40.35%.
GUIrilla-See comes in 0.7B, 3B, and 7B, trained on synthetic data only. The 7B model reaches 94.73% on ScreenSpot-v2 and, at publication, the best macOS score in the field from 4.2K images - UI-TARS 2B needed an estimated 20 million.
Click accuracy in percent, as published. UI-TARS 7B is the reference on ScreenSpot-v2 - trained on an estimated 20 million images against GUIrilla-See's 4.2K.
| Model | ScreenSpot-v2 | ScreenSpot-Pro | Pro · macOS |
|---|---|---|---|
| GUIrilla-See 0.7B | 53.55% | 7.34% | 7.95% |
| GUIrilla-See 3B | 89.54% | 29.35% | 32.62% |
| GUIrilla-See 7B | 94.73% | 35.36% | 41.39% |
| UI-TARS 7B reference | 91.60% | – | – |
MacArena puts computer-use agents in the ring on real macOS: 421 tasks across first- and third-party apps - 372 ported from OSWorld and macOSWorld and fixed for the Mac, 49 written for it - every one human-verified, run in fresh virtual machines on Apple silicon, and scored by execution-based checks of files, app state, and shell output, never by a judge model.
221 ported from OSWorld and 151 from macOSWorld, each reviewed by hand and fixed for macOS, plus 49 new macOS-native tasks across 20 apps with hand-crafted evaluation scripts.
The only earlier macOS benchmark ran on x86 virtual machines and covered mostly first-party apps.
At publication, no agent cleared 37% on the tasks written for the Mac. The leader, OpenAI CUA, trails its own OSWorld result on identical tasks.
The 221 OSWorld tasks exist on both platforms, so they are a direct comparison - and every model with a public reference score does worse on the Mac: OpenAI CUA by 9.26 points, Qwen3-VL 4B by 9.84, Qwen3-VL 2B by 7.05, UI-TARS-1.5 7B by 3.23. Different looks, shortcuts, window management, and system behavior cost real accuracy.
UI-TARS-1.5 7B leads OpenAI CUA on the Linux-born OSWorld tasks (21.27% against 16.74%) and collapses on the 49 tasks written for macOS (10.20% against 36.73%). The drop concentrates in macOS-specific interfaces: on MacArena's file-management tasks every agent scored 0%. Recognizing task shapes seen in training is not the same as operating a new platform.
The 221 OSWorld and 151 macOSWorld tasks were ported to macOS and, like the 49 new ones, reviewed by hand to be executable, unambiguous, and correctly specified - and fixed where they were not. Every task is scored by checking files, app state, and shell output in a fresh copy-on-use VM, never by a judge model.
Success rates with a 15-step limit, two runs per task, as published in "MacArena: Benchmarking Computer Use Agents on an Online macOS Environment" (2nd AIWILD Workshop @ ICML 2026). All code and virtual machines are released.
The agent itself is a dependency-free Swift package - the port of the Python research agent that collected GUIrilla's trajectories - built to be embedded in Mac software and to run on the device.
The same package sits behind the Computer Use screen in Eney, in development. Permissions stay with the host app: Accessibility and Screen Recording are granted to it, never to the model. Every run is recorded as a trajectory - numbered screenshots, a JSON-lines log, and a results table - the layout the research agent used to collect training data.
ScreenCaptureKit captures the full screen or the focused window at native resolution, an accessibility walker produces the same JSON tree macapptree does, and clicks, drags, scrolls, hotkeys, and Unicode typing go through system events - with the model's coordinates mapped to screen points in one place.
UI-TARS-1.5 7B is the verified default - served by vLLM, or on the Mac through LM Studio. General models use the same action grammar, and OpenAI's native computer tool is supported as well.
Every action passes a delegate before it happens: show it, approve it, or stop the run. A run can pause on a question and resume with the answer in the same conversation.
With UI-TARS-1.5 7B served locally on an M5 with 24 GB of memory, the model loads in about 9 seconds and a step takes 17 to 35 seconds depending on the screenshot budget - the whole loop without leaving the device.
Three peer-reviewed workshop papers, with the code, models, datasets, and virtual machines released alongside them.
Computer Use is a research preview. Working on agents for the Mac, on accessibility, or on evaluation? Tell us what you are building and we will talk about data, benchmarks, and integration.
Email the team [email protected]The short answers - the rest of the page has the detail.
Computer Use is MacPaw Research's stack for AI agents that operate the Mac: Screen2AX rebuilds an app's accessibility tree from a screenshot, GUIrilla produces training data and grounding models from real apps, MacArena measures agents on real macOS tasks, and a Swift agent performs the actions. It is a research preview.
It captures a screenshot and reads the accessibility tree of the focused window through the macOS Accessibility API. When an app ships no usable tree - about two thirds of Mac apps are incomplete - Screen2AX rebuilds one from the pixels with an F1 score of 79%.
It can. The Swift agent talks to any OpenAI-compatible endpoint, including LM Studio on the same Mac. With UI-TARS-1.5 7B served locally on an M5 with 24 GB of memory, a step takes 17 to 35 seconds and nothing leaves the device. Cloud models, including OpenAI's computer tool, are supported when the host chooses them.
The agent performs one action per turn and asks its host before each one, so the app can show the step, require approval for risky actions, or stop the run. Permissions for Accessibility and Screen Recording belong to the host app, never to the model.
Three workshop papers (MacArena at ICML 2026, GUIrilla at ICLR 2026, and Screen2AX), the MacArena and GUIrilla code and virtual machines, macapptree, the GUIrilla-See models and detectors on Hugging Face, and the GUIrilla, Screen2AX, and UiPad datasets.