Computer Use

Research preview

AI that operates the Mac

Computer Use is how an AI agent sees the screen the way macOS does, acts in real apps one verified step at a time, and gets measured on real Mac tasks. It is the research behind computer use in Eney - published, benchmarked, and built in Swift.

How an agent uses the Mac

Four stages, repeated until the task is done. Each one is a research line of its own, and each one has a number behind it.

  1. See

    Capture the screen and its structure

    A screenshot plus the accessibility tree of the focused window. When an app ships no usable tree, Screen2AX rebuilds one from the pixels.

  2. Decide

    Pick the next step

    A vision-language model reads the task, the history, and the screen, and answers with one action - in a grammar the agent can parse and check before anything happens.

  3. Act

    Click, type, scroll, press

    The action is mapped from model space to screen points and performed with native system events in the real app. One step per turn, never a batch.

  4. Check

    Look again before going on

    The next screenshot shows what the step did. Risky actions wait for the user, and finished runs are scored by evaluators that inspect files and app state - not by another model's opinion.

Sees the screen like macOS does

macOS describes every window as a tree of elements - the structure screen readers and agents rely on. Only about a third of Mac apps ship it completely. Screen2AX rebuilds that tree from a screenshot.

  • A detector finds UI elements and groups, a vision-language model describes them, and the result is assembled into a tree that mirrors macOS's own accessibility structure - in real time, from a single screenshot.

Learns from real apps

When GUIrilla was built, macOS barely existed in public agent data - 0.06% of OS-Atlas, 2.45% of automatically collected desktop UI. GUIrilla crawls real apps and turns every screen into grounded tasks - no hand labeling.

Feed it an application bundle. It installs the app, walks every reachable state, and leaves behind a structured graph of the entire interface - then cleans up after itself.

  • An app goes in

    GUIrilla takes an application bundle, installs it, and uninstalls it once the crawl is done.

  • Every reachable state gets walked

    Exploration runs through the accessibility tree, with handlers for pop-ups, menu unrolling, element and graph ordering, and empty or invisible elements. Together with the agents they raised task discovery 5× in Stocks and 3× in Maps.

  • Three agents keep it realistic

    An input agent fills fields with plausible data, an order-and-login agent queues irreversible actions last and pings a human when a login is required, and a task agent turns each action and its UI context into a task.

  • A graph and tasks come out

    A structured interface graph of the whole app, and one grounded task per action: a screenshot, an accessibility state, and a single click or type.

  • The data

    GUIrilla-Task: 27,171 tasks across 1,108 apps and 23 genres, about 4,200 unique full-desktop screens, each a screenshot, an accessibility state, and one grounded click or type. MacApp Trees: 561 GB of structured interaction graphs, released with the dataset.

  • The gold set

    GUIrilla-Gold: 1,283 tasks verified by five accessibility-expert annotators, reaching 90.26% task alignment. Task wording matters, too: training on tasks rewritten around functional role gave 53.55% accuracy against 40.35%.

  • The models

    GUIrilla-See comes in 0.7B, 3B, and 7B, trained on synthetic data only. The 7B model reaches 94.73% on ScreenSpot-v2 and, at publication, the best macOS score in the field from 4.2K images - UI-TARS 2B needed an estimated 20 million.

GUIrilla-See grounding accuracy

Click accuracy in percent, as published. UI-TARS 7B is the reference on ScreenSpot-v2 - trained on an estimated 20 million images against GUIrilla-See's 4.2K.

ModelScreenSpot-v2ScreenSpot-ProPro · macOS
GUIrilla-See 0.7B 53.55%7.34%7.95%
GUIrilla-See 3B 89.54%29.35%32.62%
GUIrilla-See 7B 94.73%35.36%41.39%
UI-TARS 7B reference91.60%––

Measured on real Macs

MacArena puts computer-use agents in the ring on real macOS: 421 tasks across first- and third-party apps - 372 ported from OSWorld and macOSWorld and fixed for the Mac, 49 written for it - every one human-verified, run in fresh virtual machines on Apple silicon, and scored by execution-based checks of files, app state, and shell output, never by a judge model.

  • 421

    Human-verified tasks

    221 ported from OSWorld and 151 from macOSWorld, each reviewed by hand and fixed for macOS, plus 49 new macOS-native tasks across 20 apps with hand-crafted evaluation scripts.

  • 50

    macOS applications

    The only earlier macOS benchmark ran on x86 virtual machines and covered mostly first-party apps.

  • 36.73%

    Best macOS-native score

    At publication, no agent cleared 37% on the tasks written for the Mac. The leader, OpenAI CUA, trails its own OSWorld result on identical tasks.

MacArena's own tasks, written for macOS Ported from OSWorld and macOSWorld Success rate · higher is better ↑

Success rates with a 15-step limit, two runs per task, as published in "MacArena: Benchmarking Computer Use Agents on an Online macOS Environment" (2nd AIWILD Workshop @ ICML 2026). All code and virtual machines are released.

Acts in real apps

The agent itself is a dependency-free Swift package - the port of the Python research agent that collected GUIrilla's trajectories - built to be embedded in Mac software and to run on the device.

The same package sits behind the Computer Use screen in Eney, in development. Permissions stay with the host app: Accessibility and Screen Recording are granted to it, never to the model. Every run is recorded as a trajectory - numbered screenshots, a JSON-lines log, and a results table - the layout the research agent used to collect training data.

  • Native perception and input

    ScreenCaptureKit captures the full screen or the focused window at native resolution, an accessibility walker produces the same JSON tree macapptree does, and clicks, drags, scrolls, hotkeys, and Unicode typing go through system events - with the model's coordinates mapped to screen points in one place.

  • Any vision-language model

    UI-TARS-1.5 7B is the verified default - served by vLLM, or on the Mac through LM Studio. General models use the same action grammar, and OpenAI's native computer tool is supported as well.

  • The user stays in the loop

    Every action passes a delegate before it happens: show it, approve it, or stop the run. A run can pause on a question and resume with the answer in the same conversation.

  • Runs on the Mac

    With UI-TARS-1.5 7B served locally on an M5 with 24 GB of memory, the model loads in about 9 seconds and a step takes 17 to 35 seconds depending on the screenshot budget - the whole loop without leaving the device.

Published, not promised

Three peer-reviewed workshop papers, with the code, models, datasets, and virtual machines released alongside them.

Talk to the team

Build with Computer Use

Computer Use is a research preview. Working on agents for the Mac, on accessibility, or on evaluation? Tell us what you are building and we will talk about data, benchmarks, and integration.

Email the team [email protected]

Frequently asked questions

The short answers - the rest of the page has the detail.

What is Computer Use?

Computer Use is MacPaw Research's stack for AI agents that operate the Mac: Screen2AX rebuilds an app's accessibility tree from a screenshot, GUIrilla produces training data and grounding models from real apps, MacArena measures agents on real macOS tasks, and a Swift agent performs the actions. It is a research preview.

How does the agent see the screen?

It captures a screenshot and reads the accessibility tree of the focused window through the macOS Accessibility API. When an app ships no usable tree - about two thirds of Mac apps are incomplete - Screen2AX rebuilds one from the pixels with an F1 score of 79%.

Does it run on-device?

It can. The Swift agent talks to any OpenAI-compatible endpoint, including LM Studio on the same Mac. With UI-TARS-1.5 7B served locally on an M5 with 24 GB of memory, a step takes 17 to 35 seconds and nothing leaves the device. Cloud models, including OpenAI's computer tool, are supported when the host chooses them.

How is it kept safe?

The agent performs one action per turn and asks its host before each one, so the app can show the step, require approval for risky actions, or stop the run. Permissions for Accessibility and Screen Recording belong to the host app, never to the model.

What is published?

Three workshop papers (MacArena at ICML 2026, GUIrilla at ICLR 2026, and Screen2AX), the MacArena and GUIrilla code and virtual machines, macapptree, the GUIrilla-See models and detectors on Hugging Face, and the GUIrilla, Screen2AX, and UiPad datasets.