AI Potluck
Model components / Evaluation code

OSWorld

XLang AI Lab

Real-computer environment and execution-based harness for benchmarking multimodal computer-use agents on 369 open-ended tasks across Ubuntu, Windows and macOS. Agents drive a virtual desktop and are scored by 134 execution-based evaluation functions against full trajectories, running under VMware, VirtualBox or Docker with KVM.

Eight Google Drive tasks may be excluded, leaving a 361-task run, and the project site now redirects to osworld-v1.xlang.ai. It ships both the task set and the runner; this record is the harness. Verified 2026-08-13 via the xlang-ai/OSWorld repository, its license endpoint and the project site.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(github.com/xlang-ai/OSWorld)
core-gated
ungated(full harness runs locally on VMware, VirtualBox or Docker, no paid tier)

Apache-2.0, full harness + environment public. Scored as the evaluation harness; the bundled 369-task set is also openly redistributable.

Adoption

3 medium confidence
3.0

Being cited as the standard in frontier release reports is the top adoption signal for a harness, and OSWorld scores appear across computer-use launches through 2026: OpenAI (CUA 38.1%, GPT-5.4 75.0%) and Anthropic (Claude Sonnet 4.5 61.4%, Sonnet 4.6 72.5%, Opus 4.6 72.7%). The harness is run by a concentrated set of frontier labs and agent teams rather than by a mass install base, so the band is directional, and the 3.1k GitHub stars corroborate only. Two caveats on the evidence: the Anthropic Sonnet 4.6 post carries its OSWorld figure in a chart rather than in prose, so the 72.5% above is not quotable from that page; and os-world.github.io now redirects to osworld-v1.xlang.ai, where the leaderboard lives, though the cited URL still resolves through the redirect.

  • https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-08-13

    The Sonnet 4.6 release post reports OSWorld results and charts Sonnet's progress on it across sixteen months. The benchmark is cited, which is the standard-citation signal; the numeric score is in the chart rather than the prose.

  • https://os-world.github.io/ recorded 2026-06-04

    best-model vs human baselines; ongoing updates with new LLMs/VLMs

  • https://osworld-v1.xlang.ai/ recorded 2026-08-13

    the site os-world.github.io now redirects here; it still publishes the human-vs-model baselines - 'While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success' for the original v1 setting - and still lists the model families evaluated (Operator, GPT, Gemini, Claude, UI-TARS, Agent-S, Qwen, CogAgent).

Capability

5 high confidence
5.0

Frontier-defining on the computer-use and agentic axis. Execution-based evaluation in a real environment is the hardest eval modality, and OSWorld is its reference harness: 369 tasks (361 excluding the eight Google Drive ones), 134 execution-based evaluation functions, and multi-platform virtualization across VMware, VirtualBox, Docker and AWS.

Verified 2026-08-13