OSWorld
XLang AI LabReal-computer environment and execution-based harness for benchmarking multimodal computer-use agents on 369 open-ended tasks across Ubuntu, Windows and macOS. Agents drive a virtual desktop and are scored by 134 execution-based evaluation functions against full trajectories, running under VMware, VirtualBox or Docker with KVM.
Eight Google Drive tasks may be excluded, leaving a 361-task run, and the project site now redirects to osworld-v1.xlang.ai. It ships both the task set and the runner; this record is the harness. Verified 2026-08-13 via the xlang-ai/OSWorld repository, its license endpoint and the project site.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(github.com/xlang-ai/OSWorld)
- core-gated
- ungated(full harness runs locally on VMware, VirtualBox or Docker, no paid tier)
Apache-2.0, full harness + environment public. Scored as the evaluation harness; the bundled 369-task set is also openly redistributable.
- https://github.com/xlang-ai/OSWorld recorded 2026-06-04
Apache-2.0 license, public source, computer-use agent eval harness
- https://os-world.github.io/ recorded 2026-06-04
369 tasks (361 excl. Google Drive), real-computer execution-based evaluation
- https://raw.githubusercontent.com/xlang-ai/OSWorld/main/README.md recorded 2026-08-11
README carries the Apache-2.0 badge and gives full setup instructions for running the benchmark locally under VMware, VirtualBox, or Docker with KVM. AWS and Azure appear only as optional providers for parallel runs. No paid tier, commercial edition, or license key is mentioned.
- https://api.github.com/repos/xlang-ai/OSWorld/license recorded 2026-08-13
GitHub's license endpoint reports spdx_id Apache-2.0 for xlang-ai/OSWorld and returns the Apache-2.0 text as the LICENSE body.
- https://raw.githubusercontent.com/xlang-ai/OSWorld/main/README.md recorded 2026-08-13
Apache-2.0 badge, full local setup under VMware, VirtualBox or Docker with KVM, AWS and Azure optional for parallel runs, and no paid tier, commercial edition or license key.
Adoption
3 medium confidenceBeing cited as the standard in frontier release reports is the top adoption signal for a harness, and OSWorld scores appear across computer-use launches through 2026: OpenAI (CUA 38.1%, GPT-5.4 75.0%) and Anthropic (Claude Sonnet 4.5 61.4%, Sonnet 4.6 72.5%, Opus 4.6 72.7%). The harness is run by a concentrated set of frontier labs and agent teams rather than by a mass install base, so the band is directional, and the 3.1k GitHub stars corroborate only. Two caveats on the evidence: the Anthropic Sonnet 4.6 post carries its OSWorld figure in a chart rather than in prose, so the 72.5% above is not quotable from that page; and os-world.github.io now redirects to osworld-v1.xlang.ai, where the leaderboard lives, though the cited URL still resolves through the redirect.
- https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-08-13
The Sonnet 4.6 release post reports OSWorld results and charts Sonnet's progress on it across sixteen months. The benchmark is cited, which is the standard-citation signal; the numeric score is in the chart rather than the prose.
- https://os-world.github.io/ recorded 2026-06-04
best-model vs human baselines; ongoing updates with new LLMs/VLMs
- https://osworld-v1.xlang.ai/ recorded 2026-08-13
the site os-world.github.io now redirects here; it still publishes the human-vs-model baselines - 'While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success' for the original v1 setting - and still lists the model families evaluated (Operator, GPT, Gemini, Claude, UI-TARS, Agent-S, Qwen, CogAgent).
Capability
5 high confidenceFrontier-defining on the computer-use and agentic axis. Execution-based evaluation in a real environment is the hardest eval modality, and OSWorld is its reference harness: 369 tasks (361 excluding the eight Google Drive ones), 134 execution-based evaluation functions, and multi-platform virtualization across VMware, VirtualBox, Docker and AWS.
- https://os-world.github.io/ recorded 2026-06-04
369 execution-based real-computer tasks, multi-OS, agent-trajectory evaluation
- https://github.com/xlang-ai/OSWorld recorded 2026-06-04
harness + baselines for Operator/GPT/Gemini/Claude/UI-TARS/Agent-S
- https://osworld-v1.xlang.ai/ recorded 2026-08-13
369 tasks with an official note that the 8 Google Drive tasks may be excluded for a 361-task run; 134 execution-based evaluation functions; 'reliable, reproducible setup and evaluation scripts'.
- https://raw.githubusercontent.com/xlang-ai/OSWorld/main/README.md recorded 2026-08-13
harness runs the environment under VMware, VirtualBox, Docker/KVM, AWS or Azure and evaluates full agent trajectories against any VLM/LLM agent.
Verified 2026-08-13