OSWorld
XLang AI LabScalable, real-computer environment for benchmarking multimodal agents on 369 open-ended tasks across Ubuntu, Windows and macOS; agents drive a virtual desktop and are scored by execution-based test scripts. Picked because it is the leading harness for computer-use agents (Claude Computer Use, OpenAI Operator both publish OSWorld scores). 2.9k stars, 86 contributors, 55 commits in last 90 days.
OSWorld: scalable real-computer environment + execution-based evaluation harness for multimodal/computer-use agents (NeurIPS 2024). Apache-2.0. 369 tasks (361 excl. 8 problematic Google Drive tasks). De-facto standard for computer-use agents (GPT-5.4 75.0%, Claude Opus 4.6 72.7%, Claude Sonnet 4.6 72.5% on OSWorld-Verified cited in 2026 releases). Confirmed live June 2026. Dual artifact: ships both the task set and the runner; registry placed it in evaluation_code (the harness); note the dataset overlaps benchmark_eval_data.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(github.com/xlang-ai/OSWorld)
Apache-2.0, full harness + environment public. Scored as the evaluation harness; the bundled 369-task set is also openly redistributable.
- https://github.com/xlang-ai/OSWorld recorded 2026-06-04
Apache-2.0 license, public source, computer-use agent eval harness
- https://os-world.github.io/ recorded 2026-06-04
369 tasks (361 excl. Google Drive), real-computer execution-based evaluation
Adoption
3 medium confidenceTop adoption signal for a harness = cited as the standard in frontier release reports: OSWorld scores reported in OpenAI (CUA 38.1%, GPT-5.4 75.0%), Anthropic (Claude Sonnet 4.5 61.4%, Sonnet 4.6 72.5%, Opus 4.6 72.7%) computer-use launches through 2026. Reach band is directional. The harness is run by a concentrated set of frontier labs/agent teams rather than a mass install base. 2.9k stars corroborate only.
- https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-06-04
Claude Sonnet 4.6 OSWorld score reported in official model release (standard-citation signal)
- https://os-world.github.io/ recorded 2026-06-04
best-model vs human baselines; ongoing updates with new LLMs/VLMs
Capability
5 high confidenceFrontier-defining within evaluation_code for the computer-use/agentic axis. Execution-based, real-environment agent evaluation is the hardest eval modality and OSWorld is its reference harness. A 5 here is for the category's frontier-definers; OSWorld qualifies for agentic eval.
- https://os-world.github.io/ recorded 2026-06-04
369 execution-based real-computer tasks, multi-OS, agent-trajectory evaluation
- https://github.com/xlang-ai/OSWorld recorded 2026-06-04
harness + baselines for Operator/GPT/Gemini/Claude/UI-TARS/Agent-S
Unchanged since 2026-06-24 (last edited, not re-checked)