H Company's Holo4 Scores 61.7% on OSWorld 2.0 With a 27B Model. The Weights Are Not Free for Commercial Use.
The French lab's new agent models drive GUIs, code, MCP, and APIs with one calling convention and publish every benchmark trajectory. Most write-ups call them open. The license says non-commercial.
TL;DR
H Company released Holo4 on September 28 in two sizes, a 27B dense model and a 35B-A3B mixture-of-experts model, built on Qwen bases and trained to operate software through screens, code, MCP servers, and APIs with the same model and the same call. Holo4 27B scores 61.7% on OSWorld 2.0, against 81.8% for Opus 5.5; the 35B-A3B model scores 30.9%. Both are on the H Models API, with unannounced pricing, and on Hugging Face, where the Holo4-27B weights carry a CC BY-NC 4.0 non-commercial license. H Company also open-sourced every trajectory behind its benchmark scores.
Computer-use agents usually specialize: a model trained to click through screens is lost when a task needs an API, and a tool-calling model is stuck when the app has no API. H Company's Holo4 is pitched as one model for all of it. The launch numbers are strong for its size, but two details matter more for anyone planning to use it than the headline score: who ran the comparisons, and what the license allows.
What H Company released
- Holo4 27B (dense, built on Qwen3.8 27B) and Holo4 35B-A3B (mixture of experts, built on Qwen3.6 35B-A3B).
- Holotron4 Nano, a follow-up to Holotron 3 built on NVIDIA's Nemotron 3 Nano Omni.
- One interface for GUIs, code execution, MCP, and APIs, on desktops, the web, Android, code sandboxes, and business APIs, called the same way in each case.
- Weights on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF, plus access through the H Models API. Faster "DSpark" drafter checkpoints are promised "in the coming days".
The scores, and who ran them
| Model | OSWorld 2.0 | Source of the number |
|---|---|---|
| Opus 5.5 | 81.8% | H Company citing OpenAI launch chart data |
| Opus 5 | 70.2% | OpenAI launch chart, max-effort partial rewards |
| GPT-5.6 Sol | 66.2% | OpenAI launch chart, max-effort partial rewards |
| Holo4 27B | 61.7% | H Company's own harness |
| Holo4 35B-A3B | 30.9% | H Company's own harness |
H Company's main claim is cost per task, not raw score: it says Holo4 competes with frontier models on OSWorld 2.0 and AutomationBench at much lower cost. Its own notes set limits on that claim. Holo4's AutomationBench results come from the public task set in H Company's internal harness, while the other models' cost figures come from the official leaderboard, which runs on a private set; H Company says it will report private-set results once they're evaluated. Holo4's costs are priced at H Models API rates, and those rates haven't been published.
The part most coverage skipped: the license
Holo4 is being described as an open-weight release. The weights are downloadable, but the Holo4-27B model card says they are available under a Creative Commons Attribution-NonCommercial 4.0 license, even though the Qwen base model underneath is Apache-licensed. For a company that wants to run the model in its own product, that means self-hosting is not an option without a separate arrangement, and the practical route is the paid API.
How it was built
- An internal Agentic Task Factory builds interactive environments and verifiable tasks from documentation and screenshots; H Company says it has produced about 10,000 tasks across web apps, MCP servers, and desktop environments.
- A rebuilt training harness, redesigned from OSWorld 2.0 failures, added memory that tracks state over hundreds of steps and a shell on the desktop machine.
- In published side-by-side runs, Holo4 27B built a Pac-Man-style game in Godot in 68 calls and 2.4M tokens, where its Qwen base took 197 calls and 11.4M tokens. These are H Company's own selected examples.
What this means if you're choosing an agent model
Holo4 is worth testing if you need one model that can move between a screen and an API in the same workflow, and the published trajectories let you check how it actually completed tasks rather than trusting a chart. Budget for the API rather than self-hosting unless you negotiate a license, wait for published pricing before comparing costs, and remember the 20-point OSWorld 2.0 gap to Opus 5.5: on long workflows, a cheaper model that fails more often can end up costing more.
Sources
AI Industry Reporter
Priya covers model releases, industry announcements, and the gap between what labs claim and what independent evaluators actually find. She reads the primary source - the paper, the system card, the benchmark org's own statement - before writing a word.
More on AI Automation & Agents
Manus 2.0 Launches Cascade, Studio, and Cue, an App That Gives Personal Agents Their Own Phone Numbers and Wallets
Priya Nair · 5 min
Meta Enterprise Platform: What Meta Will Actually Sell Businesses, and What It Hasn't Said Yet
Priya Nair · 5 min
Nvidia's Open Agent Safety Platform: What You Can Install Today, What Needs BlueField-4, and How It Would Stop a Rogue Agent
Priya Nair · 5 min