Click a screen to replay the run
Holo4: powering generalist computer-use agents
Click, code and call tools across desktop, web and mobile. Holo4 27B scores 85.2% on OSWorld at $0.08 per task.
Research · September 28, 2026 · 5 min read
Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.
Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.
Models built for every interface
Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.
Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.
| Benchmark | Holo41 27B | Holo41 35B-A3B | Qwen3.8 27B | Frontier2 |
|---|---|---|---|---|
| OSWorldComputer tasks | 85.2%$0.08 | 80.8%$0.05 | 84.3%$0.22 |
|
| OSWorld 2.0Long computer workflowsScore, success, cost | 61.7%41.5%$1.22 | 30.9%12.3%$0.61 | 48.0%19.4%$3.49 |
|
| ALE-CLI3Expert workflowsScore, success, cost | 44.1%19.4%$0.82 | 30.9%13.5%$0.29 | 43.5%19.0%– |
|
| AutomationBench4Business automation | 45.4%$0.05 | 34.5%$0.02 | 40.3%*$0.09 |
|
| AndroidWorldPhone apps | 85.1%$0.08 | 77.6%$0.07 | 81.9%$0.13 |
|
| Agentic Task Factory5 | ||||
| Web47 web apps | 80.2% | 70.8% | 77.6%* | – |
| MCP14 tool servers | 89.4% | 85.4% | 74.2%* | – |
| Desktop17 desktop apps | 72.0% | 64.8% | 68.9%* | – |
- 1Holo4 in our harness: mean over 2 to 4 runs, single run on OSWorld 2.0 and ALE-CLI.
- 2Public scores, not run by us, across different harnesses and effort levels. OSWorld: Fable 5 from the official leaderboard, GPT-5.5 and Qwen3.8 Max as reported by OpenAI and Qwen.
- OSWorld 2.0: Opus 5.5 at max effort in Anthropic's harness, GPT-6 Astra at max effort on the 82-task offline subset, Muse Spark 1.3 as reported by Meta. Costs read off the providers' charts.
- AndroidWorld: all three as measured by Qwen for the Qwen3.8 release. AutomationBench: public-set scores from the AutomationBench README, costs from the official leaderboard, which runs on the private set.
- 3ALE-CLI: the 105-task Linux split of Agents' Last Exam, average score and pass rate. Qwen3.8 27B and frontier scores and costs from the official leaderboard, at each model's best effort setting. Holo4 excludes attempts that reached reference answers left on the test machine: 2 of 105 for 27B, 1 for 35B-A3B.
- 4AutomationBench: score on the 600 public v1.0.6 tasks; 480 of them fall in the split we collected training data from, the other 120 are held out.
- Score on the 120 held-out tasks: 49.3% for Holo4 27B and 31.7% for Holo4 35B-A3B, against 40.3% for Qwen3.8 27B and 13.1% for Qwen3.6 35B-A3B. We report the public-set score until Holo4 is evaluated on the official private set.
- 5Generated by our Agentic Task Factory and not used in training, pooling in-domain and out-of-domain test sets.
- *Qwen3.8 27B in our harness, where no public score exists.
- Under each score: cost per task in USD, at H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen3.8 27B, from the tokens of our runs.
Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.
Competitive with the frontier, at a fraction of the cost
On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.
- Holo
- Base model
- Reported
- Closed frontier
AI that does work
Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model.
84 calls · 1.3M tokens
Same prompt and harness for both models.
How we built Holo4
Agentic Task Factory
Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.
Flow diagram of the Agentic Task Factory, listed in order from sources to output.
Sources
- Product docsmanuals · help centers
- Screenshotsof real websites
- Real softwareself-hosted · desktop · OS
Agentic Task Factory
Build
- environment
- tasks
- verifiers
Gates
A task is kept only if its verifier
- fails on the untouched seed
- passes on the golden state
- rejects every near miss
and an agent solved it through the real interface.
Failed attempts loop back to Build: attempt, audit, harden.
Runtime
1 Docker image
- human
- noVNC
- agent
- CDP / MCP
- score
- state delta or tool trace
Output
10,000 tasks
Training
- 01
Supervised fine-tuning
127B tokens. About three quarters are successful agentic trajectories from our Agentic Task Factory: desktop (45%), web (14%), MCP and API (12%) and mobile (3%). The rest covers multimodal reasoning, GUI grounding, and text-only tool use and coding.
- 02
Two RL experts
Asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model: one for desktop and web, one for terminal, MCP and API.
- 03
One merged model
Both experts merge back into the fine-tuned model with equal weight and no further training. The result combines the general skills from supervised fine-tuning with each expert's specialized skills.
Harness
Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.
- Run the shell on the desktop machine itself
- Context passthrough during memory compaction
- Scale model serving; fix grading setup and adapter crashes
- New system prompt; tool and environment fixes (200 → 300 steps · 2 → 4 h)
- Fix a vision bug in serving; trim the prompt (300 → 500 steps · 4 → 6 h)
- Drive the desktop from the shell: one tool for code and clicks
- Holo4 27B release 61.7%
- Reported scores, top to bottom: Opus 5.5 81.8% · Opus 5 70.2% · GPT-5.6 Sol 66.2% · Opus 4.8 54.8% · Qwen3.8 27B 48.0% · Sonnet 4.6 41.5%
- Opus 5 (70.2%) and GPT-5.6 Sol (66.2%) use max-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart, as in the cost-performance plot. Other reference scores come from model cards and the official leaderboard. Task releases, subsets and harnesses vary.
Holotron4 Nano
Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.
The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.
Holotron benchmarks
OSWorldCUA GUI
OSWorld 2.0CUA + Code
AutomationBenchMCP
PinchBenchTerminal
ALE (Linux, Code)Terminal
Gains are absolute percentage-point improvements over Nemotron 3 Nano Omni.
These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.
Run it yourself
Both sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.
We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.
New
Flagship
Holo4 27B
Dense, 27B
Best accuracy on long, multi-step tasks across web, desktop, and mobile.
- Context
- 256K
- Input / 1M · $0.04 cached
- $0.40
- Output / 1M
- $3.00
New
Fast
Holo4 35B-A3B
MoE, 35B total, 3B active
Close to 27B accuracy, cheaper and faster. Open weights under Apache 2.0.
- Context
- 256K
- Input / 1M · $0.03 cached
- $0.30
- Output / 1M
- $2.00
Get started with Holo4
Open weights, a hosted API, and every run behind the numbers.