Introducing Holo4: powering generalist computer-use agents

Click a screen to replay the run

Newsroom

Holo4: powering generalist computer-use agents

Click, code and call tools across desktop, web and mobile. Holo4 27B scores 85.2% on OSWorld at $0.08 per task.

Research · September 28, 2026 · 5 min read

Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.

Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.

Models built for every interface

Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.

Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.

BenchmarkHolo41
27B
Holo41
35B-A3B
Qwen3.8
27B
Frontier2
OSWorld85.2%$0.0880.8%$0.0584.3%$0.22
  • Fable 586.0%
  • GPT-5.578.7%
  • Qwen3.8 Max86.1%
OSWorld 2.0Score, success, cost61.7%41.5%$1.2230.9%12.3%$0.6148.0%19.4%$3.49
  • Opus 5.581.8%48.7%$8.48
  • GPT-6 Astra73.5%–$9.07
  • Muse Spark 1.366.9%32.0%–
ALE-CLI3Score, success, cost44.1%19.4%$0.8230.9%13.5%$0.2943.5%19.0%–
  • Opus 5.563.7%34.3%$8.22
  • GPT-6 Astra61.4%33.3%$5.31
  • Muse Spark 1.357.5%33.3%$2.44
AutomationBench445.4%$0.0534.5%$0.0240.3%*$0.09
  • Opus 550.3%$3.05
  • GPT-5.6 Sol45.8%$0.67
  • Kimi K346.7%$0.43
AndroidWorld85.1%$0.0877.6%$0.0781.9%$0.13
  • Fable 588.8%
  • GPT-5.6 Sol77.6%
  • Qwen3.8 Max85.3%
Agentic Task Factory5
Web80.2%70.8%77.6%*–
MCP89.4%85.4%74.2%*–
Desktop72.0%64.8%68.9%*–
  1. 1Holo4 in our harness: mean over 2 to 4 runs, single run on OSWorld 2.0 and ALE-CLI.
  2. 2Public scores, not run by us, across different harnesses and effort levels. OSWorld: Fable 5 from the official leaderboard, GPT-5.5 and Qwen3.8 Max as reported by OpenAI and Qwen.
  3. OSWorld 2.0: Opus 5.5 at max effort in Anthropic's harness, GPT-6 Astra at max effort on the 82-task offline subset, Muse Spark 1.3 as reported by Meta. Costs read off the providers' charts.
  4. AndroidWorld: all three as measured by Qwen for the Qwen3.8 release. AutomationBench: public-set scores from the AutomationBench README, costs from the official leaderboard, which runs on the private set.
  5. 3ALE-CLI: the 105-task Linux split of Agents' Last Exam, average score and pass rate. Qwen3.8 27B and frontier scores and costs from the official leaderboard, at each model's best effort setting. Holo4 excludes attempts that reached reference answers left on the test machine: 2 of 105 for 27B, 1 for 35B-A3B.
  6. 4AutomationBench: score on the 600 public v1.0.6 tasks; 480 of them fall in the split we collected training data from, the other 120 are held out.
  7. Score on the 120 held-out tasks: 49.3% for Holo4 27B and 31.7% for Holo4 35B-A3B, against 40.3% for Qwen3.8 27B and 13.1% for Qwen3.6 35B-A3B. We report the public-set score until Holo4 is evaluated on the official private set.
  8. 5Generated by our Agentic Task Factory and not used in training, pooling in-domain and out-of-domain test sets.
  9. *Qwen3.8 27B in our harness, where no public score exists.
  10. Under each score: cost per task in USD, at H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen3.8 27B, from the tokens of our runs.

Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.

Competitive with the frontier, at a fraction of the cost

On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.

OSWorld 2.0Average partial score (%) · higher is better
  • Holo
  • Base model
  • Reported
  • Closed frontier
0%5%10%15%20%25%30%35%40%45%50%55%60%65%70%75%80%$0.01$0.03$0.10$0.30$1$3$10$30$100Holo4 27BHolo4 35B-A3BGPT-6 Astra · maxOpus 5 · maxGPT-6 Luna · maxQwen3.8 MaxSonnet 4.6 · maxQwen3.6 35B-A3BKimi K2.6 · 500 stepsQwen3.7-Plus · thinking
← Lower costUSD per task · log scale

AI that does work

Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model.

84 calls · 1.3M tokens

Same prompt and harness for both models.

Prompt

Build a 3D model of the Eiffel Tower in FreeCAD, at a scale of 1 mm to 1 metre, centred on the origin and aligned to the X and Y axes. Work to this design. The tower is square in plan at every height, never round. Its half-width, measured from the central axis out to the corner, is 62.5 mm at ground level, 32.5 mm at height 57, 17.5 mm at height 115, and 9.35 mm at height 276. Between those heights the half-width follows a smooth curve that falls steeply near the ground and gently higher up, never a straight line. Four identical legs, one per quadrant, each a square column whose outer corner follows that profile. Each leg is 14 mm across at the ground and tapers to 4 mm at height 276. The legs stand apart from the ground up to the first platform, and converge as they rise. Nothing fills the space between them: the tower is open, and you can see straight through it from every side. Three platforms, each a solid square slab centred on the axis: 72 mm across and 4 mm thick at height 57; 40 mm across and 3 mm thick at height 115; 22 mm across and 3 mm thick at height 276. A mast from height 276 to 324, square, 8 mm across at its base tapering to 2 mm at the tip. Every part must be a closed solid with non-zero volume, and no part may fill the space between the legs.

How we built Holo4

Agentic Task Factory

Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

Flow diagram of the Agentic Task Factory, listed in order from sources to output.

Sources

  • Product docsmanuals · help centers
  • Screenshotsof real websites
  • Real softwareself-hosted · desktop · OS

Agentic Task Factory

Build

  • environment
  • tasks
  • verifiers

Gates

A task is kept only if its verifier

  • fails on the untouched seed
  • passes on the golden state
  • rejects every near miss

and an agent solved it through the real interface.

Failed attempts loop back to Build: attempt, audit, harden.

Runtime

1 Docker image

human
noVNC
agent
CDP / MCP
score
state delta or tool trace

Output

10,000 tasks

4kweb apps
3kMCP servers
3kdesktop · OS

Training

  1. 01

    Supervised fine-tuning

    127B tokens. About three quarters are successful agentic trajectories from our Agentic Task Factory: desktop (45%), web (14%), MCP and API (12%) and mobile (3%). The rest covers multimodal reasoning, GUI grounding, and text-only tool use and coding.

  2. 02

    Two RL experts

    Asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model: one for desktop and web, one for terminal, MCP and API.

  3. 03

    One merged model

    Both experts merge back into the fine-tuned model with equal weight and no further training. The result combines the general skills from supervised fine-tuning with each expert's specialized skills.

Harness

Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.

Automatic harness engineeringOSWorld 2.0 average partial score (%)
0%20%40%60%80%Aug 1Aug 15Sep 1Sep 15Opus 5.5, 81.8%, Anthropic system card, max effort, 5 runsOpus 5.581.8%Opus 5, 70.2%, OpenAI GPT-6 Sol and Luna launch chart, max effort, partial reward, v2026.08.08 offline setOpus 570.2%GPT-5.6 Sol, 66.2%, OpenAI GPT-6 Sol and Luna launch chart, max effort, partial reward, v2026.08.08 offline setGPT-5.6 Sol66.2%Opus 4.8, 54.8%, Official OSWorld 2.0 leaderboard, max effortOpus 4.854.8%Qwen3.8 27B, 48.0%, Qwen model cardQwen3.8 27B48.0%Sonnet 4.6, 41.5%, Official OSWorld 2.0 leaderboard, max effortSonnet 4.641.5%Run the shell on thedesktop machine itselfContext passthrough duringmemory compactionScale model serving; fix gradingsetup and adapter crashesNew system prompt; tool andenvironment fixes(200 → 300 steps · 2 → 4 h)Fix a vision bug in serving;trim the prompt(300 → 500 steps · 4 → 6 h)Drive the desktop from the shell:one tool for code and clicksHolo4 27B release61.7%
  1. Run the shell on the desktop machine itself
  2. Context passthrough during memory compaction
  3. Scale model serving; fix grading setup and adapter crashes
  4. New system prompt; tool and environment fixes (200 → 300 steps · 2 → 4 h)
  5. Fix a vision bug in serving; trim the prompt (300 → 500 steps · 4 → 6 h)
  6. Drive the desktop from the shell: one tool for code and clicks
  7. Holo4 27B release 61.7%
  8. Reported scores, top to bottom: Opus 5.5 81.8% · Opus 5 70.2% · GPT-5.6 Sol 66.2% · Opus 4.8 54.8% · Qwen3.8 27B 48.0% · Sonnet 4.6 41.5%
  • Opus 5 (70.2%) and GPT-5.6 Sol (66.2%) use max-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart, as in the cost-performance plot. Other reference scores come from model cards and the official leaderboard. Task releases, subsets and harnesses vary.

Holotron4 Nano

Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.

The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.

Holotron benchmarks

Nemotron 3 Nano OmniHolotron4 Nano

OSWorldCUA GUI

21.0 →76.3+55.3

OSWorld 2.0CUA + Code

0.2 →7.9+7.7

AutomationBenchMCP

19.4 →35.6+16.2

PinchBenchTerminal

84.7 →88.6+3.9

ALE (Linux, Code)Terminal

0.6 →8.5+7.9

Gains are absolute percentage-point improvements over Nemotron 3 Nano Omni.

These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.

Run it yourself

Both sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.

We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.

New

Flagship

Holo4 27B

Dense, 27B

Best accuracy on long, multi-step tasks across web, desktop, and mobile.

85.2%on OSWorld
Context
256K
Input / 1M$0.04 cached
$0.40
Output / 1M
$3.00

New

Fast

Holo4 35B-A3B

MoE, 35B total, 3B active

Close to 27B accuracy, cheaper and faster. Open weights under Apache 2.0.

80.8%on OSWorld
Context
256K
Input / 1M$0.03 cached
$0.30
Output / 1M
$2.00