Trending: On-device modelsSearch
iHeartGeek
iTECH

Holo4 agents click, type and write their own code

H company's Holo4 models drive software through a desktop, code, MCP tools or an API without being told which, and the company has published the benchmark trajectories behind its scores alongside open FP16, FP8 and GGUF weights.

An announcement graphic reading Introducing Holo4 on a cream background, with a black wireframe globe orbited by a lime green and a violet marker, and the H company logo in the lower left

H company has released Holo4, a new family of agentic models that drive software through whatever interface is available, whether that is a graphical desktop, code, Model Context Protocol tools or a plain API. It comes in two sizes, a 27B dense model and a 35B-A3B mixture of experts, both available through the H Models API, alongside an updated Holotron 4 Nano. The weights are open, released in FP16, FP8 and GGUF formats on Hugging Face, and the company has published the trajectories behind its benchmark scores so the runs can be replayed step by step.

The interface argument is the point of the release. Most agentic models are trained for one surface: a model that lives behind a screen is useless when the task needs a shell, and a model that only calls tools is stuck in front of software that has no API. H company says a single real business task can need both, so Holo4 uses whichever fits, and the same model is invoked the same way whether it is running on a desktop, in a browser, on Android, inside a code sandbox or against a business API.

What the numbers actually say

On OSWorld 2.0, the hardest academic benchmark for desktop control, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, while the smaller 35B-A3B variant reaches 30.9%. That is a real gap to the frontier, and H company says so rather than burying it. The counter-argument it makes is cost: the models trail the strongest closed systems on long workflows while running orders of magnitude fewer parameters, and the company publishes cost-per-task comparisons against its Qwen base models on both OSWorld 2.0 and AutomationBench, the benchmark it uses for API work.

The methodology is disclosed in unusual detail for a model launch, including the caveats. Costs on OSWorld 2.0 are estimated from the input and output tokens of each agentic run, the Qwen figures come from Alibaba Cloud list prices, the closed-model points are drawn from the official leaderboard, and H company notes that releases, harnesses and task subsets differ between entries. AutomationBench scores for rivals are taken from the public set while the leaderboard runs the private set, so Holo4 will be reported on the private set only once it is evaluated.

Trained on work, not on demos

Holo4 was trained with supervised and reinforcement learning across a large set of environments and tasks, including ones generated by the company's own Agentic Task Factory. The published examples are deliberately professional rather than playful: building a scaled model of the Eiffel Tower in FreeCAD to a detailed geometric specification, extruding the H company logo as a 3D solid, and building a playable Pac-Man-style game in Godot with a maze, pellets, scoring and three chasing ghosts. Each example is paired with the same task run on Qwen3.8 27B, the base model, under the same prompt and harness, along with the call and token counts.

The comparison is unflattering to nobody in particular and useful for buyers. On the Eiffel Tower task Holo4 27B took 84 calls and 1.3 million tokens against Qwen's 60 calls and 1.0 million, and on the logo task it took fewer calls, 94 against 118, while using fewer tokens, 1.5 million against 1.9 million. Agentic coding is less about a single clever answer and more about how a model recovers when a step goes wrong, and publishing the trajectories is the only way for anyone outside the lab to check that.

Our opinion

The most useful sentence in the announcement is the one admitting Holo4 trails the strongest closed models on long workflows. Model launches have an unhappy habit of leading with a cherry-picked benchmark, and a 20-point OSWorld deficit against Opus is not a rounding error. What makes it defensible is that the company immediately explains who should care: teams paying per task at volume, for whom a cheaper model that is right most of the time beats an expensive one that is right slightly more often.

The interface-agnostic claim is the part that will decide whether this matters. Agents that can only click, or only call tools, hit a wall as soon as real work involves a legacy desktop application with no API, which in most large organisations is immediately. If Holo4 genuinely switches between GUI, code, MCP and API mid-task without being told, that is a more interesting capability than a few benchmark points. The open trajectories are how that gets checked, and it is good that they are there to check.