AI Models

Holo4: An Open Model Built to Use Your Computer, Not Just Chat

Holo4 is a new computer-use VLM family built to click, type, run code, and call tools. Here is what the 27B and 35B-A3B models are actually for.

Approximately 5 min read

What is Holo4 for?

Holo4 is a vision-language model family designed to operate software, not just talk about it.

With an agent harness around the model, Holo4 can look at screenshots, decide what to do next, click and type in graphical interfaces, write or execute code, and call structured tools such as MCP or APIs.

That makes Holo4 a useful example of a model category that is increasingly important in 2026: computer-use models whose output is an action, not merely an answer.

H Company released Holo4 on September 28, 2026. The family currently includes a 27B dense model and a 35B-A3B Mixture-of-Experts model, with multiple precision and quantized variants.

The useful mental model: Holo4 is an action model

A conventional multimodal model can inspect a screenshot and explain what it sees.

A computer-use model has to go further.

The working loop looks more like this:

screenshot or tool result
        ↓
      Holo4
        ↓
click / type / code / tool call
        ↓
new application state
        ↓
      Holo4

H Company’s open hai-agents harness provides the execution layer around the model. It sends screenshots and tool results to Holo4, executes requested actions, then returns the new state to the model.

The important distinction is that Holo4 is not itself a complete desktop automation platform.

The checkpoint is the decision-making model. A useful computer agent still needs an environment that can capture application state, expose tools, execute actions, manage permissions, and return results safely.

Two main Holo4 models

Holo4-27B

Holo4-27B is a 27B dense vision-language model built on the Qwen3.8-27B architecture.

Its published configuration supports a maximum context length of 262,144 tokens.

H Company distributes the family in BF16, FP8, NVFP4, and Q4 GGUF forms. The official Q4_K_M GGUF is about 16.9 GB, which makes the weights substantially more approachable for local experimentation than the original BF16 checkpoint.

The important licensing limitation is that the 27B weights use CC BY-NC 4.0, so they are non-commercial.

Holo4-35B-A3B

Holo4-35B-A3B uses the Qwen3.6-35B-A3B MoE architecture.

The model has about 35B total parameters while using an A3B-style sparse architecture, so only a smaller portion of the network is active for each token.

This version is also published in BF16, FP8, NVFP4, and GGUF variants.

Unlike the 27B checkpoint, the official Holo4-35B-A3B weights are released under the Apache License 2.0.

That difference matters for anyone evaluating Holo4 for commercial or enterprise experimentation.

What would you actually use it for?

Holo4 is most interesting when a workflow crosses interfaces.

A normal automation system prefers APIs because APIs are deterministic and structured. But many useful applications still expose important operations only through a browser or desktop interface.

A computer-use agent can combine both approaches.

For example:

That makes Holo4 relevant to browser automation, desktop automation, agentic RPA, software testing, engineering tools, and workflows that span multiple applications.

H Company demonstrates Holo4 operating applications including FreeCAD and Godot. Those examples are useful because the tasks require long sequences of stateful interactions rather than a single screenshot question.

Why this is different from another VLM

The interesting part of Holo4 is not that it can see a screen.

Many models can do that.

The important capability is choosing between different kinds of actions.

A robust software agent should not click through ten menus if it can make one API call. It should not call an API if the needed function exists only in a graphical application. It should be able to move between both.

Holo4 is trained around that hybrid interface model: GUI interaction, code execution, and structured tool calls are all possible actions in the same workflow.

That is closer to how practical automation systems have to operate.

How good is it?

H Company reports 61.7% on OSWorld 2.0 for Holo4-27B and 30.9% for Holo4-35B-A3B.

The company also reports results on AutomationBench and its own held-out Agentic Task Factory workflows, and publishes evaluation trajectories so individual runs can be inspected.

Those numbers should still be treated as vendor-reported benchmark results, not as a universal ranking of computer-use models.

Agent benchmarks are especially sensitive to the execution harness, available tools, environment versions, retry policy, task definitions, and inference settings. H Company also notes that some comparison results come from public leaderboards while others were produced with its own harness.

For this model category, the harness is part of the system.

Can you run Holo4 locally?

The weights are open and downloadable, and H Company publishes quantized variants including GGUF.

That does not mean loading Holo4 into an ordinary chat frontend gives you a working computer agent.

For real computer use, you still need an execution layer that can:

The official model cards point users toward H Company’s hai-agents tooling for this loop.

So the practical architecture is:

Holo4 model
    +
agent harness
    +
browser / desktop / mobile environment
    +
permission boundaries
    =
computer-use agent

That final line is important. The model is the reasoning and action-selection component, not the entire product.

Who should care?

Holo4 is worth watching if you work on:

It is much less compelling if all you need is a conversational local LLM.

There are simpler models for that job.

Holo4 exists for a different reason: it is designed to do work across software interfaces instead of only explaining what a user should do next.

That is what makes it a useful model to watch.

Sources and further reading

Continue reading