Holo4 agent release splits the license and the score
The Holo4 agent release is two checkpoints with two licenses and two very different scores. The company post supplies the headline numbers. OpenTools reads the model cards and refuses to treat those numbers as a bake-off.

in this block
The Holo4 agent release landed in public on September 28, 2026, as a pair of vision-language agents from H Company. The Hugging Face post says both sizes, a 27B dense model and a 35B-A3B mixture of experts, are on the H Models API, with weights in BF16, FP8, NVFP4, and 4-bit GGUF, plus a smaller Holotron4 Nano built on Nemotron 3 Nano Omni. OpenTools, which says it read the launch post and the model cards on September 29, adds the split that the launch prose does not: the 27B card is CC BY-NC 4.0, and the 35B-A3B card is Apache 2.0.
TL;DR - On OSWorld 2.0 the company post gives Holo4 27B a score of 61.7% and Holo4 35B-A3B a score of 30.9%, against 81.8% for Opus 5.5. OpenTools calls the Holo4 figures partial-reward scores and attaches mean model costs of $1.22 and $0.61 per task. - OpenTools says the 27B weights are noncommercial under CC BY-NC 4.0, while the lower-scoring 35B-A3B weights are Apache 2.0. A hosted API contract is a third document neither page reprints. - AutomationBench numbers in OpenTools are company-run on a public set: 45.4% and 34.5%, with a held-out slice of 49.3% and 31.7%. H says it has not yet reported the private leaderboard set. Do not line those percents up against private-set scores.
What actually happened
H's post describes one model family that is supposed to click and type on a screen, write and run code, and call MCP or API tools, on desktops, the web, Android, a code sandbox, and business APIs. The post says most agents are trained for one interface. Holo4, it says, was trained with supervised and reinforcement learning, including tasks from an internal factory it counts at about 10,000.
The score line everyone will screenshot is OSWorld 2.0. H writes that Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and that Holo4 35B-A3B reaches 30.9%, "with orders of magnitude fewer parameters and at a much lower cost." The same post warns that releases, harnesses, and task subsets differ across the comparison chart. It says Holo4's own points use H Models API rates, that other points come from a launch chart or the official leaderboard, and that Holo4 was left off a line drawn through closed models. The chart is not one experiment. OpenTools sharpens it. It calls 61.7% and 30.9% partial-reward scores, puts the mean model cost at $1.22 and $0.61 per task, and notes that OSWorld 2.0 has 108 long tasks whose site says a skilled person needs a median of about 1.6 hours. Partial credit is not a finished job. OpenTools also says H's public trajectory card lists 106 OSWorld 2.0 runs for each checkpoint, two fewer than 108, and that the sources it kept do not say whether those two were dropped from the score or only from the traces.
A second chart in the company post mentions Opus 5 at 70.2% and GPT-5.6 Sol at 66.2%, and it says those two use max-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart. Those are not the 81.8% Opus 5.5 figure. Mixing Opus 5 and Opus 5.5 because the names look related is how a screenshot lies. This article keeps them in separate sentences.
OpenTools is also where the license lives, because the launch post opened here does not state one. It says the Holo4-27B model card labels the weights CC BY-NC 4.0, reuse for noncommercial purposes under that license's conditions, and the Holo4-35B-A3B card labels its weights Apache 2.0. Both cards, OpenTools says, list a maximum configured context of 262,144 tokens, which is a configuration limit and not evidence the model stays reliable across that whole window. The higher score and the permissive license are on different checkpoints. That is the actual product decision in the Holo4 agent release, and it is easy to miss if you only read the 61.7% line.
What the benchmark footnotes change
AutomationBench is the other number, and it is mostly OpenTools, because the company post describes the chart without printing the percents in the text that loaded. OpenTools says H reports strict public-set pass rates of 45.4% for the 27B and 34.5% for the 35B-A3B on AutomationBench v1.0.6, at $0.05 and $0.02 per task, measured in H's harness. It says 480 of 600 public tasks belong to a split H collected training data from. On the 120-task slice H says it held out, the figures are 49.3% and 31.7%, against 40.3% for Qwen3.8-27B and 13.1% for Qwen3.6-35B-A3B in that same harness. Held-out still favors H's post-training over the stated bases. It is still H's measurement. OpenTools adds that the public set and the private set behind the official leaderboard are separate, and that H says it has not reported Holo4 on the private set. Putting 45.4% next to a private-leaderboard number from another lab would invent a comparison neither page makes.
The traces are the part that is genuinely more open than a screenshot. OpenTools says the trajectories dataset lists 7,366 agent runs across OSWorld, OSWorld 2.0, AndroidWorld, AutomationBench, PinchBench, and Agents' Last Exam, and that the dataset is Apache 2.0 while upstream task content keeps its own license. It also says credentials and personal data are masked, some screenshots are replaced, and a few tasks are omitted. Public traces make a self-run auditable. They do not make it independent. The company post says you can replay steps at trajectories.hcompany.ai or download them from Hugging Face. This article did not replay them.
Holotron4 Nano is a transfer claim. The company says the same stack on Nemotron 3 Nano Omni improves GUI and tool workflows, and that the loaded text's point gains were in a figure this article will not invent. DSpark drafter checkpoints were still described as coming later. That is a plan, not a file.
What a reader should not collapse
The Holo4 agent release is not "open weights beat Opus." Read the Holo4 agent release as a pair of checkpoints, not as one brand. The company's own OSWorld line has the 27B at 61.7% and Opus 5.5 at 81.8%, and OpenTools says the 61.7% is partial reward from H's harness. The release is also not "the good one is Apache." OpenTools puts Apache 2.0 on the checkpoint with the 30.9% score. If your constraint is commercial self-hosting, you are looking at the weaker published OSWorld number. If your constraint is the higher published number, you are looking at a noncommercial weight license, and you still have to read the API terms before you treat the hosted button as the same permission.
Context length is the other collapse. 262,144 tokens on a model card, as OpenTools describes it, is a maximum configuration. It is not a result. Cost is a third. A lower dollar figure per task at $0.61 versus $1.22 is H's mean model cost on its own runs. It does not include a human cleaning up a half-finished desktop. OpenTools says that cleanup can dominate the token bill. That is a warning, not a measured invoice.
For a closed model these charts keep using as a yardstick, see the Sonnet 5.5 note. For another name that shows up in the partial-reward footnote, see the GPT-6 Sol note. Neither note is a license for Holo4, and neither reruns OSWorld.
What to do as a reader (not a trade)
This is not investment advice. There is no token in the two pages. If you are choosing a checkpoint, write two columns. Column one is the score you actually need, partial or strict, on a task set you froze yourself. Column two is the license on that exact repository, not the family name. The Holo4 agent release gives you a reason to test, which is all OpenTools claims for it. It does not give you a reason to skip the test because 61.7% looks close to a frontier name from across a chart H already marked as mixed.
Pin the quantization, the harness, and the step limit if you do run it. The company post says task subsets differ. Your rerun will not be their chart unless you copy more than the weights. The primary page is H's launch post. The second page is OpenTools' license and score guide, which is the one that separates noncommercial weights from the Apache sibling and partial credit from a finished task.
Not financial advice. DYOR, ser.