Holo4: The Generalist Computer-Use Agent Bridging GUI, Code, and APIs
Holo4 登場:融合 GUI、程式碼與 API 的全能型電腦操作 AI 代理

H company has launched Holo4 (available in 27B dense and 35B MoE) and Holotron4 Nano. Unlike traditional single-interface agents, Holo4 natively combines GUI navigation, code execution, MCP, and API calls across desktops, web, Android, and sandboxes. It achieves highly competitive performance on OSWorld 2.0 and AutomationBench at a fraction of the cost of frontier closed-source models.
Key points
Multi-Interface Integration
Bridges the gap between GUI and API, enabling a single model to click screens, write code, and call MCP/APIs dynamically based on the task.
Built for Real Workflows
Trained on over 10,000 real-world tasks generated by the Agentic Task Factory, ensuring viability for actual business automation.
Exceptional Cost-Efficiency
Scores 61.7% on OSWorld 2.0, competing closely with frontier closed models while using orders of magnitude fewer parameters and cost.
Upgraded Execution Harness
Reengineered execution loop provides the agent with reliable memory for hundreds of steps alongside a native desktop shell.
How it works
Why it matters
Most agentic models are siloed into either GUI-only or API-only interfaces. Holo4 breaks this barrier by serving as a true generalist computer-use agent. By performing complex cross-platform tasks at a lower cost, it makes enterprise automation highly practical. It also demonstrates a transferrable training recipe that can turn smaller models into highly capable agents.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Enterprise Leader
How to use it
- 1Automated 3D CAD modeling (e.g., using FreeCAD to build precise parts or logos from text prompts)
- 2Autonomous software development and playtesting (e.g., building and running games independently in Godot)
- 3Hybrid business workflows (simultaneously interacting with web GUIs, query databases via MCP, and running backend APIs)
Limitations & caveats
- Performance on extremely long workflows (61.7% on OSWorld 2.0) still trails frontier closed models like Opus 5.5 (81.8%)
- Performance on the private evaluation sets of certain API benchmarks (e.g., AutomationBench) is still pending official run verification
Related

The Agent Said It Was Done, the Database Disagreed: Introducing ThinkingBox
為什麼 Agent 的對話會騙人?微軟與 HF 推出 ThinkingBox 評測真實資料庫狀態
While traditional benchmarks only grade tool calls, ThinkingBox evaluates agents based on their actual terminal backend states and unintended side effects.

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料
AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸
KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.