Google Announces Gemini 4 Argon: Frontier Model with 1M Output Tokens and Long-Horizon Reasoning
Google 發表新一代前沿模型 Gemini 4 Argon:具備 100 萬 Token 輸出與強大自主 Agent 推理能力

Google DeepMind has introduced Gemini 4 Argon, initially rolling out to trusted cyber defenders. Argon expands the output token limit from 64K to 1 million, giving the model the capacity to generate deep, multi-step reasoning trajectories. Google has deployed Argon internally to automate large-scale codebase migrations (up to 800K+ lines of C/C++ to Rust) and optimize data center memory by over 300 TiB. It achieves state-of-the-art results on benchmarks like DeepSWE v1.1 and LVBench, while introducing rigorous mitigations against alignment risks and prompt injections.
Key points
1M Output Token Capacity
Expands output capacity to 1 million tokens, allowing the model to generate extensive reasoning paths to solve complex problems in one go.
Large-Scale Code Migration
Automates refactoring of up to 800K+ lines of code to Rust inside Google, achieving a state-of-the-art 77.9% on DeepSWE v1.1.
Autonomous Cyber Defense
Trained for proactive defense, allowing trusted partners to autonomously discover and patch vulnerabilities, tying for first on CWE-bench v1.
Active Alignment Monitoring
Deploys alignment mitigations that actively monitor the model's chain-of-thought and actions, halting execution if it deviates from user intent.
How it works
Why it matters
This model marks a significant shift from conversational AI to highly autonomous, long-horizon Agents. The 1M output token capacity enables single-trajectory resolution of vast tasks like kernel-level code migration. Crucially, Google's deployment of chain-of-thought monitoring and hardened sandbox execution establishes a practical safety blueprint for aligning highly capable future AI systems.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Policy Maker
How to use it
- 1Large-scale codebase migrations, such as refactoring legacy C/C++ libraries into memory-safe Rust code.
- 2Autonomous cyber vulnerability remediation, identifying and patching exposures via black-box penetration testing.
- 3Complex, multi-step enterprise tasks, including long-video analysis, multi-step financial research, and legal drafting.
Limitations & caveats
- Due to advanced agent capabilities, high-risk training and evaluation require isolated, sealed sandbox environments.
- Code migrations for mission-critical systems still require manual auditing and emulation testing before production release.
Related
Thinking Before Thinking: Scaling AI Agents with Meta-Reasoning
後設推理:讓 AI 代理在動手前先「思考如何思考」
This paper introduces agentic meta-reasoning, a framework using a controller to dynamically allocate compute budget, improving long-horizon task execution.
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
以「後設技能」為核心的 AI4AI:為 Agent 設計最佳執行環境的測試期學習架構
This study introduces a test-time AI-for-AI framework where a Builder model learns Meta-Skills to design optimal execution environments (harnesses) for a target agent without training weights, boosting performance on unseen tasks.
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
This paper introduces TokenCast, a real-time method that accurately forecasts dynamic LLM Agent token consumption within 32.8 ms without extra LLM calls, using composable cost representations.