AI Daily ·
Google Unveils Gemini 4 Argon and Breakthroughs in Brain-to-Text Decoding
今日 AI 重點
Today's top AI developments highlight Google's announcement of Gemini 4 Argon, a frontier model featuring an unprecedented 1 million output token limit and advanced long-horizon reasoning. Additionally, researchers achieved a breakthrough in non-invasive brain-to-text decoding by eliminating the 'timing shortcut' to significantly reduce word error rates. Other notable advancements include SCAPO for token-level credit optimization and Cogentic for multi-agent automated proof discovery.
01Google DeepMindAI AgentGoogle Announces Gemini 4 Argon: Frontier Model with 1M Output Tokens and Long-Horizon Reasoning
Google DeepMind has introduced Gemini 4 Argon, initially rolling out to trusted cyber defenders. Argon expands the output token limit from 64K to 1 million, giving the model the capacity to generate deep, multi-step reasoning trajectories. Google has deployed Argon internally to automate large-scale codebase migrations (up to 800K+ lines of C/C++ to Rust) and optimize data center memory by over 300 TiB. It achieves state-of-the-art results on benchmarks like DeepSWE v1.1 and LVBench, while introducing rigorous mitigations against alignment risks and prompt injections.
- 02arXivLLM
SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
In reinforcement learning with verifiable rewards (RLVR), standard GRPO assigns a uniform outcome-derived advantage to all tokens, which can mistakenly reinforce spurious prompt features. To address this, researchers developed SCAPO (Semifactual Credit-Augmented Policy Optimization). It utilizes semifactual prompt interventions to calculate token probability drift. By reducing the credit advantage of highly unstable tokens during early training, SCAPO dramatically reduces prompt sensitivity. Evaluation on Qwen3 base models shows significant improvements in AIME benchmarks and out-of-distribution mathematical reasoning tasks.
- 03arXivAI Research
Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
A prominent 2025 non-invasive brain-to-text model by d'Ascoli et al. was found to score 22.0% accuracy on synthetic data containing zero brain signals, compared to 22.3% on real recordings. This occurred because joint encoding of overlapping time windows leaked word durations (timing shortcuts). To resolve this, researchers introduced SimpleB2T, which processes each word's window independently. By blocking this temporal loophole, the model is forced to decode actual neural signals. Combined with pretrained LLMs and multi-observation aggregation, SimpleB2T achieved a breakthrough Word Error Rate (WER) of 36.6% using non-invasive recordings.
- 04arXivAI Research
EvoDuet: Co-Evolving Web Search and Task Solving for LLM-Based Scientific Discovery
Evolutionary search with LLMs often stalls when missing critical external knowledge. Simply adding search tools can lead to redundant page retrievals. EvoDuet addresses this by co-evolving search queries and task solutions under fixed parameters. In each iteration, a retrieval gate assesses the model's knowledge gap, an inner loop refines queries to rank documents, and an outer loop parallel-generates solution candidates. EvoDuet raised Gemini-3.8-Flash's OpenEvolve discovery gain from 61.3% to 82.3%, though smaller models like Qwen3.5-9B showed no improvement.
- 05arXivAI Research
Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
Single-shot LLM generation struggles with open math problems that require long-horizon reasoning and conjecture exploration. To bridge this gap, researchers developed Cogentic, a multi-agent harness powered by Gemini. An orchestrator coordinates independent provers across different mathematical directions. Their outputs undergo adversarial verification by specialized components, and verified steps are saved to a persistent ledger for future rounds. Cogentic successfully solved five open problems in online learning, auction theory, and mechanism design, all verified by experts.
- 06arXivLLM
Scaling Laws for Looped Mixture of Experts: Combining Recurrence and Sparsity
This paper presents the Loop Scaling Laws, the first framework to jointly model recurrence and sparsity alongside model size and data. While recurrence increases computational depth without adding parameters, MoE expands model capacity without increasing active compute. The study shows that sparsity yields ~3x active-parameter efficiency, and recurrence delivers ~2x total-parameter efficiency on reasoning tasks. At a trillion-token scale, a Looped MoE matches the performance of a non-looped MoE twice its size, while enabling dynamic test-time scaling.
07NVIDIA DeveloperAI HardwareNVIDIA Expands AI Storage Acceleration with cuObject and SCADA Server SDK
Traditional storage access paths that copy data through the server's CPU memory limit AI training and inference speeds. NVIDIA is addressing this by expanding xio-sig to include cuObject (for object storage) and introducing the SCADA Server SDK (for fine-grained, GPU-initiated storage requests). These open-standard solutions enable GPUs to access storage directly over RDMA without CPU overhead. Tech leaders like Google Cloud, Microsoft, and IBM are backing these initiatives to build an interoperable high-performance storage ecosystem.
- 08arXivAI Research
Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance
In clinical diagnostics, extreme class imbalance makes traditional accuracy-based prompt optimization ineffective, as models can score high by simply predicting negative. To solve this, the authors propose Ranking-PE, which reformulates prompt search using pairwise ordering. By optimizing how well candidate prompts rank positive instances over negative ones, it directly targets empirical AUROC. Tested on the MIMIC dataset, Ranking-PE boosts AUROC by +5.8 pp on Qwen3-VL-8B and +16.2 pp on MedGemma-4B without any additional model-call overhead.
- 09arXivVideo AI
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
While video editing has progressed, modifying local 'scene text' (such as storefront signs or whiteboard text) while maintaining temporal stability remains highly challenging. This paper presents ViTeX-Bench, a benchmarking suite consisting of 387 real-world 720p videos and a 3-axis evaluation protocol spanning 13 metrics. The authors also release ViTeX-Edit-14B, a reference video editor fine-tuned with motion-aligned glyph conditioning, setting a new baseline for temporal text stability.
- 10arXivAI Research
VideoMSN: Turning Image Classifiers into Efficient Video Learners via Super-Images
This paper introduces VideoMSN, a decoder-free Masked Siamese Network framework that reframes video representation learning. By formatting video frames into grid-like 'super images', VideoMSN leverages standard image ViTs (like DINO-v3 or DeiT-v3) with two distinct masking strategies: spatial patch masking and temporal frame masking. This design prevents cross-frame info leakage and aligns embeddings using a masked Siamese loss. VideoMSN achieves state-of-the-art performance on UCF101 and Kinetics-400 while requiring up to 160x fewer video pretraining epochs than existing baselines.