CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
CLIFT:基於共形自我驗證的網頁 Agent 訓練與測試端擴展技術
Training open-source web agents suffers from sparse binary rewards, while using frontier LLMs as step-by-step judges is too expensive and unavailable at deployment. CLIFT solves this via conformal self-verification. During training, the agent answers verification questions about its actions, filtered by a conformal certifier to guide reward design. At test time, this frozen verification bank enables Conformal Trajectory Selection (CTS) across multiple rollouts, achieving judge-free test-time scaling.
Key points
Self-Verification Mechanism
The agent answers natural-language verification questions about its own web-browsing rollouts, turning trajectories into structured evidence.
Compositional Conformal Certifier
Filters verification signals aligning with training-time judges and assigns weights using polarity-aware lift to blend into step-level rewards.
Judge-Free Test-Time Scaling
Freezes the certified bank for Conformal Trajectory Selection (CTS), using a conservative majority vote to select the best trajectory without external LLM calls.
Cross-Model Generalization
Achieves SOTA on WebArena Infinity; verification banks trained with open models transfer successfully to GPT-5.5 on VisualWebArena.
How it works
Why it matters
Traditional web agent RL relies heavily on expensive frontier LLMs for step-by-step evaluation. CLIFT proves that self-verification can serve as both a fine-grained reward generator and a judge-free test-time scaling mechanism. This significantly reduces the cost of deploying web agents and enables seamless knowledge transfer from open-source models to proprietary ones like GPT-5.5.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Web Agent RL training: Using fine-grained step-level rewards to solve the convergence issue of sparse success/failure rewards.
- 2Offline deployment & test-time scaling: Sampling multiple rollouts and selecting the best path without calling external paid APIs.
Limitations & caveats
- Relies on an initial training-time judge to calibrate the conformal certifier; biases in the initial judge may affect the verification bank quality.
- Sampling multiple trajectories during test-time scaling (CTS) increases inference latency and local computational resource consumption.
Related
Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?
The study introduces the BOTTLED benchmark to evaluate if LLM agents can autonomously package their capabilities into low-cost, task-specific artifacts, revealing that strong zero-shot performance does not guarantee successful bottling.
AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力
AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen world model, boosting the robustness of a 4B agent against adaptive prompt injections by 33.6%.
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
點陣圖秒變任意比例流程圖!「One Figure, Every Canvas」以 Agent 協同管線自動排版且支援 draw.io 編輯
This research introduces an agentic pipeline that automatically reformats raster flowcharts into various aspect ratios while maintaining structural fidelity, outputting editable draw.io XML files.