AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力
Web agents are highly vulnerable to prompt injections planted on third-party pages. AdvSim2Real solves this by co-evolving a task curriculum, an injection adversary, and the web agent itself within a frozen web world model. The adversary is rewarded for triggering "success flips" (turning success to failure), while the curriculum dynamically adjusts task difficulty. This approach significantly enhances both general task completion and robustness, achieving a 33.6% relative improvement under unseen frontier-model attacks when transferred to a real browser.
Key points
Three-way Co-evolution
Simultaneously co-evolves the task curriculum, injection adversary, and the agent inside a frozen web world model to create a dynamic adversarial loop.
Success Flip Reward
The adversary is rewarded specifically for "success flips," meaning it must successfully trick the agent into failing an otherwise solvable task.
Adaptive Task Curriculum
The curriculum generator prioritizes tasks with roughly 50% agent success rate to maintain an optimal learning gradient.
Sim2Real Transfer
The 4B agent trained in simulation shows excellent Sim2Real transfer, boosting task completion by 33.6% under unseen frontier-model attacks in real browsers.
How it works
Why it matters
Web agents must process untrusted third-party web content, making prompt injection an inherent and severe security risk. Traditional defenses often hurt the agent's general capabilities. AdvSim2Real demonstrates that a lightweight 4B model can achieve remarkable robustness against unseen, frontier-class adversaries without sacrificing base performance. By proving successful Sim2Real transfer, this work opens up a practical path for deploying highly secure yet computationally efficient autonomous web agents.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Training highly robust web agents that navigate untrusted third-party websites without executing malicious injected prompts.
- 2Conducting cost-effective security evaluations and red-teaming of AI agents using a simulated web world model.
Limitations & caveats
- The training effectiveness heavily relies on the accuracy and fidelity of the frozen web world model.
- Co-evolutionary training introduces high computational complexity and potential training instabilities that require careful hyperparameter tuning.
Related
Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?
The study introduces the BOTTLED benchmark to evaluate if LLM agents can autonomously package their capabilities into low-cost, task-specific artifacts, revealing that strong zero-shot performance does not guarantee successful bottling.
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
點陣圖秒變任意比例流程圖!「One Figure, Every Canvas」以 Agent 協同管線自動排版且支援 draw.io 編輯
This research introduces an agentic pipeline that automatically reformats raster flowcharts into various aspect ratios while maintaining structural fidelity, outputting editable draw.io XML files.
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
MemPilot:為多模態 AI Agent 打造的按需動態記憶整理框架
MemPilot is a runtime memory curation framework for LLM agents that uses a reinforcement learning policy to dynamically balance performance, cost, and latency.