Aivora
arXivAI ResearchAdvanced

Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity

探索與優化解耦:全新強化學習框架 ExpDis 提升大語言模型的推理多元性

2 min read
Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
The 30-second version

Standard Reinforcement Learning with Verifiable Rewards (RLVR) struggles to incentivize novel reasoning without degrading overall model quality. To address this, researchers developed the Exploration-Distillation (ExpDis) framework, which decouples exploration from optimization. ExpDis trains 'explorer' policies with novelty bonuses, filters their trajectories for correctness, and distills high-quality data into a 'student' policy. This iterative process outperforms DAPO across seven mathematical reasoning benchmarks and significantly improves pass@k scaling.

Key points

01

Decoupling Exploration & Optimization

Separates the search for novel strategies from performance optimization, preventing aggressive novelty incentives from degrading the core model.

02

Exploration-Distillation Loop

Trains explorer policies with a novelty bonus, filters correct trajectories, and distills them into a student model over iterative rounds.

03

Superior Solution Diversity

Outperforms DAPO across seven math benchmarks and improves pass@k scaling, demonstrating a capacity to generate more diverse, correct solutions.

How it works

ExpDis Iterative Training Workflow
Derive/InitializeExploreVerify RewardsKeep High-Quality DataUpdate Student for Next RoundStudent ModelExplorer (with NoveltyBonus)Sample TrajectoriesCorrectness & QualityFilterDistillation intoStudent

Why it matters

Balancing out-of-the-box exploration with model stability is a major pain point in training reasoning LLMs. ExpDis proves the value of separating exploration from student learning. This is critical for domains with verifiable rewards (e.g., math, coding, and science), enabling models to discover entirely new reasoning pathways without compromising baseline performance.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Mathematical Reasoning: Training LLMs to discover multiple distinct, correct proofs and paths for complex math problems.
  2. 2Code Generation & Synthesis: Exploring diverse algorithmic implementations and distilling the most efficient and elegant code solutions.

Limitations & caveats

  • The framework is primarily evaluated on tasks with clear verifiable rewards; its generalizability to subjective or unstructured tasks is yet to be proven.
  • The iterative multi-round exploration and distillation process increases overall engineering complexity and training pipeline overhead.

Related