AdviSD: Steering Frontier LLMs via Targeted Multi-Turn Self-Distillation
微型顧問的逆襲:AdviSD 透過目標多輪自我蒸餾精準引導前沿大模型
The paper introduces AdviSD, where a small trainable advisor (e.g., Qwen3-8B) generates natural language advice to guide frozen, black-box frontier LLMs (e.g., Gemini, Claude). To prevent ineffective corrections from degrading training, AdviSD couples outcome-based RL with targeted self-distillation. The advisor scores executor responses with and without advice, distilling only from high-impact decisions. AdviSD outperforms advisor-GRPO by up to 6.4% on BFCL-v3 and 5.1 points on EnvScaler, demonstrating strong cross-model generalization.
Key points
Lightweight Steering
Uses a small, trainable advisor (Qwen3-8B) to guide frozen, closed-source frontier models via natural-language advice, bypassing heavy fine-tuning.
Selective Self-Distillation
Compares execution outputs with and without advice, training only on high-impact decisions to avoid noise from useless corrections.
Zero Extra Rollout Overhead
Requires no executor likelihoods or additional executor rollouts, significantly reducing computational overhead during training.
Cross-Model Generalization
Outperforms advisor-GRPO on BFCL-v3 and EnvScaler, and generalizes well to out-of-domain tasks and different executor model families.
How it works
Why it matters
This research addresses the challenges of customizing massive, closed-source frontier LLMs. By training a small, external advisor to guide frozen models via natural language, developers can steer LLM behavior without weight access or expensive API tuning. AdviSD's selective distillation ensures efficient training, showing a practical, low-cost path for enterprise domain-specific alignment.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Improving complex tool-use accuracy for API-based closed-source LLMs
- 2Aligning behaviors and seamlessly migrating across different LLM versions
- 3Low-cost multi-turn guidance for domain-specific tasks
Limitations & caveats
- Highly dependent on the executor's inherent capability to understand and execute natural-language advice.
- Multi-turn reflection and evaluation overhead may be less suitable for ultra-low-latency, single-turn inference applications.
Related
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸
STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化
LeapQuant is a training-free quantization method that achieves near-lossless 8-bit recurrent state quantization for linear attention, delivering up to 1.47x end-to-end inference speedup while preserving FP32 accuracy.
Telescopic Language Models: One Training Run for Endless Compute Budgets
伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM
This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.