Aivora
arXivLLMAdvanced

AdviSD: Steering Frontier LLMs via Targeted Multi-Turn Self-Distillation

微型顧問的逆襲:AdviSD 透過目標多輪自我蒸餾精準引導前沿大模型

2 min read
AdviSD: Steering Frontier LLMs via Targeted Multi-Turn Self-Distillation
The 30-second version

The paper introduces AdviSD, where a small trainable advisor (e.g., Qwen3-8B) generates natural language advice to guide frozen, black-box frontier LLMs (e.g., Gemini, Claude). To prevent ineffective corrections from degrading training, AdviSD couples outcome-based RL with targeted self-distillation. The advisor scores executor responses with and without advice, distilling only from high-impact decisions. AdviSD outperforms advisor-GRPO by up to 6.4% on BFCL-v3 and 5.1 points on EnvScaler, demonstrating strong cross-model generalization.

Key points

01

Lightweight Steering

Uses a small, trainable advisor (Qwen3-8B) to guide frozen, closed-source frontier models via natural-language advice, bypassing heavy fine-tuning.

02

Selective Self-Distillation

Compares execution outputs with and without advice, training only on high-impact decisions to avoid noise from useless corrections.

03

Zero Extra Rollout Overhead

Requires no executor likelihoods or additional executor rollouts, significantly reducing computational overhead during training.

04

Cross-Model Generalization

Outperforms advisor-GRPO on BFCL-v3 and EnvScaler, and generalizes well to out-of-domain tasks and different executor model families.

How it works

AdviSD System Architecture and Selective Self-Distillation Workflow
Pass taskProvide adviceExecution outputFilter high-impact correctionsUpdate advisor parametersUser Task InputTrainable Advisor(Qwen)Frozen ExecutorEvaluate Diff (With vsWithout Advice)SelectiveSelf-Distillation

Why it matters

This research addresses the challenges of customizing massive, closed-source frontier LLMs. By training a small, external advisor to guide frozen models via natural language, developers can steer LLM behavior without weight access or expensive API tuning. AdviSD's selective distillation ensures efficient training, showing a practical, low-cost path for enterprise domain-specific alignment.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Improving complex tool-use accuracy for API-based closed-source LLMs
  2. 2Aligning behaviors and seamlessly migrating across different LLM versions
  3. 3Low-cost multi-turn guidance for domain-specific tasks

Limitations & caveats

  • Highly dependent on the executor's inherent capability to understand and execute natural-language advice.
  • Multi-turn reflection and evaluation overhead may be less suitable for ultra-low-latency, single-turn inference applications.

Related

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
arXivLLM

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States

STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸

STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.

2 min read
Telescopic Language Models: One Training Run for Endless Compute Budgets
arXivLLM

Telescopic Language Models: One Training Run for Endless Compute Budgets

伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM

This research introduces Telescopic Language Models (TLM), which use stochastic prefix supervision to enable a single Transformer to act as a valid language model at any layer depth, serving diverse compute budgets from a single training run.

2 min read