Queen: A 4B Chess-Language Model that Plays at Grandmaster Level and Explains Its Moves
懂棋藝也會說人話:4B 參數模型 Queen 達到特級大師水準並能清晰解釋每步棋
While traditional chess engines play flawlessly but cannot explain their moves, standard language models can write explanations but lack playing strength. To bridge this, researchers built 'Queen,' a 4B chess-language model. By linking a silent expert chess encoder to an instruction-tuned LM via cross-attention, and optimizing via an iterative distillation algorithm mimicking the Bellman update, Queen boosted its Elo from 1782 to 2697 over 7 iterations. It achieves Grandmaster-level play alongside natural, highly coherent strategy explanations.
Key points
Hybrid Encoder-Decoder Architecture
Integrates a silent expert chess encoder with an instruction-tuned LM via cross-attention, using a QA curriculum to extract domain-specific chess concepts.
Natural-Language Bellman Update
Introduces an iterative distillation algorithm where the model analyzes future positions, synthesizes them into current explanations, and distills them back.
Grandmaster-Level Elo Surge
In only 7 iterations, Queen gained over 900 Elo points (reaching 2697), easily outperforming massive frontier models with 1000x more parameters.
High-Quality Coherent Explanations
Language model evaluations show that Queen's tactical explanations are highly fluent, approaching the coherence of GPT-5.6-Sol (high).
How it works
Why it matters
This research successfully addresses the gap between black-box expert systems and explainable AI. The underlying framework of Queen serves as a generalizable recipe: whenever a high-performing 'silent' domain encoder exists (e.g., in robotics, games, or computer use), we can now interface it with LMs. This allows specialized systems to not only make superhuman decisions but also explain their logic coherently to humans, drastically improving reliability and trust.
Who it affects
- AI Researcher
- AI Developer
- Product Manager
- Student & Learner
How to use it
- 1Interactive chess tutoring and tactical game analysis
- 2Providing explainable decision-making frameworks for robotics and black-box scientific models
Limitations & caveats
- Highly dependent on the pre-existence of a high-quality 'silent expert' encoder for the target domain.
- The iterative self-distillation and future state analysis require multi-step simulation, incurring substantial computational costs.
Related
Less Decoder is More Encoder: Extracting Robust 3D Geometric Representations via Novel View Synthesis
減少解碼器反而增強編碼器:從新視角合成中提煉強大三維幾何表徵
This paper reveals how expressive decoders dilute geometric learning in Novel View Synthesis, and proposes SNAP—a self-supervised framework that restricts decoders to force encoders to learn robust 3D representations.
4DCodeBench: Benchmarking AI Agents on 4D Inverse Graphics and Dynamic Scene Code Generation
4DCodeBench:評估 AI Agent 動態場景 4D 反向圖形學與程式碼生成能力的全新基準
4DCodeBench is a new benchmark designed to evaluate AI agents' ability to reconstruct 4D dynamic scenes from videos by generating executable graphics and physics code.
What Should World Models Forget? Stratified Retention for Continual Adaptation
世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力
This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.