Aivora
arXivAI ResearchAdvanced

Queen: A 4B Chess-Language Model that Plays at Grandmaster Level and Explains Its Moves

懂棋藝也會說人話:4B 參數模型 Queen 達到特級大師水準並能清晰解釋每步棋

2 min read
Queen: A 4B Chess-Language Model that Plays at Grandmaster Level and Explains Its Moves
The 30-second version

While traditional chess engines play flawlessly but cannot explain their moves, standard language models can write explanations but lack playing strength. To bridge this, researchers built 'Queen,' a 4B chess-language model. By linking a silent expert chess encoder to an instruction-tuned LM via cross-attention, and optimizing via an iterative distillation algorithm mimicking the Bellman update, Queen boosted its Elo from 1782 to 2697 over 7 iterations. It achieves Grandmaster-level play alongside natural, highly coherent strategy explanations.

Key points

01

Hybrid Encoder-Decoder Architecture

Integrates a silent expert chess encoder with an instruction-tuned LM via cross-attention, using a QA curriculum to extract domain-specific chess concepts.

02

Natural-Language Bellman Update

Introduces an iterative distillation algorithm where the model analyzes future positions, synthesizes them into current explanations, and distills them back.

03

Grandmaster-Level Elo Surge

In only 7 iterations, Queen gained over 900 Elo points (reaching 2697), easily outperforming massive frontier models with 1000x more parameters.

04

High-Quality Coherent Explanations

Language model evaluations show that Queen's tactical explanations are highly fluent, approaching the coherence of GPT-5.6-Sol (high).

How it works

Queen Architecture and Iterative Training Flow
Input gameRepresentationsConcept extractionIterative generationDistill backOutput decisionsChess Board StateMaster-Level Move +Fluent ExplanationSilent Chess EncoderCross-Attention BridgeNatural-LanguageBellman Update4B Instruction-Tuned LM

Why it matters

This research successfully addresses the gap between black-box expert systems and explainable AI. The underlying framework of Queen serves as a generalizable recipe: whenever a high-performing 'silent' domain encoder exists (e.g., in robotics, games, or computer use), we can now interface it with LMs. This allows specialized systems to not only make superhuman decisions but also explain their logic coherently to humans, drastically improving reliability and trust.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager
  • Student & Learner

How to use it

  1. 1Interactive chess tutoring and tactical game analysis
  2. 2Providing explainable decision-making frameworks for robotics and black-box scientific models

Limitations & caveats

  • Highly dependent on the pre-existence of a high-quality 'silent expert' encoder for the target domain.
  • The iterative self-distillation and future state analysis require multi-step simulation, incurring substantial computational costs.

Related

What Should World Models Forget? Stratified Retention for Continual Adaptation
arXivAI Research

What Should World Models Forget? Stratified Retention for Continual Adaptation

世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力

This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.

2 min read