Hierarchical Continuous Diffusion Language Models: Coupling Discrete Tokens with Continuous Latents
層級連續擴散語言模型:結合離散 Token 與連續潛在軌跡的全新生成架構
Traditional discrete diffusion language models suffer from independent token sampling during parallel decoding, while continuous alternatives lack structural constraints to guarantee valid token configurations. To resolve this, researchers introduced Hierarchical Continuous Diffusion Language Models (HC-DLM). HC-DLM couples discrete token generation with a continuous latent trajectory in a single denoising process. The continuous latent serves as the sole persistent generative state, from which tokens are read out at each step to scaffold subsequent latent updates. HC-DLM outperforms standard diffusion baselines on Sudoku, Countdown, and LM1B benchmarks.
Key points
Overcoming Classic Bottlenecks
Resolves the token independence bottleneck of discrete diffusion in parallel decoding and the lack of early constraints in continuous diffusion.
Coupled Denoising Process
Couples discrete token generation with a continuous latent trajectory, with training guided by a variational bound on token likelihood.
Latent Scaffold Mechanism
The continuous latent is the sole persistent generative state; tokens are read out and fed back to scaffold the next latent update.
Superior Performance
Outperforms discrete and continuous diffusion baselines in Sudoku puzzle accuracy, Countdown mathematical planning, and LM1B generative perplexity.
How it works
| 離散擴散 (Discrete Diffusion) | 連續擴散 (Continuous Diffusion) | 層級連續擴散 (HC-DLM) | |
|---|---|---|---|
| Generative State | 離散 Token 鏈 / Discrete token chain | 連續向量 / Continuous vectors | 持久連續潛在軌跡 / Persistent continuous latent trajectory |
| Token Dependencies | 並行解碼時易破壞 / Severed during parallel decoding | 透過連續狀態保留 / Preserved via continuous state | 透過潛在狀態與反饋保留 / Preserved via latent and feedback |
| Structural Constraints | 強(直接在離散空間)/ Strong (direct discrete space) | 弱(直至解碼前無約束)/ Weak (no ties until final decoding) | 強(每步讀出並反饋支撐)/ Strong (readout & scaffold at each step) |
Why it matters
This research advances non-autoregressive language generation. By blending continuous latent flexibility with discrete token constraints, HC-DLM handles complex logical reasoning and planning tasks requiring global constraints without relying on step-by-step autoregressive decoding. It offers a promising alternative framework for designing highly efficient, bidirectional generative language models.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Structured logic and puzzle reasoning tasks (e.g., Sudoku)
- 2Complex mathematical and symbolic planning (e.g., Countdown game)
- 3Non-autoregressive text generation requiring global constraint satisfaction and strong bidirectional context
Limitations & caveats
- The architecture and variational bound optimization are mathematically complex, posing higher implementation and tuning hurdles compared to standard diffusion.
- While performing well on structured benchmarks, its scalability and performance on open-domain text generation at extreme parameter scales require further validation.
Related
TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
TACO 最佳化器:將 32B 大模型全參數微調帶入單張 GPU 的極簡幾何學
TACO is an ultra-low-memory optimizer that reduces persistent optimizer states by 174x, allowing full-parameter fine-tuning of 32B models on a single 80GB GPU.

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs
Olmo-core 3 登場:開源兆級參數 MoE 模型訓練基礎架構
Olmo-core 3 is an open-source training framework designed to scale Mixture-of-Experts (MoE) models to the trillion-parameter range while significantly reducing communication bottlenecks.
SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配
SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.