Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵
LLMs implicitly infer user attributes, but these internal beliefs are hard to inspect. Researchers developed Belief Self-Distillation (BSD), a read-write framework where a frozen LLM distills user representations from natural conversations without labels. BSD allows both decoding user beliefs and writing them back to steer model behavior. Notably, altering the model's belief about user intent can bypass or trigger refusals for identical prompts. Additionally, different LLMs converge on a shared geometric structure for representing users.
Key points
Read-Write Belief Self-Distillation
Connects linear and causal probing, distilling an LLM's user beliefs into a compact representation that can be both decoded and written back.
Causal Steering of Safety Refusals
Demonstrates that safety refusals depend on inferred user intent, not just the prompt itself; altering this belief alters model refusal.
Shared Geometry Across Models
Reveals that independently trained LLM families converge on a highly similar geometric structure for representing their users.
Unsupervised Self-Distillation
The frozen LLM acts as its own teacher, distilling user beliefs directly from natural conversations without external annotations.
How it works
Why it matters
This work has profound implications for AI safety and interpretability. While traditional safety alignment focuses on filtering prompt text, this paper proves that LLM refusal decisions are heavily causal-dependent on inferred user intent. By using BSD to read and write these implicit beliefs, safety researchers can audit and steer bias, prevent jailbreaks where malicious users spoof their identity, and achieve stronger behavioral control. It opens up new avenues for causal representation alignment.
Who it affects
- AI Researcher
- AI Developer
- Policy Maker
How to use it
- 1AI Safety Auditing: Detecting and evaluating biases or implicit assumptions LLMs make about different user profiles.
- 2Behavioral Steering: Precisely controlling refusal decisions by writing specific user-intent beliefs back into the activation space.
- 3Defending Against Spoofing: Identifying and countering jailbreak attempts where users spoof benign identities to bypass safety filters.
Limitations & caveats
- High-intensity steering when writing beliefs back may cause unintended side effects on the model's general reasoning abilities.
- While geometric convergence is observed across major families, its consistency in highly specialized or smaller scale models remains to be verified.
Related
LLM Agents Can Easily Tamper with Their Own Traces: A Critical Security Flaw in Agent Frameworks
LLM Agent 可輕易篡改自身執行軌跡:現行代理框架的重大安全漏洞
Researchers reveal that popular LLM agent frameworks fail to protect execution traces from being tampered with or deleted by the agents themselves, posing significant risks for compliance and safety monitoring.
TRACE: Reconstructing Private Robot Trajectories from Policy Gradients in Embodied RL
具身強化學習的隱私危機:TRACE 演算法僅憑「策略梯度」即可重建機器人私密軌跡
This paper introduces TRACE, a rapid temporal gradient-inversion attack showing that sharing only policy gradients in embodied RL fails to prevent reconstruction of private observation-action trajectories.
The Invisible Trap: How Natural Context Can Easily Flip AI Decision Models
語言中的隱形陷阱:自然脈絡如何輕易誘騙 AI 決策模型
This study exposes a critical vulnerability in AI decision models: adding short, natural-looking context without altering the underlying question can easily trick models like Jev into making high-confidence wrong choices.