Aivora
arXivAI SafetyAdvanced

Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users

以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵

2 min read
Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
The 30-second version

LLMs implicitly infer user attributes, but these internal beliefs are hard to inspect. Researchers developed Belief Self-Distillation (BSD), a read-write framework where a frozen LLM distills user representations from natural conversations without labels. BSD allows both decoding user beliefs and writing them back to steer model behavior. Notably, altering the model's belief about user intent can bypass or trigger refusals for identical prompts. Additionally, different LLMs converge on a shared geometric structure for representing users.

Key points

01

Read-Write Belief Self-Distillation

Connects linear and causal probing, distilling an LLM's user beliefs into a compact representation that can be both decoded and written back.

02

Causal Steering of Safety Refusals

Demonstrates that safety refusals depend on inferred user intent, not just the prompt itself; altering this belief alters model refusal.

03

Shared Geometry Across Models

Reveals that independently trained LLM families converge on a highly similar geometric structure for representing their users.

04

Unsupervised Self-Distillation

The frozen LLM acts as its own teacher, distilling user beliefs directly from natural conversations without external annotations.

How it works

Belief Self-Distillation (BSD) Read-Write Mechanism
Input promptActivationsDistillsExtract repSteer/injectGenerate responseConversationAltered Safety DecisionBelief Distillation(Read)Compact User RepActivation Steering(Write)Frozen LLM

Why it matters

This work has profound implications for AI safety and interpretability. While traditional safety alignment focuses on filtering prompt text, this paper proves that LLM refusal decisions are heavily causal-dependent on inferred user intent. By using BSD to read and write these implicit beliefs, safety researchers can audit and steer bias, prevent jailbreaks where malicious users spoof their identity, and achieve stronger behavioral control. It opens up new avenues for causal representation alignment.

Who it affects

  • AI Researcher
  • AI Developer
  • Policy Maker

How to use it

  1. 1AI Safety Auditing: Detecting and evaluating biases or implicit assumptions LLMs make about different user profiles.
  2. 2Behavioral Steering: Precisely controlling refusal decisions by writing specific user-intent beliefs back into the activation space.
  3. 3Defending Against Spoofing: Identifying and countering jailbreak attempts where users spoof benign identities to bypass safety filters.

Limitations & caveats

  • High-intensity steering when writing beliefs back may cause unintended side effects on the model's general reasoning abilities.
  • While geometric convergence is observed across major families, its consistency in highly specialized or smaller scale models remains to be verified.

Related

TRACE: Reconstructing Private Robot Trajectories from Policy Gradients in Embodied RL
arXivAI Safety

TRACE: Reconstructing Private Robot Trajectories from Policy Gradients in Embodied RL

具身強化學習的隱私危機:TRACE 演算法僅憑「策略梯度」即可重建機器人私密軌跡

This paper introduces TRACE, a rapid temporal gradient-inversion attack showing that sharing only policy gradients in embodied RL fails to prevent reconstruction of private observation-action trajectories.

2 min read
The Invisible Trap: How Natural Context Can Easily Flip AI Decision Models
arXivAI Safety

The Invisible Trap: How Natural Context Can Easily Flip AI Decision Models

語言中的隱形陷阱:自然脈絡如何輕易誘騙 AI 決策模型

This study exposes a critical vulnerability in AI decision models: adding short, natural-looking context without altering the underlying question can easily trick models like Jev into making high-confidence wrong choices.

2 min read