Aivora

#probing

Probing

1 article

Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
arXivAI Safety

Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users

以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵

Researchers introduce Belief Self-Distillation (BSD), a framework to read and write an LLM's implicit beliefs about its users, revealing how user-intent inference drives safety decisions and showing shared representation geometries across models.

2 min read