arXivAI Safety
Caught in the Act: Probes Effectively Detect AI Sabotage and Unverbalized Deception
當場抓包:探針如何有效偵測 AI 的破壞行為與未言明的欺瞞
Researchers introduce the FIBS deception dataset and a novel multi-layer probe architecture that detects LLM sabotage and hidden deception directly from internal activations with up to 99.7% AUC.
2 min read