Aivora

#deception-detection

Deception Detection

1 篇文章

當場抓包:探針如何有效偵測 AI 的破壞行為與未言明的欺瞞
arXivAI 安全

當場抓包:探針如何有效偵測 AI 的破壞行為與未言明的欺瞞

Caught in the Act: Probes Effectively Detect AI Sabotage and Unverbalized Deception

研究團隊發布目前最大的 AI 欺瞞資料集 FIBS 與新型白盒探針架構,能直接從模型內部神經激活狀態偵測隱瞞意圖與破壞行為,對隱蔽目標的辨識率高達 99.7% AUC。

2 分鐘閱讀