Aivora
arXivAI ResearchAdvanced

Why Forget-Only Unlearning Needs Memorization

「僅去記憶」的兩難:為什麼機器去學習反而需要模型「記住更多」?

2 min read
Why Forget-Only Unlearning Needs Memorization
The 30-second version

Machine unlearning aims to delete specific training data so the model behaves as if it were retrained from scratch. This paper investigates "forget-only" unlearning, where the algorithm only has access to the trained model and the forget targets. The authors demonstrate that because different datasets can yield identical models but require distinct outputs after deletion, forget-only unlearning has fundamental limitations. To succeed, the model must intentionally memorize more training information during its initial learning phase than standard training requires.

Key points

01

Forget-Only Constraint

In forget-only unlearning, the deletion algorithm receives only the trained model and the forget targets, with no retained training data or extra training information.

02

Same Model, Different Outputs

Different datasets can produce the exact same trained model but require very different outputs after the same forget examples are removed.

03

The Memorization Paradox

To handle arbitrary deletions, models must retain more information than usual. For threshold learners, the required info can be the entire dataset, though standard training keeps only one boundary point.

How it works

Standard Training vs. Forget-Only Ready Training
標準訓練 (Standard)支援「僅去記憶」訓練 (Forget-Only Ready)
Stored Information僅保留決策邊界 (例如單一邊界點)需額外記住多餘的資料細節 (甚至整個資料集)
Unlearning Accuracy極低,無法精準還原重新訓練的效果極高,可精準模擬從頭重新訓練
Space & Privacy Trade-off節省空間,但難以完全抹除特定資料的影響空間開銷大,但能隨時安全且精準地執行刪除

Why it matters

This work exposes a fundamental paradox: to successfully "forget" later without keeping the original dataset, a model must "memorize" more details during initial training. This provides crucial theoretical bounds for designing privacy-preserving and compliant machine learning systems.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager

How to use it

  1. 1Designing privacy-compliant ML models and training algorithms that satisfy GDPR "Right to be Forgotten" mandates.
  2. 2Assessing the feasibility and accuracy limits of unlearning on edge devices without access to the full training set.

Limitations & caveats

  • The findings are primarily based on theoretical analysis and lower bounds (such as threshold learners); empirical overhead in complex deep models like LLMs remains to be explored.

Related