Memory Fossils in Large Language Models: A Mechanistic Framework for Quantifying Residual Latent Knowledge Post-Unlearning

The concept of the Memory Fossils framework is presented to discover any potential leftover information existing in LLMs after implementing machine unlearning. The framework distinguishes three types of residues: Activation Fossils, Attention Fossils, and Gradient Fossils. It also proposes deriving a score called Fossil Score by combining different approaches, such as probing, causal reconstruction, and output deviations. The results indicate a low output recall does not guarantee complete suppression of knowledge in question, indicating a need to examine internal representations along with behavior.