FOMO: Forget the Concept, Don’t Miss Out on the Scene in Selective Video Unlearning

Łukasz Rudnik1, Agnieszka Polowczyk1,2, Alicja Polowczyk1,2, Przemysław Spurek1,2

1Jagiellonian University    2IDEAS Research Institute

Videos on this page contain content that may be disturbing, including blood, violence and nudity.

Unlearning results for identity, motion, nudity and gore.
We present FOMO: a selective video unlearning method that erases only the unwanted concept, while keeping scene composition and dynamics close to the original. It removes harmful concepts like nudity, the likeness of public figures, and disturbing or violent content, and was specifically designed to handle motion unlearning, where the concept is not an object in the frame but an action carried across the whole clip.

Abstract

The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesirable material, making concept unlearning an essential mitigation. Existing methods mainly target static visual concepts, such as objects, identities, or unsafe appearance, largely overlooking motion unlearning. Furthermore, these approaches often pay little attention to preserving the surrounding scene. As a result, successful concept removal may unintentionally alter the background, composition, or overall video dynamics. We argue that effective unlearning should ideally change only what is targeted, while minimizing unnecessary changes to the remaining scene. In this work, we introduce FOMO, to the best of our knowledge the first training-based selective video unlearning method that directly treats preservation of the original scene as a priority. We formulate unlearning around two complementary objectives: what to change and what to preserve. Our method localizes concept-related representations and modifies them, while the preservation mechanism maintains non-target scene information without requiring auxiliary data. Beyond simply erasing unwanted concepts, FOMO explicitly redirects the generation toward a specified safe alternative. We further extend this formulation to motion unlearning, where the concept is defined by temporal behavior rather than a fixed spatial region. Our solution achieves effective unlearning across unsafe content, object, and motion concepts, while achieving the best trade-off between concept removal and scene preservation.

Method

Overview of the FOMO training pipeline.

BibTeX

TODO