NeuralEraser

Links

Project Overview

The goal of NeuralEraser is to improve both our understanding and control over AI systems, particularly image models, to both improve their alignment with global human values and eliminate their potential to do wrongdoing. NeuralEraser uses sparse autoencoders to interpretably identify and robustly erase knowledge of specific artistic styles within stable diffusion models therefore ensuring that it is impossible for the model to generate specific images no matter what jailbreaking mechanisms are used on it. This significantly enhances fine-grained controllability and interpretability in generative image models by allowing the user to selectively and robustly unlearn specific styles in image models.

Project Status

  • [ CURRENTLY IN DEVELOPMENT ]
  • Currently working on making the method more rebust against whitebox methods as detailed in the following paper: [link]