Back to news
AI Ethics
Jul 13, 2026

MIT researchers develop method to detect AI models generating illegal content without outputting it

Jul 13, 2026
AI Summary

A team of researchers from MIT and Thorn has created a new auditing technique to identify AI models capable of generating illegal content, such as child sexual abuse material (CSAM), without needing to prompt the models. This method, which achieved 100% accuracy in tests, aims to enhance safety measures for AI systems and protect children from exploitation.

MIT researchers develop method to detect AI models generating illegal content without outputting it
  • The rise of generative AI has led to concerns about its misuse for creating illegal content, including CSAM. Reports of AI-generated CSAM increased from 67,000 in 2024 to over 1.5 million in 2025 according to the National Center for Missing and Exploited Children.
  • Traditional methods for testing AI models for harmful outputs are not feasible for CSAM due to legal restrictions on generating such content.
  • Researchers from MIT, led by Vinith Suriyakumar, developed a new approach that analyzes the internal modifications made to AI models during fine-tuning, specifically using a technique called Gaussian probing.
  • This method examines how a model's internal structure is altered without generating any outputs, allowing for the identification of models that have been adapted to produce harmful content.
  • The technique demonstrated 100% accuracy in identifying models specialized for generating CSAM during testing.
  • The researchers emphasize the importance of scalability and cost-effectiveness in implementing this auditing method, as thousands of AI model variations are released online monthly.
  • Future research aims to evaluate this technique on a broader range of models and potentially detect harmful capabilities in base models before they are adapted for malicious use.
generative aichild safetyauditingillegal contentresearch