Generative AI
Aug 4, 2026
Mistral launches Shieldstral, a new multimodal safety classifier with open weights
Aug 4, 2026
AI Summary
Mistral has introduced Shieldstral, a 3B open-weights multimodal safety classifier that outperforms larger models in content moderation tasks. It allows for adaptive policy evaluation without the need for retraining, providing calibrated safety scores for both text and images.
- Shieldstral is a 3B open-weights multimodal safety classifier designed for content moderation.
- It frames moderation as a policy-adaptive question-answering task, allowing users to input plain-language policies at inference time.
- The model can evaluate text and images simultaneously, providing a unified interface for safety assessments.
- Shieldstral has been shown to match or exceed the performance of models up to seven times its size across various safety benchmarks, including text safety and refusal detection.
- It operates efficiently on a single 16GB NVIDIA GPU and is released under the Apache 2.0 license.
- The model's design allows for flexibility in adapting to different safety definitions and contexts without requiring retraining.
- Shieldstral was developed using a diverse dataset and incorporates techniques to unify and calibrate safety data from various sources.
- The model aims to improve content moderation by adapting to specific contexts rather than relying on a fixed set of harm categories.
mistralmultimodalopen-weightsmodelmoderation