Search papers, labs, and topics across Lattice.
Shieldstral is a 3B-parameter multimodal safety classifier that achieves state-of-the-art performance on text safety benchmarks, outperforming larger models by nearly 7脳. By framing content moderation as a binary question-answering task, it consolidates various moderation tasks into a unified framework, allowing for the integration of heterogeneous safety datasets. The model's design, supported by a robust dataset of 54.1M samples, demonstrates that smaller, adaptive models can effectively compete with their larger counterparts in safety classification tasks.
A 3B-parameter model outperforms larger competitors in safety classification by reframing content moderation as a simple yes/no question.
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.