Safety4 articles

Safety

Articles

  • Mistral launches Shieldstral, a 3B open-weights policy-adaptive safety classifier

    Mistral AI released Shieldstral on Monday, a 3-billion-parameter open-weights safety classifier that matches or outperforms guard models up to seven times its size. The model is available under Apache 2.0 and runs on a single 16 GB GPU. What sets Shieldstral apart from typical guardrail models is its approach to content moderation. Rather than baking a fixed taxonomy of harm categories into the model weights — which forces developers to retrain whenever their safety requirements change — Shield

    1 min
  • Fable 5 comes back on a shorter leash

    Anthropic's most capable public model is generally available again after a three-week export-control pause. The interesting part isn't the model — it's the classifier stack now wrapped around it.

    1 min