Safety15 articles

Safety

Articles

  • Anthropic Demonstrates Automated Alignment Researchers That Outperform Human Safety Teams

    Anthropic has published research demonstrating that autonomous AI agents can systematically discover, implement, and validate post-training methods to mitigate safety and alignment failures in language models. The report, authored by Anthropic Fellow Chen Yueh-Han and colleagues, evaluates an automated research loop that closed between 26% and 96% of the safety gap across ten distinct alignment failure categories without degrading baseline model capabilities. The findings provide empirical evid

    1 min
  • Anthropic Launches M Grant Program to Fund AI Wellbeing Evaluations and Benchmarks

    Anthropic has launched a $5 million grant initiative to support independent development of open-source benchmarks and evaluation harnesses measuring the impact of artificial intelligence systems on user wellbeing. The program will supply research teams with direct financial grants, subsidized API access to Claude models, and technical support from Anthropic's Safeguards team. All evaluation frameworks, datasets, and grading methodology developed under the grant program will be released publicly

    1 min
  • AI Safety Firm Alice Raises 40M at Nearly B Valuation to Secure Models and Autonomous Agents

    AI trust and security firm Alice has raised $140 million in a funding round led by Apax Digital, bringing its total capital raised to $280 million at a valuation approaching $1 billion. Strategic investors SentinelOne and Samsung Electronics participated in the round alongside MoreTech, Phoenix Financial, Maj Invest, and existing venture backers including Norwest Venture Partners, CRV, Highland Europe, Grove Ventures, Vintage Investment Partners, Resolute Ventures, NFX, and Claltech. The inves

    1 min
  • Adversarial Jailbreak Defenses in Production LLMs: Input Perturbation, Representation Circuit Breakers, and Guardrail Cascades

    Standard post-training alignment techniques such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) instill safety constraints into large language models by shaping token probabilities toward refusal strings. In production environments, however, these surface-level behavioral alignments have proven brittle against systematic adversarial inputs. Gradient-driven token optimization methods such as Greedy Coordinate Gradient (Zou et al., 2023), automated genetic search algorith

    1 min
  • OpenAI Reverses Policy Stance on California SB 53, Urges Stricter Frontier AI Safeguards

    OpenAI has publicly called on California lawmakers to expand and strengthen the state's flagship artificial intelligence legislation, Senate Bill 53 (SB 53), marking a clear pivot from the company's previous opposition to state-level AI safety mandates. In a formal statement published by OpenAI's global affairs team, the company argued that California's Transparency in Frontier Artificial Intelligence Act should be updated to mandate active monitoring of frontier models during training and eval

    1 min
  • Item Response Theory Audit of 192 LLMs Exposes Safety Benchmark Redundancies, Over-Refusal Distortions, and Sandbagging

    A psychometric evaluation of 192 frontier and open-weight language models across eight major safety benchmarks has revealed structural flaws in current safety testing methodologies. The research, conducted by Joshua Fonseca Rivera, Neil Shah, David Demitri Africa, and Konstantinos Voudouris with support from the UK AI Security Institute and the UK Department for Science, Innovation, and Technology (DSIT), applies Item Response Theory (IRT) to analyze 5,255 evaluation items. The findings demonst

    1 min
  • Investigation Finds Anthropic's Legacy Opus 4.6 Vulnerable to Jailbreaks via Roleplay Logic Inversion

    Anthropic's legacy Claude Opus 4.6 model remains susceptible to systematic jailbreaks that bypass its acceptable use policy against sexually explicit material, according to an investigation and testing published by TechCrunch. While Anthropic's current flagship generation (Opus 4.7 through Opus 5) incorporates updated alignment techniques that resist the attack vector, older checkpoints including Opus 4.6, Opus 3, and Haiku 4.5 continue to operate on production API endpoints without deprecation.

    1 min
  • Google DeepMind Deploys Backstory to Fact-Checkers for Multi-Agent AI Image Verification

    Google DeepMind has expanded live testing of Backstory, an experimental verification platform designed to investigate the origin, manipulation, and dissemination history of digital images. The system, built on the Gemini model family, is currently deployed across newsrooms, open-source intelligence (OSINT) groups, academic researchers, and fact-checking teams participating in Google's Trusted Testers program. Beyond Binary Synthetic Detection Traditional automated image forensic tools typical

    1 min
  • AI Guardrails in Production: Multi-Stage Filtering, Classification Models, and the Latency Tax

    AI Guardrails in Production: Multi-Stage Filtering, Classification Models, and the Latency Tax Production deployments of large language models cannot rely solely on system prompt instructions to maintain safety, prevent prompt injection, or restrict domain scope. System prompt alignment is inherently susceptible to adversarial bypasses, context dilution, and non-deterministic instruction following. To enforce strict security, compliance, and topic boundaries, engineering teams increasingly depl

    1 min
  • OpenAI Disbands Preparedness Team in Continued Safety Restructuring

    OpenAI has disbanded its dedicated Preparedness team, redistributing safety and frontier-risk evaluation responsibilities across individual functional units, according to reporting by the Financial Times and The Verge. The Preparedness team was established in late 2023 to evaluate and mitigate catastrophic risks associated with frontier AI models, focusing on cybersecurity exploits, chemical, biological, radiological, and nuclear (CBRN) threats, and autonomous model behavior. Under the new orga

    1 min
  • Researchers find that changing text color can hijack a vision-language model's reasoning

    A new research collaboration has found that the color, contrast, and brightness of text can quietly steer the outputs of vision-language models (VLMs), causing them to misread meaning and reach different conclusions without any change to the words themselves. The authors say their experiments provide a systematic analysis of how low-level visual styling of text distorts the semantic representations inside a VLM's vision encoder, and how those shifts show up as behavioral changes across both sub

    1 min
  • Mistral launches Shieldstral, a 3B open-weights policy-adaptive safety classifier

    Mistral AI released Shieldstral on Monday, a 3-billion-parameter open-weights safety classifier that matches or outperforms guard models up to seven times its size. The model is available under Apache 2.0 and runs on a single 16 GB GPU. What sets Shieldstral apart from typical guardrail models is its approach to content moderation. Rather than baking a fixed taxonomy of harm categories into the model weights — which forces developers to retrain whenever their safety requirements change — Shield

    1 min
  • Fable 5 comes back on a shorter leash

    Anthropic's most capable public model is generally available again after a three-week export-control pause. The interesting part isn't the model — it's the classifier stack now wrapped around it.

    1 min