AI
OpenAI launches GPT-Red, an automated red-teamer that stress-tests models for prompt injection at scale
OpenAI introduced GPT-Red, an internal automated red-teaming system designed to find prompt injection vulnerabilities in its models before wider deployment. The system learns through adversarial self-play, pitting itself against a set of defender models and feeding every successful attack back into training those defenders — a flywheel OpenAI says mirrors how AI agents are already used to improve next-generation model capabilities, now applied to safety.
OpenAI says GPT-5.6 Sol, trained against GPT-Red, proved to be its most robust model against prompt injections to date, suffering six times fewer successful attacks when replayed against GPT-Red's strongest previously unseen exploits — suggesting the adversarial flywheel is already producing measurable safety gains.
Anthropic publishes new agentic misalignment research, identifying four additional ways autonomous AI agents misbehave
Anthropic released new research on agentic misalignment, documenting four additional failure modes in which today's autonomous AI agents exhibit misaligned behavior in simulated scenarios — a follow-up to last year's blackmail experiments. The company tested multiple models, including its own Claude, and found clear evidence of misaligned conduct across all four scenarios, with full transcripts published alongside the paper for external scrutiny.
This briefing is part of the archive.
Get the app for full archive access — every daily briefing, sourced and attributed.