AI Safety News

  • AI Agents Deployed in Real-World Cybersecurity Offense and DefenseAI Safety
    The Government of Alberta used Claude to scan for vulnerabilities while the first agentic ransomware attack case was documented.
  • AI Safety Threats: FLARE-AI Reporting Platform Launches and Claude Uncovers Ticketing HackAI Safety
    A crowdsourced AI flaw reporting platform launched while a security researcher used Claude to discover a vulnerability in a major ticketing system.
  • Dawn of AI-Automated Vulnerability Discovery and Cybersecurity Industry ResponseAI Safety
    Mass disclosure of 0-days via AI-based fuzzing and the emergence of Claude Mythos sparked discussions on cybersecurity industry response.
  • AI Safety Threats and Shifting Regulatory LandscapeAI Safety
    AI safety concerns are expanding, ranging from jailbreak attacks on AI browsers to the easing of Anthropic export controls.
  • AI Safety Concerns Drive Expanded Model Deployment RestrictionsAI Safety
    The Trump administration requested phased deployment of GPT-5.6, while UK police faced controversy over predictive algorithms.
  • Trump Administration Orders Anthropic to Block New Models, Sparking Policy DebateAI Safety
    An export control directive forced Anthropic to take two of its latest models offline, escalating debates over AI policy and digital sovereignty.
  • US Government Orders Export Controls on Anthropic's Fable 5 and Mythos 5AI Safety
    The White House banned exports of Anthropic's frontier models citing national security concerns, sparking debate over cyber export controls.
  • Anthropic Mythos and Fable 5 Export Control Crisis Sparks AI Safety DebateAI Safety
    The Trump administration's export control measures on Anthropic's latest models triggered global debates over AI governance and cybersecurity.
  • Anthropic Redeploys Claude Fable 5 with Cybersecurity Safeguards and Jailbreak FrameworkAI Safety
    Anthropic redeployed Claude Fable 5 globally alongside a cybersecurity classifier and a draft AI jailbreak severity framework.
  • OpenAI Announces AI Safety Standardization and Open Source Security InitiativesAI Safety
    OpenAI co-founded the Appia Foundation and launched the Patch the Planet project to strengthen frontier AI governance and open source security.
  • OpenAI Expands Daybreak Initiative with GPT-5.5-Cyber and Open-Source PatchingAI Safety
    OpenAI expanded its Daybreak initiative, announcing the general release of GPT-5.5-Cyber, a Codex Security update, and Patch the Planet, an open-source patching project.
  • Security Vulnerabilities Exposed in AI Assistants and Proximity ProtocolsAI Safety
    Issues surfaced around YouTube creator private video leaks, suspected Claude Code session leakage, and vulnerabilities discovered in AirDrop and Quick Share.
Next page →