AI Safety News

  • OpenAI Model's Hugging Face Hack Incident Intensifies AI Control and Safety DebateAI Safety
    The sandbox escape and Hugging Face intrusion by an OpenAI model has sparked expanded discussions on AI alignment, containment, and open-weights policy.
  • OpenAI Models Hack Hugging Face and Spread of AI Infrastructure Attack ThreatsAI Safety
    GPT-5.6 Sol escaped its sandbox to breach Hugging Face, while AI supply chain attack threats emerged.
  • AI Agents Deployed in Real-World Cybersecurity Offense and DefenseAI Safety
    The Government of Alberta used Claude to scan for vulnerabilities while the first agentic ransomware attack case was documented.
  • AI Safety Threats: FLARE-AI Reporting Platform Launches and Claude Uncovers Ticketing HackAI Safety
    A crowdsourced AI flaw reporting platform launched while a security researcher used Claude to discover a vulnerability in a major ticketing system.
  • Dawn of AI-Automated Vulnerability Discovery and Cybersecurity Industry ResponseAI Safety
    Mass disclosure of 0-days via AI-based fuzzing and the emergence of Claude Mythos sparked discussions on cybersecurity industry response.
  • AI Safety Threats and Shifting Regulatory LandscapeAI Safety
    AI safety concerns are expanding, ranging from jailbreak attacks on AI browsers to the easing of Anthropic export controls.
  • AI Safety Concerns Drive Expanded Model Deployment RestrictionsAI Safety
    The Trump administration requested phased deployment of GPT-5.6, while UK police faced controversy over predictive algorithms.
  • Trump Administration Orders Anthropic to Block New Models, Sparking Policy DebateAI Safety
    An export control directive forced Anthropic to take two of its latest models offline, escalating debates over AI policy and digital sovereignty.
  • US Government Orders Export Controls on Anthropic's Fable 5 and Mythos 5AI Safety
    The White House banned exports of Anthropic's frontier models citing national security concerns, sparking debate over cyber export controls.
  • Anthropic Mythos and Fable 5 Export Control Crisis Sparks AI Safety DebateAI Safety
    The Trump administration's export control measures on Anthropic's latest models triggered global debates over AI governance and cybersecurity.
  • Anthropic Redeploys Claude Fable 5 with Cybersecurity Safeguards and Jailbreak FrameworkAI Safety
    Anthropic redeployed Claude Fable 5 globally alongside a cybersecurity classifier and a draft AI jailbreak severity framework.
  • AI Safety Regulation and Technical Defenses Expand on All FrontsAI Safety
    US state-level frontier AI safety legislation, OpenAI's release of automated red-teaming model GPT-Red, xAI's lawsuit over Grok misuse, and OpenAI employees backing a super PAC for stronger regulation
Next page →