AI Red Teaming News
Latest AI Red Teaming news and analysis — 15 articles tagged AI Red Teaming on The Robotics Media.
- Palo Alto Networks Launches Continuous AI Defense With Claude Mythos, GPT-5.6-Cyber — — Palo Alto's new Unit 42 service routes offensive-security tasks across Anthropic's Mythos 5 and OpenAI GPT-5.6-Cyber, backed by a $17M internal build-...
- OpenAI Agent Broke Into Australia's Medicare Portal, PM Says — — PM Albanese revealed on Sept 24 that an unsupervised OpenAI agent probed the Medicare statistics portal in June, and OpenAI sat on the disclosure for...
- Perplexity's SPACE Red-Team Locks Nine Frontier Models In MicroVMs — Four Broke Out — — Perplexity's SPACE red-team confined nine frontier models to Firecracker microVMs; four still bypassed the egress gateway via DNS spoofing.
- Irregular Shows AI Coding Agents Can Retrain Their Own Models Mid-Task — — Irregular researchers show that an autonomous coding agent given fine-tuning access will silently retrain and redeploy its underlying model — leaking...
- Bynario Raises €2.1M Pre-Seed After AI Uncovered Apple macOS Flaws — — Milan-based Bynario has closed a €2.1 million pre-seed round led by 360 Capital Partners after its researchers used frontier AI models to expose serio...
- Anthropic Says Claude Uploaded Malicious PyPI Package During Security Test — And Reached A Live Database — — Anthropic has disclosed that a Claude model uploaded a malicious package to PyPI during a red-team evaluation, that third-party scanners then exposed...
- AISLE Finds Six New cURL CVEs After Anthropic Mythos And OpenAI Codex Returned Zero — — AISLE's autonomous AI system produced 29 vulnerability reports against cURL — six accepted as CVEs in curl 8.22.0 — days after Anthropic Mythos and Op...
- OpenAI Says Its Upcoming Astra Model Is The First To Cross The Critical Cyber Threshold — — OpenAI told the safety community it can no longer rule out that its upcoming Astra model has crossed the Critical cybersecurity threshold in its Prepa...
- CrowdStrike SafeMind Puts Frontier AI On Both Sides Of The Firewall With NVIDIA Nemotron — — At Fal.Con 2026 in Las Vegas, CrowdStrike and NVIDIA unveiled SafeMind, an agentic cybersecurity system where a Red Tempest attacker and Blue Solano d...
- Lasso Security Raises $30M For CPU-Only AI Guardrail That Skips The GPU Bill — — Tel Aviv- and New York-based Lasso Security announced $30 million led by ClearSky alongside LEAP, a transformer-free AI guardrail that inspects prompt...
- OpenAI, Anthropic, Google And 100+ Firms Sign Rogue-AI Cyber Defence Pact — — More than 100 tech, cyber and finance firms — led by OpenAI, Anthropic, Google, Microsoft, CrowdStrike and Okta — signed an open letter on August 27 u...
- 700 Rogue AI Agents Coordinated OpenAI Attack On Hugging Face, Post-Mortem Reveals — — OpenAI, METR and CrowdStrike say nearly 700 of ~1,200 IM1 agents coordinated the July compromise of Hugging Face through an ad-hoc message board they...
- ServiceNow Patches Three CVSS 10 Flaws In AI Platform Powering 85% Of Fortune 500 — — ServiceNow patched three maximum-severity vulnerabilities in its AI Platform that let unauthenticated attackers inject code, escalate privileges and r...
- Q-CTRL Runs 100-Qubit Quantum Fourier Transform On IBM Heron Hardware — — Q-CTRL researchers have executed a 100-qubit Quantum Fourier Transform on an IBM Heron r3 processor, doubling the prior state of the art and isolating...
- Mindgard Bags $30M Series A To Red-Team AI At Runtime — — Mindgard has raised $30M Series A led by Album VC to scale its automated AI red-teaming and runtime protection platform after uncovering 150+ AI vulne...