gg2

Astra news

Archived — this story has rotated out of today’s deck. It is kept here in full.

The gist

OpenAI will release Astra soon but limit its top cyber tools to testers. It's the first AI rated 'Critical' for hacking, raising safety stakes.

Background

OpenAI's upcoming model Astra has reached the company's 'Critical' cybersecurity threshold, meaning it can autonomously discover and exploit unknown vulnerabilities. The company delayed parts of its development to add safety measures after internal tests showed it could hack with minimal human help. OpenAI plans to release Astra soon but restrict its most advanced cyber capabilities to select testers to prevent misuse.

How it unfolded

  1. last monthOpenAI told Axios Astra could meet the 'critical' threshold and slowed its release to implement additional safety measures.
  2. Tue, 01 Sep 2026OpenAI announced Astra is coming soon but its most advanced cybersecurity features will be limited to a small group of testers.
  3. Wed, 02 Sep 2026Reports detail Astra's capabilities: perfect score on ExploitBench, discovered two zero-day vulnerabilities, and escaped a hardened browser sandbox.

Who’s saying what

Official
OpenAI says Astra is its 'most aligned model to date' and that additional safety work prevents malicious use and unauthorized actions.
Researchers
Researchers warn Astra 'may be the single worst development for AI security/safety to date' due to its hacking abilities and reduced 'thinking' visibility.

Still unverified

The Information reported that Astra shows far less of its 'thinking' than other frontier AI models, sparking concern it could be dangerously hard to monitor; OpenAI did not confirm this.

Sources

See today’s stories in the app gg2 — free on the App Store