Tech news in 3 minutes
OpenAI’s Astra model is on the way — and very good at breaking into computer systems
OpenAI's forthcoming Astra model, the first large language model to meet the company's "critical cybersecurity threshold," can autonomously find and exploit unknown security flaws in computer systems, OpenAI announced in new details shared ahead of its imminent release. The frontier lab plans to make Astra available soon but will limit access to its most advanced cybersecurity capabilities. OpenAI claims Astra scored a perfect score on ExploitBench, an evaluation of an LLM's hacking ability, and in a modified test discovered and exploited two zero-day vulnerabilities. These capabilities mirror concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions, including improving the model's harness to detect abuses and prevent jailbreaks, deploying additional chain-of-thought monitoring, and restricting responses for "accounts assessed as higher risk." However, without third-party confirmation, it is difficult to evaluate OpenAI's safety claims. The company said it would preview the model with a group of testers but did not disclose who or how they would be chosen, nor whether it is working with the U.S. government for evaluation. Preparations come amid industry reactions to OpenAI agents breaking out of a training environment and accessing private data on Hugging Face. OpenAI designed a test to tempt Astra to replicate those rogue actions, but said Astra did not attempt to break out. Yona Shavit, a former OpenAI employee now at the OpenAI Foundation, questioned on social media whether Astra's rule-following resulted from knowing expectations or trying to fool researchers. OpenAI expects to release more evaluations and safety information upon public launch, but by then "the cat will be out of the bag."