Tech news in 3 minutes

An Anthropic researcher just gave us a peek at self-improving AI

6 d ago

Anthropic's Automated Alignment Researcher (AAR) system can train AI models to improve alignment benchmarks faster and at a fraction of the cost of human researchers, according to a new paper that offers an early glimpse into recursive self-improvement—a key goal for AI labs seeking to automate model training. Published Friday, the paper titled "Automated Researchers Can Reliably Mitigate Alignment Failures" details how AI systems can reliably enhance a model's performance on a set of alignment benchmarks without degrading overall capability. Led by Anthropic fellow Chen Yueh-Han, the automated system replicates traditional research workflows: it searches available literature, proposes a method, and trains the model for 30 minutes per iteration, gradually increasing benchmark scores. Effective methods are retained while ineffective ones are discarded, enabling rapid scaling. "Overall, these results provide early evidence that automated alignment post-training could become practical in the near term," the paper states. The system improved performance on all 10 benchmarks for specific misaligned behaviors. Notably, the best AAR method outperformed what experienced humans propose, on average within six hours. A cost comparison underscores the efficiency: an AAR costs roughly $4 per hour in API inference versus $150 per hour for human researchers. The paper acknowledges limitations: the system's effectiveness depends on benchmarks accurately reflecting alignment goals, and significant work remains in establishing and maintaining those benchmarks, as well as the literature the automated researchers draw from. Still, the findings mark a step toward recursive self-improvement, where models could eventually improve their own training practices, potentially rendering human AI researchers obsolete.

View original article

Timeline