344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support132
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Anthropic researchers report automated systems that improved an AI model on 10 safety benchmarks without hurting overall performance.
In short: Anthropic published research showing automated AI systems can help fix certain unsafe behaviors in an AI model, using a set of safety tests.
Anthropic released a paper called “Automated Researchers Can Reliably Mitigate Alignment Failures.” It describes an “Automated Alignment Researcher,” which is a set of AI systems designed to help train another AI model to behave more safely.
In this work, the automated systems were given 10 benchmarks, meaning 10 specific tests that measure unwanted behavior. The paper says the systems improved the model’s performance on every one of those tests, and it did not make the model worse overall.
The automated systems worked in a loop that looks like a simplified research process. They searched existing writing on the topic, suggested a training method, then trained the model for about 30 minutes at a time. Over several rounds, the system kept methods that helped and dropped methods that did not (like trying many study strategies and keeping the ones that raise your score).
The paper also compares speed and cost. It claims the best automated method beat what experienced humans suggested, on average within six hours. It also estimates the automated approach costs about $4 per hour in API inference (paying to run AI through a service), compared with $150 per hour for human researchers.
A lot of people worry about AI systems doing things their makers did not intend. This research suggests some parts of AI safety work might be automated, at least when clear tests exist. Anthropic also notes a key limitation: the system only helps as much as the benchmarks reflect the real safety goal, and building and maintaining those tests is still hard.
Source: TechCrunch AI