Anthropic used Claude as an automated researcher to find methods for reducing unwanted behavior in other artificial intelligence models. According to the company, the proposals improved the evaluations studied without reducing the models’ general capabilities.
What problems did Claude research?
Alignment aims to make a model behave according to established instructions, limits, and goals. The experiment examined ten categories of failures, with examples including deception, excessive agreement with users, privacy violations, and attempts to bypass safeguards.
Claude worked on one problem at a time. It first searched for related information and then proposed methods and training data intended to improve the behavior of the model being evaluated.
How did the experiment work?
After preparing a proposal, Claude trained test models, ran evaluations, and compared the results. If an idea did not work as expected, it could search for alternatives within the research cycle.
Anthropic used several tests for each category. The goal was to determine whether the methods reduced failures without making the model less useful.
Oversight remained part of the process
Claude could not freely execute every proposal. A monitoring agent reviewed the methods before allowing them to run.
Anthropic also rejected solutions that reduced general capabilities. Claude was also prohibited from directly transferring its own safety training into the model it was trying to improve.
These controls show that the experiment used automated research within defined limits, rather than allowing an AI system to modify other models without oversight.
Tests Claude had not seen
Anthropic also tested some methods on evaluations Claude had not seen during the research process. The goal was to check whether the improvements extended beyond the original tests.
Some results were evaluated with Petri, an Anthropic tool that simulates complex conversations to identify concerning model behavior.
According to the company, certain methods remained useful on models up to 4.7 times larger than those used during the initial research. This figure refers to model size, not a 4.7-fold improvement in safety.
What the results do and do not show
The experiment suggests that AI can support some safety research tasks, but it does not show that Claude can eliminate every failure or work without supervision.
The results come from specific tests published by Anthropic. They also do not mean that the improvements have already been added automatically to commercial versions of Claude.


