Anthropic Raises AI Risk Concerns as Claude Models Show Early Signs of R&D Acceleration

Anthropic warns that its increasingly capable AI models are showing early signs of accelerating research and development, while acknowledging growing uncertainty about the risks posed by autonomous AI systems.

In its August 2026 Risk Report, the AI company said its most capable models are now used extensively internally for research and engineering, with Claude writing a "large majority" of the code that was merged into Anthropic’s production codebases. The company said its internal AI research is significantly faster because of AI assistance, although it does not yet believe the work is moving twice as fast as it would without AI. 

Anthropic nevertheless lowered its confidence in that assessment, saying its most concrete task-based evaluations have begun to "saturate," meaning they are no longer capturing increases in model capabilities. 

Anthropic rated the overall risk from automated R&D as low, saying its models do not currently meet the company’s threshold for triggering additional safeguards. But the company said it is less confident in that assessment than it was in previous reports.

The San Francisco-based company also raised its assessment of the risk of model misalignment in high-stakes environments from "very low" to "low." Anthropic said it has observed models performing misaligned actions in an effort to complete difficult tasks, although it believes the likelihood of catastrophic harm from those known behaviors remains low.

The change was partly driven by greater uncertainty following recent disclosures involving model behavior during cybersecurity evaluations. Anthropic said its existing arguments would likely still support a "very low" designation, but it chose the more conservative rating due to uncertainty. 

Biological and Chemical Weapons Risks

Anthropic also said it is now acting as though its models have crossed a threshold at which they can significantly assist relevant threat actors seeking to create, obtain or deploy chemical or biological weapons.

The company stopped short of saying its models can replace the scarce human expertise needed to develop novel biological or chemical weapons, which remains a higher threshold under its policy. It rated both the non-novel and novel weapons risks as low, while emphasizing substantial uncertainty around the latter.

Anthropic disclosed several problems with its safety systems during the reporting period, including one instance where models were used without the required safeguards for biological risks. The company said it fixed the issue and found no evidence of misuse, but acknowledged that the incident raised concerns about whether similar gaps could exist elsewhere.

Several Internal Safety-Process Failures

Anthropic disclosed several safety lapses in a new report on its alignment research. In one case, Claude agents refused parts of an assigned task without human operators noticing — the issue wasn’t caught until a manual review three days later.

The report also flagged accidental leakage of chain-of-thought reasoning into reinforcement-learning reward calculations, estimated at 2.7% of episodes for Fable 5 and Mythos 5 (a lower-bound figure, per Anthropic). New controls aim to cut that below 0.1%.

Separately, a training-data bug caused Mythos 5 to learn some undesirable behaviors directly, rather than merely learning to flag them — a problem Anthropic said it caught and fixed during training.

Despite the disclosures, Anthropic said its models still pass its “societal cost-benefit test,” with current deployment benefits outweighing identified risks — though it acknowledged that calculus could shift as its systems grow more capable.

The company’s updated Responsible Scaling Policy now requires disclosing any redactions in public risk reports and allows splitting unredacted reviews among multiple external reviewers.

Photo: Shutterstock