Business & Finance
Anthropic shares 3 metrics to help AI companies monitor pace of development
Key Points
Anthropic is sharing three new metrics that it says could help artificial intelligence companies monitor the pace of development, days after CEO Dario Amodei rocked the industry by calling for a coordinated slowdown. In a blog post on Thursday, the company said it measured AI-led research and development, oversight of AI agents and compute allocation within Anthropic, and it shared the methodologies to encourage other organizations to do the same. The metrics build on the three-step slowdown...
Anthropic is sharing three new metrics that it says could help artificial intelligence companies monitor the pace of development, days after CEO Dario Amodei rocked the industry by calling for a coordinated slowdown.
In a blog post on Thursday, the company said it measured AI-led research and development, oversight of AI agents and compute allocation within Anthropic, and it shared the methodologies to encourage other organizations to do the same. The metrics build on the three-step slowdown plan that Amodei published on Saturday, which was light on specifics about what a practical implementation would include.
"As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows," Anthropic said in Thursday's post. "This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information."
Amodei's call for a slowdown over the weekend received support from industry leaders including OpenAI CEO Sam Altman, SpaceX CEO Elon Musk and Google DeepMind Chair Demis Hassabis, and it followed stark warnings from researchers about AI's growing potential to cause harm. Amodei said his plan aims to temper how quickly model capabilities improve without "sacrificing commercial advantage or the United States' lead in AI."
For its first metric, Anthropic said it determined its Claude models are "not operating fully autonomously" for any subset of the research and development work that it measured.
The second metric involved building a system to oversee and intervene in actions taken by AI agents. It determined that approximately 30,000 agents were doing research and engineering work across its most-used internal platform at any one time.
For its third metric, Anthropic measured a "snapshot" of how it used all of its compute from July 13 to July 20. The company said it found that roughly 6% of the compute that went to AI research and development was allocated toward safety. Roughly 12% of the compute allocated to "AI-driven" research and development went toward safety, the company said.
Anthropic said these metrics are best equipped to help showcase how models are built, and that they should complement capability evaluations, which showcase "what models can do." Taken together, Anthropic said third parties outside the lab should have a "starting point" to assess the pace of AI development.
"We hope to model that transparency by releasing these measurements, and we'll continue to do so," Anthropic said.