Anthropic and OpenAI propose embedding third-party safety evaluators inside AI companies
The plan could reshape AI governance globally, but major firms like Meta and Google DeepMind have not yet committed. OpenAI and Anthropic aim to set a precedent, though the path to adoption remains unclear.

Anthropic and OpenAI have proposed embedding third-party evaluators within AI companies to assess model safety and report findings publicly. The idea, outlined in a detailed essay by Anthropic CEO Dario Amodei, suggests that these evaluators would have the authority to review AI systems, identify alignment issues, and share unfiltered results with the public. This marks a significant shift in how the AI industry approaches safety oversight, as it moves away from self-regulation toward external scrutiny.
The proposal comes amid growing concerns about the risks posed by advanced AI systems. Amodei emphasized that external evaluators could provide an objective perspective on safety and alignment, reducing the influence of corporate interests. OpenAI’s CEO, Sam Altman, has also expressed support for the initiative, suggesting that the practice could become a standard in the industry. However, major players like Meta and Google DeepMind have not yet committed to the plan, despite some interest from DeepMind’s leadership.
Amodei highlighted that the ideal scenario would involve strong regulatory frameworks mandating such oversight. He argued that without external evaluators, companies might alter their safety policies in response to public relations crises, undermining long-term safety goals. The proposal has been met with cautious optimism, though the timeline for implementation remains unclear. Some industry observers suggest that the plan could take years to materialize, given the need for consensus among major AI firms and regulators.
The potential consequences of embedding third-party evaluators are significant. Companies may face increased operational costs and the risk of reputational damage if safety issues are exposed. There is also concern about potential vendor lock-in, as firms may rely on specific evaluators for compliance. Governance challenges could arise, particularly in ensuring that evaluators remain independent and free from corporate influence. Market reactions have been mixed, with some stakeholders viewing the move as a necessary step toward accountability, while others warn of the complexities involved in implementation.
The proposal remains in its early stages, with no clear roadmap for adoption. While OpenAI and Anthropic have signaled support, broader industry participation is still uncertain. The success of the initiative will depend on the ability to establish trust between companies, evaluators, and regulators. As the debate continues, the AI industry faces a critical decision: whether to embrace external oversight or continue with self-regulation, despite the risks.