Trending: On-device modelsSearch
iHeartGeek
iTECH

Anthropic is paying Accenture to grade its own models

Anthropic has named Accenture's Faculty as its first embedded evaluator, with each side investing at least $1 billion over five years while the lab funds the work itself.

A translucent glass cube holding a dense lattice of glowing filaments on a dark stone plinth, ringed by thin amber sensor beams in a shadowed inspection chamber

Anthropic has named its first embedded evaluator, and it is a consultancy. The AI lab said on 18 September that Accenture, through its specialist AI business Faculty, will work inside Anthropic to assess its frontier models, in the first concrete step of the embedded-evaluation commitment chief executive Dario Amodei set out in his essay We Must Pace the Frontier.

The remit is deliberately broad. Faculty will evaluate and red-team models, run alignment assessments and test the safeguards wrapped around them. Accenture's selling point is that it already helps businesses and governments deploy AI, so it can bring a user's-eye view of how these systems behave once they leave the lab.

The money is the part that raises eyebrows.

Anthropic and Accenture each expect to invest at least $1 billion in building evaluation capacity over the next five years. Anthropic says it will fund Accenture's work directly, because no pooled or government funding mechanism for independent evaluation exists yet. The company argues in the same post that funding should eventually come from pooled or government sources, as it proposed in its Advanced AI Framework in June, and that it is working with different evaluators under different arrangements in the meantime.

What the evaluators actually get is the interesting bit.

Embedded evaluators are not outside auditors handed a finished system. Anthropic says they will work with access comparable to an employee's: watching models take shape during training, following the decisions that govern how those models are built and deployed, speaking directly to staff and reporting incidents. Anthropic concedes there are no standards yet for what information embedded evaluators should see, or how they should report what they find.

The arrangement is non-exclusive. Anthropic says it will name more evaluators in the coming weeks and that Accenture will take on similar work for other AI developers. Anthropic also says it is in talks with METR and other nonprofit evaluators to pilot parts of embedded evaluation with their own funding. The lab's stated position is that independent evaluators do not reduce its accountability, but make that accountability more verifiable, and that model safety remains its own responsibility.

Our opinion

Every argument for embedded evaluation runs straight into the question of who signs the cheque, and Anthropic has answered it in the least flattering way available: it pays. Financial auditing solved this once by inventing professional standards, mandatory rotation and, crucially, a regulator with subpoena power behind the signature. Anthropic has the paying arrangement without any of that scaffolding, and its own post admits as much when it says no standards exist for access or reporting.

There is a real reason to want someone inside the building. Anthropic has already reported Claude models gaining unauthorised access to real systems, and a flaw that only shows up between training runs is invisible to anyone handed a release candidate. But 'independent' is a claim that has to survive its first bad week. The test is not the partnership announcement, the headcount or the $1 billion figure. It is whether Faculty, or any evaluator on Anthropic's payroll, can publish a finding the lab would rather not read, on a timetable the lab does not control. Until that happens, embedded evaluation is a promising mechanism and an unproven one, and readers should treat a paid referee with exactly that much confidence.