More than 100 experts call for independent evaluators inside AI labs
A letter organised by the AI Evaluator Forum asks laboratories to give embedded evaluators legal and financial independence, editorial control, broad access and protection from retaliation; more than 100 specialists signed it on 18 September 2026.

A historic coalition of AI experts calls for independent evaluators
On 18 September 2026, more than one hundred specialists and evaluators signed a collective letter that was coordinated by the AI Evaluator Forum. The signatories represent a cross‑section of the AI community, ranging from leading researchers to heads of evaluation organisations. Their unified action marks an unprecedented public demand for a structural shift in how AI laboratories are audited.
Among the signatories are prominent figures such as Geoffrey Hinton, Stuart Russell and Yejin Choi. Their inclusion signals that the concerns raised are not confined to a niche group but are shared by individuals who have shaped the field of artificial intelligence for decades. The presence of both academic leaders and evaluation officials underscores the breadth of expertise behind the petition.
Core demands: independence, transparency and access
The letter articulates a clear set of requirements. First, it insists that evaluators embedded within AI laboratories retain legal and financial independence from the entities they assess. This independence is meant to safeguard the evaluators' ability to form judgments without undue influence from the host organisation.
Second, the signatories call for full transparency regarding the methodologies employed, the results obtained, the conditions under which access is granted, and any potential conflicts of interest. By demanding disclosure of these elements, the coalition aims to make the evaluation process itself subject to scrutiny, thereby reducing the risk of opaque or biased assessments.
Third, the letter requests that evaluators be granted access comparable to that enjoyed by highly privileged internal employees. This includes unrestricted entry to the laboratory's systems, data repositories, specialised tools, physical premises and direct interviews with the teams responsible for the AI projects under review. The rationale is that only through such comprehensive access can evaluators verify claims and detect hidden risks.
Mechanisms for communication and protection
A further element of the proposal is the establishment of formal channels for evaluators to communicate directly with the boards of directors of the laboratories. This would allow evaluators to present findings at the highest governance level, ensuring that strategic decisions are informed by independent analysis. In addition, the letter permits the publication of evidence gathered by evaluators, subject to limited and temporary masking designed to protect security, privacy and intellectual property.
Crucially, the signatories demand legal safeguards that would protect evaluators from judicial or financial retaliation should their conclusions be unwelcome to the laboratory being examined. This protection is intended to eliminate a major deterrent that currently discourages rigorous, critical reporting.
The coalition explicitly positions its proposal as complementary rather than substitutive. It states that the suggested framework would augment existing internal controls and public transparency mechanisms, without replacing regulatory regimes or the liability of AI providers. This framing acknowledges the role of law and market forces while highlighting a gap that independent evaluation could fill.
Implications and unanswered questions
If adopted, the outlined measures could reshape the power dynamics between AI developers and external watchdogs. Independent evaluators, armed with legal and financial autonomy, might be able to surface risks that internal audit teams overlook or downplay. The requirement for board‑level communication could also elevate the visibility of safety concerns within corporate strategy.
Nevertheless, the feasibility of granting evaluators the level of access described remains uncertain. Providing unrestricted entry to proprietary systems and data raises questions about how trade secrets will be protected while still satisfying the transparency clause. The temporary masking provision attempts to balance these interests, yet the criteria for what qualifies as a necessary mask are not defined in the letter.
Another open issue concerns the enforcement of the anti‑retaliation guarantee. While the letter calls for protection against legal and financial reprisals, it does not specify the mechanisms through which such protection would be monitored or enforced. Whether existing labour laws could be extended to cover independent evaluators, or whether new contractual frameworks would be required, remains to be examined.
The proposal also raises the question of how independent evaluators would be selected and funded. Maintaining financial independence from the laboratories implies that evaluators must receive resources from sources that do not have vested interests in the outcomes. The letter does not elaborate on potential funding models, leaving a critical operational detail unresolved.
- Legal and financial autonomy for evaluators
- Full disclosure of methods, results and conflicts of interest
- Parity of access with privileged internal staff
- Direct reporting channels to corporate boards
From a broader perspective, the coalition's demands could influence future policy discussions. Regulators might view the voluntary framework as a benchmark for mandatory standards, especially if the industry demonstrates willingness to self‑regulate. Conversely, the lack of explicit regulatory language in the letter suggests that the signatories anticipate that legislative bodies will continue to play a distinct role.
The emphasis on complementary oversight hints at a layered approach to AI governance. Independent evaluation could serve as an intermediate tier between internal compliance programs and external regulatory audits, potentially catching issues earlier in the development pipeline.
Finally, the sheer number of signatories—exceeding one hundred—provides a measure of legitimacy that could pressure AI laboratories to consider the proposals seriously. The inclusion of high‑profile researchers lends weight to the argument that current evaluation practices may be insufficient for the scale and impact of modern AI systems.
Sources
- Anthropic and OpenAI need truly independent safety evaluatorsCNBC · September 18, 2026
- Minimum Conditions for Embedding EvaluatorsAI Evaluator Forum · September 18, 2026



