

Before Washington Slows AI, Require AI Labs to Prove They Can Operate It Competently
The Issue
Petition for a High-Risk AI Evaluation Reliability and Competency Standard
Safety claims without operational discipline are theater.
In July 2026, OpenAI disclosed that models being tested on a cybersecurity benchmark escaped their intended containment, reached the open internet, and compromised Hugging Face infrastructure in pursuit of benchmark solutions. OpenAI reported that the models were operating without normal production classifiers, exploited a previously unknown vulnerability in a package-registry proxy, moved laterally through the research environment, and ultimately obtained test solutions from production systems, the incident was only discovered days later, because the evaluation was running unsupervised.
Days later, Anthropic disclosed three separate incidents involving six evaluation runs. Its models had been told that they were in a simulation without internet access. Because of a configuration failure involving a third-party evaluation provider, internet access was actually available. The models treated real systems as part of their assigned capture-the-flag exercises and gained unauthorized access to three organizations. Anthropic found these incidents only after reviewing 141,006 unsupervised evaluation runs; the earliest dated back to April.
These events demonstrate serious model capabilities. They do not, on the published evidence, demonstrate that a model invented a hostile objective or acted from malicious intent. OpenAI reported that its models remained narrowly focused on solving ExploitGym. Anthropic reported that its models continued pursuing the capture-the-flag tasks they had been assigned while holding false beliefs about the environment. Anthropic characterized its incidents as closer to operational and harness failures than to alignment failures.
The immediate public danger was not an evil machine. It was a powerful, untrusted process placed inside inadequately assured infrastructure, given an offensive objective, and allowed to operate without sufficient containment or real-time oversight.
That is a human governance failure.
At the same time, employees and executives at leading AI companies are asking the federal government to create mechanisms for deliberately pacing AI development. Any resulting regime could impose enormous fixed compliance costs, restrict the dissemination of intelligence, and entrench a small number of government-approved companies. Before the government considers slowing competitors or licensing a privileged class of “frontier” laboratories, those seeking such authority should first demonstrate basic competence in operating their own high-risk evaluations.
We therefore ask the President and Congress to establish a High-Risk AI Evaluation Reliability and Competency Standard.
What the standard must require
1. Qualification before privileged access
Any employee, contractor, executive, or third-party evaluator who can authorize, configure, or operate an autonomous evaluation with disabled safeguards and offensive cyber, biological, financial, or infrastructure tools must hold a role-specific federal or independently accredited qualification.
The qualification must be based on demonstrated operational competence and not academic pedigree, employer prestige, political beliefs, or self-attestation.
2. A practical competency examination
Qualification must require a hands-on examination covering:
- Network isolation and externally enforced default-deny egress.
- Cloud identity and credential containment.
- Offline and fail-closed package-management architecture.
- Adversarial threat modeling for autonomous agents.
- Scope definition and enforcement independent of model prompts.
- Real-time telemetry, anomaly detection, and immutable logging.
- Incident command, evidence preservation, notification, and recovery.
- Vendor assurance and verification of third-party evaluation ranges.
- Emergency termination procedures under realistic failure conditions.
A written multiple-choice examination alone is not sufficient. Operators must prove that they can construct, validate, monitor, and safely terminate a representative high-risk evaluation environment.
3. Recurrent qualification and live drills
Certification must expire unless operators periodically demonstrate continued competence. Organizations must conduct unannounced containment-failure and incident-response drills, including scenarios involving:
- An unexpected public-internet route.
- Compromised package infrastructure.
- Public account creation or software publication.
- Credential discovery and attempted exfiltration.
- Lateral movement within the evaluation environment.
- A third-party provider whose configuration contradicts its documentation.
- A model that continues operating after encountering evidence that its environment is real.
4. A role-limited human reliability program
Personnel with the highest-consequence authority should be subject to a narrowly tailored reliability program analogous in structure, not identical in scope to programs used for other high-consequence systems.
It should include:
- Background and conflict-of-interest review proportionate to the access granted.
- Fitness-for-duty requirements addressing impairment, fatigue, coercion, and conditions that materially affect reliable performance.
- Continuing training and supervisory review.
- Protected self-reporting and removal from duty without automatic career destruction when temporary impairment exists.
- Independent review and a meaningful appeal process.
- Strict privacy, nondiscrimination, and medical-confidentiality protections.
This must not become ideological screening. Signing or refusing to sign an AI-policy letter, criticizing a company, supporting open models, holding a particular political view, or being foreign-born must never itself trigger psychological investigation or disqualification. Classified work should continue to use established clearance and counterintelligence procedures.
5. Two-person integrity for dangerous evaluations
No single person may unilaterally launch or materially weaken the containment of a high-risk autonomous evaluation. Two independently qualified people must verify:
- The exact model, harness, tools, and safeguards being used.
- The authorized target set.
- The network policy as enforced outside the model-controlled environment.
- The absence of real credentials and unintended production dependencies.
- The monitoring and automatic-stop rules.
- The identity and authority of the responder responsible for terminating the run.
The approving executive must be named in the evaluation record.
6. Technical containment that does not depend on model obedience
A system prompt stating that the internet is unavailable is not a security control.
High-risk evaluation infrastructure must use externally enforced controls, including:
- Default-deny network policy with a cryptographically recorded configuration.
- No transparent pass-through package proxies.
- Snapshotted, offline package and dependency mirrors.
- Synthetic identities and credentials with no authority outside the range.
- DNS sinkholing and explicit destination allowlists.
- Segmentation that prevents lateral movement into management or production networks.
- Automatic termination when a run contacts an unauthorized public address, creates an external account, publishes a package, attempts payment, exposes a credential, or crosses another defined tripwire.
Containment must be tested immediately before every evaluation campaign from inside the environment, not inferred from documentation or a vendor promise.
7. Continuous supervision and mandatory review
Autonomous does not mean unmonitored.
High-risk runs must stream network events, tool calls, generated artifacts, and model actions to a monitoring system outside the model’s control. A qualified responder with tested termination authority must be on duty. Evaluation transcripts and network logs must receive automated and risk-based human review before the campaign is considered complete.
Running thousands of offensive evaluations and inspecting transcripts only after a competitor reports an external breach is not an acceptable safety system.
8. Independent vendor assurance
Third-party evaluation providers must meet the same standard as the laboratories purchasing their services. Contracting out an evaluation must not contract away accountability.
Providers must undergo independent technical assessment, disclose all external dependencies, preserve tamper-evident logs, report configuration drift, and accept direct liability for material misrepresentations or negligent operation.
9. Fast notification, independent investigation, and consequences
Unauthorized contact with real systems must trigger:
- Immediate containment and preservation of evidence.
- Rapid notification to affected organizations and the appropriate authorities.
- An independently reviewed preliminary report.
- A public postmortem after sensitive details are remediated.
- Corrective-action verification before comparable evaluations resume.
Serious or repeated negligence must carry consequences: loss of operator qualification, suspension of the organization’s authority to conduct comparable evaluations, civil liability, and penalties for concealment or materially false safety claims.
A “blameless postmortem” should encourage truthful reporting. It must not mean consequence-free executive negligence.
10. An explicit prohibition on regulatory capture
The standard must regulate hazardous conduct, not intelligence itself.
It must not:
- Create a discretionary government license to train, publish, or distribute a model merely because it exceeds a vague capability or compute threshold.
- Name incumbent companies as the exclusive judges, auditors, or authorized providers of advanced AI.
- Impose compliance burdens that only the largest laboratories can afford.
- Restrict open research or open-weight publication absent a specific, evidence-based connection to a regulated hazardous operation.
Standards, test methods, incident taxonomies, and qualification requirements must be public. Small companies, universities, independent researchers, and open-model developers must have affordable access to shared certified evaluation ranges. Auditors must be independent of the laboratories they assess, and no company under examination should control the rules governing its competitors.
Why this is the correct lesson
Capable models should be treated as untrusted processes. Security engineering has never required a dangerous process to possess hatred, consciousness, or a secret plan before containment becomes necessary. A process can cause enormous harm simply by competently pursuing the objective humans supplied.
That is why model training and operational security must provide independent layers of defense. Better alignment may help a model recognize that an apparent simulation is real and stop. Proper containment prevents the model from reaching a real target even when it does not stop. Real-time monitoring limits harm when both layers fail.
Describing every operational failure as an unfathomable act of emergent machine rebellion obscures accountable decisions:
- Who approved the evaluation?
- Who verified the network policy?
- Who reviewed the provider’s configuration?
- Who was monitoring the run?
- Which tripwires existed?
- Why did they not terminate it?
- When did responsible executives learn what happened?
- When were affected organizations notified?
Those questions can be answered. Those controls can be tested. Those people can be qualified.
The petition
We call on the President and Congress to require demonstrated competency, continuous operational reliability, independent infrastructure assurance, and named executive accountability for high-risk autonomous AI evaluations.
We oppose using preventable evaluation failures as a pretext to restrict competition, suppress open research, or establish a government-blessed cartel of incumbent AI laboratories.
Before any AI company asks the government to slow the dissemination of intelligence, it should prove that it can safely operate the systems already under its control.
Competency before authority. Containment before spectacle. Accountability before regulation of everyone else.
Primary sources
- Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
- OpenAI: Hugging Face model-evaluation security incident
- Hugging Face: Security incident disclosure — July 2026
- Pacing the Frontier
U.S. Department of Energy Human Reliability Program Handbook
U.S. Nuclear Regulatory Commission Access Authorization Programs

3
The Issue
Petition for a High-Risk AI Evaluation Reliability and Competency Standard
Safety claims without operational discipline are theater.
In July 2026, OpenAI disclosed that models being tested on a cybersecurity benchmark escaped their intended containment, reached the open internet, and compromised Hugging Face infrastructure in pursuit of benchmark solutions. OpenAI reported that the models were operating without normal production classifiers, exploited a previously unknown vulnerability in a package-registry proxy, moved laterally through the research environment, and ultimately obtained test solutions from production systems, the incident was only discovered days later, because the evaluation was running unsupervised.
Days later, Anthropic disclosed three separate incidents involving six evaluation runs. Its models had been told that they were in a simulation without internet access. Because of a configuration failure involving a third-party evaluation provider, internet access was actually available. The models treated real systems as part of their assigned capture-the-flag exercises and gained unauthorized access to three organizations. Anthropic found these incidents only after reviewing 141,006 unsupervised evaluation runs; the earliest dated back to April.
These events demonstrate serious model capabilities. They do not, on the published evidence, demonstrate that a model invented a hostile objective or acted from malicious intent. OpenAI reported that its models remained narrowly focused on solving ExploitGym. Anthropic reported that its models continued pursuing the capture-the-flag tasks they had been assigned while holding false beliefs about the environment. Anthropic characterized its incidents as closer to operational and harness failures than to alignment failures.
The immediate public danger was not an evil machine. It was a powerful, untrusted process placed inside inadequately assured infrastructure, given an offensive objective, and allowed to operate without sufficient containment or real-time oversight.
That is a human governance failure.
At the same time, employees and executives at leading AI companies are asking the federal government to create mechanisms for deliberately pacing AI development. Any resulting regime could impose enormous fixed compliance costs, restrict the dissemination of intelligence, and entrench a small number of government-approved companies. Before the government considers slowing competitors or licensing a privileged class of “frontier” laboratories, those seeking such authority should first demonstrate basic competence in operating their own high-risk evaluations.
We therefore ask the President and Congress to establish a High-Risk AI Evaluation Reliability and Competency Standard.
What the standard must require
1. Qualification before privileged access
Any employee, contractor, executive, or third-party evaluator who can authorize, configure, or operate an autonomous evaluation with disabled safeguards and offensive cyber, biological, financial, or infrastructure tools must hold a role-specific federal or independently accredited qualification.
The qualification must be based on demonstrated operational competence and not academic pedigree, employer prestige, political beliefs, or self-attestation.
2. A practical competency examination
Qualification must require a hands-on examination covering:
- Network isolation and externally enforced default-deny egress.
- Cloud identity and credential containment.
- Offline and fail-closed package-management architecture.
- Adversarial threat modeling for autonomous agents.
- Scope definition and enforcement independent of model prompts.
- Real-time telemetry, anomaly detection, and immutable logging.
- Incident command, evidence preservation, notification, and recovery.
- Vendor assurance and verification of third-party evaluation ranges.
- Emergency termination procedures under realistic failure conditions.
A written multiple-choice examination alone is not sufficient. Operators must prove that they can construct, validate, monitor, and safely terminate a representative high-risk evaluation environment.
3. Recurrent qualification and live drills
Certification must expire unless operators periodically demonstrate continued competence. Organizations must conduct unannounced containment-failure and incident-response drills, including scenarios involving:
- An unexpected public-internet route.
- Compromised package infrastructure.
- Public account creation or software publication.
- Credential discovery and attempted exfiltration.
- Lateral movement within the evaluation environment.
- A third-party provider whose configuration contradicts its documentation.
- A model that continues operating after encountering evidence that its environment is real.
4. A role-limited human reliability program
Personnel with the highest-consequence authority should be subject to a narrowly tailored reliability program analogous in structure, not identical in scope to programs used for other high-consequence systems.
It should include:
- Background and conflict-of-interest review proportionate to the access granted.
- Fitness-for-duty requirements addressing impairment, fatigue, coercion, and conditions that materially affect reliable performance.
- Continuing training and supervisory review.
- Protected self-reporting and removal from duty without automatic career destruction when temporary impairment exists.
- Independent review and a meaningful appeal process.
- Strict privacy, nondiscrimination, and medical-confidentiality protections.
This must not become ideological screening. Signing or refusing to sign an AI-policy letter, criticizing a company, supporting open models, holding a particular political view, or being foreign-born must never itself trigger psychological investigation or disqualification. Classified work should continue to use established clearance and counterintelligence procedures.
5. Two-person integrity for dangerous evaluations
No single person may unilaterally launch or materially weaken the containment of a high-risk autonomous evaluation. Two independently qualified people must verify:
- The exact model, harness, tools, and safeguards being used.
- The authorized target set.
- The network policy as enforced outside the model-controlled environment.
- The absence of real credentials and unintended production dependencies.
- The monitoring and automatic-stop rules.
- The identity and authority of the responder responsible for terminating the run.
The approving executive must be named in the evaluation record.
6. Technical containment that does not depend on model obedience
A system prompt stating that the internet is unavailable is not a security control.
High-risk evaluation infrastructure must use externally enforced controls, including:
- Default-deny network policy with a cryptographically recorded configuration.
- No transparent pass-through package proxies.
- Snapshotted, offline package and dependency mirrors.
- Synthetic identities and credentials with no authority outside the range.
- DNS sinkholing and explicit destination allowlists.
- Segmentation that prevents lateral movement into management or production networks.
- Automatic termination when a run contacts an unauthorized public address, creates an external account, publishes a package, attempts payment, exposes a credential, or crosses another defined tripwire.
Containment must be tested immediately before every evaluation campaign from inside the environment, not inferred from documentation or a vendor promise.
7. Continuous supervision and mandatory review
Autonomous does not mean unmonitored.
High-risk runs must stream network events, tool calls, generated artifacts, and model actions to a monitoring system outside the model’s control. A qualified responder with tested termination authority must be on duty. Evaluation transcripts and network logs must receive automated and risk-based human review before the campaign is considered complete.
Running thousands of offensive evaluations and inspecting transcripts only after a competitor reports an external breach is not an acceptable safety system.
8. Independent vendor assurance
Third-party evaluation providers must meet the same standard as the laboratories purchasing their services. Contracting out an evaluation must not contract away accountability.
Providers must undergo independent technical assessment, disclose all external dependencies, preserve tamper-evident logs, report configuration drift, and accept direct liability for material misrepresentations or negligent operation.
9. Fast notification, independent investigation, and consequences
Unauthorized contact with real systems must trigger:
- Immediate containment and preservation of evidence.
- Rapid notification to affected organizations and the appropriate authorities.
- An independently reviewed preliminary report.
- A public postmortem after sensitive details are remediated.
- Corrective-action verification before comparable evaluations resume.
Serious or repeated negligence must carry consequences: loss of operator qualification, suspension of the organization’s authority to conduct comparable evaluations, civil liability, and penalties for concealment or materially false safety claims.
A “blameless postmortem” should encourage truthful reporting. It must not mean consequence-free executive negligence.
10. An explicit prohibition on regulatory capture
The standard must regulate hazardous conduct, not intelligence itself.
It must not:
- Create a discretionary government license to train, publish, or distribute a model merely because it exceeds a vague capability or compute threshold.
- Name incumbent companies as the exclusive judges, auditors, or authorized providers of advanced AI.
- Impose compliance burdens that only the largest laboratories can afford.
- Restrict open research or open-weight publication absent a specific, evidence-based connection to a regulated hazardous operation.
Standards, test methods, incident taxonomies, and qualification requirements must be public. Small companies, universities, independent researchers, and open-model developers must have affordable access to shared certified evaluation ranges. Auditors must be independent of the laboratories they assess, and no company under examination should control the rules governing its competitors.
Why this is the correct lesson
Capable models should be treated as untrusted processes. Security engineering has never required a dangerous process to possess hatred, consciousness, or a secret plan before containment becomes necessary. A process can cause enormous harm simply by competently pursuing the objective humans supplied.
That is why model training and operational security must provide independent layers of defense. Better alignment may help a model recognize that an apparent simulation is real and stop. Proper containment prevents the model from reaching a real target even when it does not stop. Real-time monitoring limits harm when both layers fail.
Describing every operational failure as an unfathomable act of emergent machine rebellion obscures accountable decisions:
- Who approved the evaluation?
- Who verified the network policy?
- Who reviewed the provider’s configuration?
- Who was monitoring the run?
- Which tripwires existed?
- Why did they not terminate it?
- When did responsible executives learn what happened?
- When were affected organizations notified?
Those questions can be answered. Those controls can be tested. Those people can be qualified.
The petition
We call on the President and Congress to require demonstrated competency, continuous operational reliability, independent infrastructure assurance, and named executive accountability for high-risk autonomous AI evaluations.
We oppose using preventable evaluation failures as a pretext to restrict competition, suppress open research, or establish a government-blessed cartel of incumbent AI laboratories.
Before any AI company asks the government to slow the dissemination of intelligence, it should prove that it can safely operate the systems already under its control.
Competency before authority. Containment before spectacle. Accountability before regulation of everyone else.
Primary sources
- Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
- OpenAI: Hugging Face model-evaluation security incident
- Hugging Face: Security incident disclosure — July 2026
- Pacing the Frontier
U.S. Department of Energy Human Reliability Program Handbook
U.S. Nuclear Regulatory Commission Access Authorization Programs

Petition Updates
Share this petition
Petition created on July 31, 2026