The Control standard v1.0 is the first published standard of Guidelight AI Standards, an independent nonprofit AI-safety standards organization founded by former OpenAI safety staff Page Hedley and Steven Adler. First released in May 2026, the standard defines minimum organizational and technical measures for preventing loss of control of AI systems during internal deployment. Guidelight published an explanation of the standard on July 15, 2026 and said it is conducting its first assessment of leading AI companies against it (Source: clear-eyed.ai).
Scope
Guidelight defines control as the technical and operational measures that constrain what an AI system can do, regardless of whether the system is aligned. On the organization's account, these measures serve two functions: they reduce catastrophic risk from a misaligned AI, and they can surface evidence of an AI's misalignment (Source: guidelight.ai). The standard addresses internal deployment — the use of AI systems inside the developing organization, before or apart from public release — the setting the AI control research agenda identifies as carrying loss-of-control risk that pre-release evaluations do not cover.
Structure
The standard sets six minimum principles, each elaborated with requirements in the full standard text (Guidelight AI Standards — Control (v1.0, May 2026); Source: guidelight.ai):
- Visibility — "Be able to see what your AI is doing during internal deployment" (tamper-evident activity logs).
- Scanning — "Scan for signs of concerning behavior."
- Stress-testing — "Stress-test the sufficiency of your scanning."
- Barriers — "Stop the AI from taking harmful actions even if it tried" (proactive barriers).
- Independent verification — "Have independent third parties verify the adequacy of your control regime."
- Breach preparedness — "Prepare for a possible breach of control" (breach-response planning).
Several requirements are stated quantitatively or as bright lines, which is what makes the standard third-party assessable. Under principle 1, "99.9% of inference tokens" must occur in a monitorable setting — logged and scannable, though the standard clarifies this means traffic is "able to be scanned," not that every request is. Risk-relevant models are defined by an existing statutory line, "the threshold for frontier models articulated by California SB 53 (10^26 FLOPs of training compute)," and internal deployment is defined broadly enough to include dogfooding, automated AI research, training-data generation, model critique, and agentic research pipelines (Guidelight AI Standards — Control (v1.0, May 2026)).
Principle 2 enumerates ten categories of concerning behaviour to scan for across both chain-of-thought and actions: deception and dishonesty; sabotage of safety-relevant work; reward hacking and specification gaming; sandbagging and capability concealment; resource and capability acquisition; self-preservation and shutdown avoidance; manipulation and influence on humans; misuse facilitation; inappropriate autonomous action; and unreliable monitorability. It also requires published, tiered response-time targets for human review and tracking of median, 90th- and 99th-percentile attainment against them (Guidelight AI Standards — Control (v1.0, May 2026)).
Principle 4 defines three tiers of action boundary — absolute (never crossable without human signoff), critical (gated by a monitor before execution), and important (trailed by a monitor able to intervene retrospectively) — and requires that changes to the monitor system itself and to control-relevant security configurations be treated as critical, closing the path by which a system could disable its own oversight. It also requires aggregate-flag halting, a rejection threshold that freezes further boundary-crossing actions until a human clears them, addressing the retry vulnerability the standard notes in semi-synchronous monitoring (Guidelight AI Standards — Control (v1.0, May 2026)).
Principle 3 requires recurring automated sting operations and monitor stress-tests, and — distinctively — testing whether models "can distinguish sting operations from real traffic," with less credence given to stings the model knows to be tests, making evaluation awareness an input to how much weight a control result carries. Principle 5 places the redaction judgment for published third-party adequacy reports with the assessor rather than the developer (Guidelight AI Standards — Control (v1.0, May 2026)).
Adoption and assessments
Guidelight stated on July 15, 2026 that it is conducting its first assessment of leading AI companies against the Control standard; results had not been published as of that date (Source: clear-eyed.ai).
Relation to other frameworks
The Control standard is one of three v1.0 standards on Guidelight's catalog, alongside Capability Testing (six principles on evaluating risk-relevant abilities, including eliciting maximum achievable performance and insulating evaluation from business pressures) and Transparency (two principles on exposing risk assessments to public scrutiny and timely incident reporting) (Source: guidelight.ai).
Where frontier-lab capability-threshold frameworks — Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework (AI Safety Cases and Frameworks) — are self-administered by the developer, the Guidelight standards are written to be assessable by a third party, and the standard itself requires independent external verification (principle 5). The standard codifies, as assessable requirements, the deployment-protocol approach developed in the AI control research agenda.
Relationships
- instance-of: AI Control (codification of the control agenda as an assessable standard)
- related: Guidelight AI Standards (publisher), Steven Adler (co-founder and Chief Scientist), AI Safety Cases and Frameworks (the lab-internal framework layer the standard complements), Alignment Auditing (adjacent third-party evaluation methodology)