Author: Nita Farahany Source: https://nitafarahany.substack.com/p/red-teaming-governance-from-practice Published: October 7, 2025
The thirteenth class in Nita Farahany's 27-class AI Law & Policy course (Class XIII of 27) examines the legal frameworks governing AI red-teaming. The class argues that the act of discovering AI vulnerabilities can itself be attacked legally as hacking, DMCA circumvention, or breach of contract, and surveys proposed safe-harbor, mandate, shield, and liability responses.
Summary of argument
The class is anchored on OpenAI's response to NYT v. OpenAI, in which OpenAI accused the New York Times of "hacking" its system by demonstrating that ChatGPT could reproduce copyrighted articles verbatim. Farahany uses this to argue that the very act of discovering AI vulnerabilities can be legally attacked. She contrasts the positions of the parties: the NYT can defend itself, whereas an independent graduate-student researcher cannot.
Key claims
Farahany identifies three legal traps for AI safety researchers:
- The Computer Fraud and Abuse Act (CFAA, 1986), which covers accessing a computer "without authorization" or "exceeding authorized access," carrying penalties of up to $500K and 5 years in prison. Van Buren v. US (2021) narrowed the statute but left ambiguity for AI testing that violates terms of service.
- DMCA Section 1201, which covers circumventing a "technological measure that effectively controls access to a work." Whether AI safety filters qualify as "protection measures" is unclear, and no court has ruled on the question for LLMs.
- Contract law and terms of service: every platform carries terms of service, and violations expose a researcher to civil liability and account bans.
The class frames a researcher's dilemma through an example: a researcher who finds that ChatGPT reveals credit-card patterns when prompted in base64 faces CFAA criminal-prosecution risk, DMCA risk, and breach-of-contract civil liability at once, with the result that few researchers can publish.
Against these traps, the class describes the Longpre et al. safe-harbor proposal, a two-part framework adapted from cybersecurity vulnerability disclosure (Coordinated Vulnerability Disclosure). Its legal safe harbor would have companies commit not to sue good-faith researchers. Its technical safe harbor would bar automated account suspension, provide clear appeals, and allow pre-registration of researchers. Farahany notes a two-tier risk in the proposal: only well-funded organizations can afford to participate, leaving individual graduate students exposed.
The class then contrasts two state approaches. California SB 53 takes a mandate approach: frontier developers must test for "catastrophic risks," defined as mass casualties (>$500M damage), critical-infrastructure cyberattacks, and weapon-design assistance. Farahany argues this misses the harms already observable — individual harms, manipulation, discrimination, and election interference — and that the law requires developers to test and disclose but not remediate, so that a developer can find a catastrophic risk, report it, and deploy anyway. Penalties cap at $1M per violation, which she calculates at 0.0002% of OpenAI's $500B valuation. The law includes whistleblower protections.
Texas HB 149 takes a shield approach: discovering violations through "adversarial testing or red-team testing" provides a defense against civil penalties (Section 552.105(e)). It also establishes a Regulatory Sandbox Program (Chapter 553) allowing AI to be tested with certain laws waived for up to 36 months. Farahany summarizes the California–Texas split as California saying "you must test and tell us what you find (but you can still deploy)" and Texas saying "if you choose to test, we won't punish you for what you find."
At the federal level, the class presents the Trump AI Action Plan and the Durbin–Hawley LEAD Act as competing visions. The Trump plan limits red-teaming to national security, covering CBRNE and cyber. The LEAD Act takes a products-liability framing under which developers are liable for "reasonable care" failures, with "design" defined to include training, testing, auditing, and fine-tuning. Farahany notes that the LEAD Act does not prescribe testing methods but makes developers liable if they do not test adequately and harm results.
The class closes on an unresolved question that, in Farahany's account, none of the frameworks answers: even with safe harbors, mandates, shields, or liability incentives, what is to be done when red-teaming finds unfixable problems.
Relationships
- part-of: Nita Farahany intro course series (Class XIII of 27)
- related: Jailbreaking and Red Teaming, Cfaa (planned), California SB 53, AI LEAD Act (S. 2937)
- previous: Inside My AI Law & Policy Class 12: Red-Teaming AI (Farahany, October 2025) next: Inside My AI Law & Policy Class 14: When Anyone Can Fake Anything (Farahany, October 2025)