AI Policy Wiki
Dashboard

Inside My AI Law & Policy Class 18: Data Privacy in an AI World (Farahany, November 2025)

medium confidence · updated 2026-06-06

Farahany's AI data privacy class. Opens with her own Google search history fed to Claude, returning an ad-targeting profile at 89% confidence. Sets out three privacy problems AI creates that traditional privacy law was not designed to handle — inference explosion, agent cascade, synthetic data generation — and tests GDPR, CCPA, and the FTC Act against each.

Author: Nita Farahany Source: https://nitafarahany.substack.com/p/data-privacy-in-an-ai-world-inside Published: November 2, 2025

This is the data-privacy installment of Nita Farahany's AI Law & Policy course (Class 18 of 27), published November 2, 2025. The class argues that AI creates three privacy problems traditional privacy law was not designed to handle — inference explosion, agent cascade, and synthetic data generation — and works each against existing regimes (GDPR, the CCPA, and Section 5 of the FTC Act). It opens with a demonstration: Farahany fed her own Google search history, printed as a PDF, to Claude, which returned an ad-targeting profile at 89% confidence, sorting her into segments labeled "Stealth Wealth Academics," "Privacy-Conscious High Earners," and "Mountain Property Owners."

Framing

The class frames AI's appetite for personal data using Ryan Calo's Senate testimony, which compared AI to the film Soylent Green: "it's made out of people." Data about people is what fuels AI systems, and on this account the incentive to scrape and collect indefinitely is the business model rather than a side effect.

The three privacy problems

Inference explosion is the first problem: a person shares roughly three things and an AI infers about thirty attributes the person never consented to share. A gym selfie, in the class's example, can yield the exact gym location, an income estimate (from the gym brand), a health-consciousness score usable by insurers, social patterns, work-schedule flexibility, and relationship status (from a wedding ring). The class cites Target's pregnancy prediction, in which the retailer identifies pregnancy from a pattern of roughly 25 products and can estimate a due date within two weeks, sometimes identifying the pregnancy before the family knows. Stanford HAI describes this as "intimate inference from ambient data."

The agent cascade is the second problem: a single instruction such as "book me a therapist" sets off a chain of data flows. The user's AI accesses calendar, location, insurance, and health-search history, then contacts therapy-platform AIs (sharing an insurance ID, which supports an income inference); those platform AIs in turn share with insurance AIs (for risk scoring), scheduling AIs, billing AIs, and analytics AIs; all of the systems update their models; and a "mental health risk score" becomes permanent. The class poses the question of at what point in this cascade the user loses meaningful consent.

The synthetic data generator is the third problem: after a three-minute work conversation with an LLM, a user asks for a professional bio, and the AI invents skills the user never mentioned, projects never worked on, and accomplishments not achieved. If that bio is scraped into training data, the false "facts" become permanent. The class connects this to GDPR Article 16's right to rectification of inaccurate data, asking whether a person can correct something that was never "collected" but rather "generated."

Testing existing law

The class works three scenarios against existing regimes.

In the first scenario, GDPR is tested against inferences: Microsoft Copilot infers that a user is job-hunting from document edits. Under GDPR Article 15 (access), the company says the inference is proprietary and not "your data." Under Article 22 (automated decision-making), a human manager reviewed the matter, so the provision does not apply. Under Article 17 (erasure), deleting the source data does not change the model's learned patterns.

In the second scenario, state law is tested against the agent cascade: a health AI shared therapy patterns with insurance AIs, and premiums rose 30%. Under the CCPA, a user can opt out of a "sale," but the company says the transfer is not a sale but service provision. A risk-benefit test is an internal process rather than an individual right, and because the cascade crosses state lines instantly, rights do not follow it.

In the third scenario, the FTC is tested against synthetic data: an AI generates a false "collaboration score" of 67/100. Under Section 5 of the FTC Act, the company disclosed that the figure is an AI-generated analytical metric rather than a factual measurement.

Proposed responses

The class presents three proposed solutions from Stanford HAI. The first is privacy by default — flipping from opt-out to opt-in, with EU cookie consent offered as an illustration of friction without control. The second is supply chain transparency — mandatory documentation of training-data sources, paired with algorithmic disgorgement, a remedy the FTC has used. The third is data intermediaries — trusts or cooperatives that act as a person's agent, enabling collective bargaining at scale and imposing fiduciary duties on platforms.

It also sets out three statutory proposals from Calo: mandate true data minimization with hard limits; treat inferences as personal data carrying the same rights as voluntary disclosures; and prohibit data misuse beyond mere deception, such as using a person's loneliness to sell to them.

Relationships