The Confessions of the Young: Childhood Self-Disclosure on the Early Web and the Limits of Consent in AI Training Corpora

DOI: 10.5281/zenodo.20471414

Abstract

Language models train on web text. Much of this text originated on early personal homepages, a publishing culture from the mid-1990s to late 2000s built on self-disclosure. Many authors were minors. Their disclosures—names, ages, schools, cities, photographs, contact handles—match categories protected by children's privacy law. Ingesting this material into training corpora violates ethics and law. Community disclosure grants no license for commercial use. Minors lacked capacity to consent; third parties cannot consent retroactively. The model dissolves inputs into parameters, rendering harm irreversible and defeating erasure rights. The early homepage concentrates the problem: candid authorship by minors without consent.

Introduction

The debate over training data focuses on copyright holders and adults. A third constituency receives less attention: children in the first decade of mass web publishing who filled personal homepages with accounts of their lives. Those disclosures sit inside language models, monetized at scale, without consent.

The early homepage isolates the mechanics of data ingestion. It was a site of self-presentation, not commerce. Minors authored it. The disclosures map directly onto protected privacy categories, yielding material that trains a model while making legal protection necessary.

The Disclosure Culture of the Early Web

Before consumer tools, web publishing required technical competence. Services like GeoCities removed these barriers, allowing millions to build their first websites.1

The underlying design of early services organized members through metaphors of place. In 1997, GeoCities arranged users into themed "neighborhoods"; "homesteaders" gathered in districts organized by interest. A member entered the online world through a community rather than a profile built from personal facts.2 On later platforms, personal disclosure became the price of entry; on the early web, disclosure remained an elective act of furnishing a built space.

The artist and theorist Olia Lialina locates the early web's significance in the possibility of building a presence outside centralized services—a place one could inhabit and develop on one's own terms.3 The visual conventions—animated graphics, under-construction icons—formed the surface of a publishing ethic oriented toward self-expression, a culture of unrefined candor.4

The young produced that candor. The early web was a native medium for adolescents, a place to record the particulars of their lives. The medium's governing convention—tell the truth about oneself—produced a documentary record of sincerity and specificity.

The Correspondence with Protected Categories

That documentary record of sincerity is precisely what modern privacy frameworks shield.

A typical homepage carried the author's name, age, school, city, sports affiliations, photographs, names of family and friends, and screen names. Under the United States Children's Online Privacy Protection Act (COPPA), "personal information" includes names, contact information, and persistent identifiers. The 2025 amendments expanded the definition to encompass biometric and government identifiers.5 Information a thirteen-year-old placed on a homepage in 1999 falls within the category COPPA shields from commercial hands absent verifiable parental consent.6

This correspondence drives both value and harm. A corporate page discloses nothing personal, training a model in register and boilerplate. A child's homepage discloses a life in a first-person voice. The features making personal pages rich as training material—natural language and specific details about a life—make their ingestion a privacy harm. The model acquires an adolescent voice because an adolescent wrote in it, by name, about her circumstances.

The law has begun to map this harm. Recent regulatory shifts—from the 2025 COPPA amendments7 to New York's Child Data Protection Act8—extend protections to older adolescents and strictly define processing to include collection, retention, and monetization. Operators require a lawful basis to process this data; model development is not one of them.9

Scraping training material constitutes processing. The operation occurs in the present, regardless of when the data appeared online. The scraper is the company's instrument; the operations are the company's.

A 1997 homepage states the author is thirteen and lives in New York. A company ingesting this page processes the data of a covered user in the present. No lawful basis applies. The child gave no consent and lacked the capacity to do so. The processing is unlawful. The page's current online availability changes nothing; the statute regulates data operations, not source availability.

"Publicly Available" Offers No License

Developers defend ingestion by claiming the material sat on the open web. This defense fails under existing law, especially concerning minors.

Public availability does not extinguish data-protection obligations. Scraping legality turns on what is collected and how it is used. Collecting personal data without a lawful basis violates privacy frameworks, even if publicly posted.10 Under the European Union's General Data Protection Regulation (GDPR), processing personal data requires a lawful basis regardless of public status; European supervisory authorities rule commercial interest cannot justify scraping.11 In 2025, the French data-protection authority issued guidance advising developers to exclude websites used by minors or containing sensitive data, and to honor opt-out mechanisms like robots.txt.12 A child's homepage is the paradigm instance.

United States law converges on this conclusion. Litigation over training-data provenance alleges developers built systems on private information taken from internet users, including children, without consent.13 Legal commentary views the open web not as a public domain, but as an aggregation of personal data subject to privacy, consumer-protection, and tort regimes. Unknown training-data provenance compounds legal exposure.14

The defense mischaracterizes the author's act. Posting to a homepage in 1999 was publication to a community of readers. It was not a license grant to an industry that did not exist for a use no one conceived. Consent in data-protection law is purpose-bound, attaching to a disclosed use at the time given. No AI training existed in 1999 to permit consent. Treating the lack of objection to an unforeseeable use as authorization constructs a false consent.

For a minor, the construction collapses. COPPA assumes a child cannot furnish valid consent to commercial collection; only a verified parent holds that authority.15 Consent deemed impossible at disclosure cannot be supplied retroactively by a third party for an unintended use. The author's minority defeats the consent on which the defense relies.

Irreversible Harm

When consent fails, the legal remedy is removal. But the physical mechanics of language models nullify this right.

The right to erasure was developed for structured databases, where records could be located and deleted.16 The law extended this right to children, requiring deletion of personal information collected through online services.17 A trained model does not retain inputs as addressable records. Training dissolves information into parameters; no entry remains to strike.18 Honoring an erasure request requires retraining the system, a remedy developers reject for individual claims.19 Once a model absorbs personal data, it remains permanent.20

This permanence transforms ingestion into a standing condition. The model metabolizes the child's disclosure into a commercial asset producing continuous value. The legal right to deletion collides with a design making deletion infeasible. European enforcement bodies prioritize action against AI developers processing children's data under erasure provisions, but enforcement arrives post-absorption. The technology blocks the legal remedy.21

The irreversibility makes the act unremediable. A system acquiring data without right, then locking that acquisition into its design, commits no procedural error. The error becomes the system.

Provenance as Obligation

Faced with this irreversibility, the industry attempts to shield itself behind scale. Developers justify ingestion by claiming ignorance of provenance—arguing they cannot identify pages authored by children or audit a corpus scaled to the open web.22

The claim fails. Inability to establish provenance provides no defense against mishandling sensitive data. In a regime conditioning lawful processing on a known basis, claiming ignorance is an admission. A practice ingesting both a child's disclosure and public text indiscriminately has resolved that consent poses no constraint.

The remedy rests on treating provenance as a material constraint. A system that cannot interrogate the origins, authorship, and consent capacity of its data cannot claim a lawful basis. If "publicly available" defines location rather than license, then a system built without the mechanism for erasure violates the law by design.

Early web authors honored an implicit bargain: they told the truth. That bargain did not include harvesting that truth into a commercial system and dissolving it beyond recovery. Strangers appended that clause retroactively to disclosures made by children who lacked the legal capacity to agree. A child's account of her life is not a public resource. The account is a person on the record. Taking it permanently, without consent and without the capacity to return it, is not a technological inevitability. It is the permanent enclosure of a life.

Works Cited

  1. Ian Milligan, "Welcome to the Web: The Online Community of GeoCities During the Early Years of the World Wide Web," in The Web as History, ed. Niels Brügger and Ralph Schroeder (London: UCL Press, 2017), 137–158.↩︎
  2. Christine Rosen, "Virtual Friendship and the New Narcissism," The New Atlantis, no. 17 (Summer 2007): 15–31.↩︎
  3. Olia Lialina, "A Vernacular Web: The Indigenous and The Barbarians," 2005, http://art.teleportacia.org/observation/vernacular/.↩︎
  4. "Cameron's World," accessed May 30, 2026, https://www.cameronsworld.net/.↩︎
  5. Federal Trade Commission, "Children's Online Privacy Protection Rule," Federal Register 90 (April 22, 2025), https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule.↩︎
  6. Federal Trade Commission, "Children's Online Privacy Protection Rule."↩︎
  7. Federal Trade Commission, "FTC Strengthens Children's Online Privacy Protection Rule," Press Release, December 20, 2023.↩︎
  8. New York Child Data Protection Act, N.Y. Gen. Bus. Law §§ 899-ee et seq. (effective June 20, 2025); New York State Office of the Attorney General, "New York Child Data Protection Act Implementation Guidance," May 19, 2025. Section 899-ee defines "covered user," "minor," "operator," "personal data," and "processing."↩︎
  9. N.Y. Gen. Bus. Law § 899-ff. For a covered user under thirteen, processing is governed by the COPPA standard; for a covered user aged thirteen to seventeen, processing is prohibited unless strictly necessary for an enumerated permissible purpose or the user has given informed consent.↩︎
  10. European Data Protection Board (EDPB), "Guidelines 8/2020 on the Targeting of Social Media Users," adopted April 13, 2021, 15.↩︎
  11. European Data Protection Board, "Guidelines 8/2020 on the Targeting of Social Media Users."↩︎
  12. Commission Nationale de l'Informatique et des Libertés (CNIL), "Artificial Intelligence: The CNIL Publishes Its Recommendations on the Development of AI Systems," 2025.↩︎
  13. Shayne Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI," Advances in Neural Information Processing Systems 36 (2023).↩︎
  14. Woodrow Hartzog, Privacy's Blueprint: The Battle to Control the Design of New Technologies (Cambridge, MA: Harvard University Press, 2018).↩︎
  15. Federal Trade Commission, "Children's Online Privacy Protection Rule."↩︎
  16. Michael Veale, Reuben Binns, and Jef Ausloos, "When Data Protection by Design and Data Subjects Rights Clash," International Data Privacy Law 8, no. 2 (2018): 105–123.↩︎
  17. Regulation (EU) 2016/679 (General Data Protection Regulation), Article 17(1).↩︎
  18. Nicholas Carlini et al., "Extracting Training Data from Large Language Models," 30th USENIX Security Symposium (2021): 2633-2650.↩︎
  19. Lucas Bourtoule et al., "Machine Unlearning," IEEE Symposium on Security and Privacy (2021): 141–159.↩︎
  20. Carlini et al., "Extracting Training Data from Large Language Models."↩︎
  21. European Data Protection Board (EDPB), "Report on the Findings of the 2025 Coordinated Enforcement Framework Action on the Right to Erasure," adopted February 2026.↩︎
  22. Longpre et al., "The Data Provenance Initiative."↩︎