Who's Getting Paid When AI Learns From Your Posts, Photos, and Private Messages
Somewhere between you typing a Reddit comment and a frontier AI model learning how humans argue about movies, money changed hands. You just didn't see it.
The AI training data economy is enormous, largely invisible, and growing fast. Analysts at Cognilytica estimated the data labeling and collection market would surpass $17 billion by 2030. That number has almost certainly been revised upward since the generative AI explosion of 2022-2023. The people funding that market? Mostly the big AI labs. The people supplying the raw material? Mostly you.
Let's pull back the curtain.
The Middlemen You've Never Heard Of
Most people assume AI companies scrape the public internet — and they do. But the more interesting story lives in the secondary market, where companies that collect data act as wholesalers to companies that train models.
Data brokers like Acxiom, LiveRamp, and LexisNexis have been in the business of aggregating personal information for decades. Originally built to serve advertisers and financial institutions, many of these firms have quietly repositioned themselves as AI-ready data suppliers. Their pitch to AI labs is straightforward: we have clean, structured, consented (sometimes) datasets that are far easier to work with than raw web scrapes.
Then there are the specialized players — companies like Scale AI, Appen, and Surge AI — that don't just collect data but annotate it. They hire human workers, often through crowdsourced platforms, to label images, transcribe audio, rank AI outputs, and flag problematic content. The humans doing this work are often in the Global South, paid pennies per task. But the companies selling that labor and the resulting datasets to OpenAI, Google, and Meta? They're doing very well.
The Apps You Trust Are in the Supply Chain Too
Here's where it gets closer to home. Several mainstream consumer platforms have updated their terms of service in ways that explicitly (or implicitly) allow user data to be used for AI training.
Adobe made headlines in 2023 when users noticed its updated terms seemed to grant the company broad rights to access user content for AI development. Adobe clarified — eventually — but the backlash exposed how buried these permissions usually are.
X (formerly Twitter) updated its privacy policy to include language permitting the use of public posts to train its own Grok AI model. You opted into that the moment you kept using the platform.
Google revised its terms in mid-2023 to state that it may use publicly available information to help train its AI models. Gmail and Docs data reportedly remain separate — but "reportedly" is doing a lot of work in that sentence.
Zoom caused a minor panic when its updated terms appeared to allow training AI on video and audio from calls. After significant user backlash, Zoom walked back the broadest interpretations, but the episode illustrated how fast these changes can slip through.
LinkedIn, owned by Microsoft (which is deeply invested in OpenAI), has a data-sharing arrangement that has raised eyebrows among privacy researchers. Microsoft's AI ambitions and LinkedIn's trove of professional behavioral data are not unrelated.
The pattern here isn't coincidence. Platforms that built their value on user-generated content are increasingly treating that content as a secondary revenue stream — one that doesn't require them to pay the people who created it.
What the AI Labs Actually Want
Not all data is equal in the training market. AI labs are particularly hungry for a few specific categories:
- Long-form human writing — forums, reviews, emails, documents. The messier and more conversational, the better for teaching models how real people communicate.
- Instructional content — how-to guides, Q&A threads, tutorials. Stack Overflow and Reddit have both had very public fights with AI companies over bulk data access.
- Multilingual text — as labs push into global markets, non-English content commands a premium.
- Code — GitHub data has been central to models like Copilot, which is why the class-action lawsuits from developers have been particularly pointed.
- Images with metadata — not just the pixels, but the context around them.
Reddit, for its part, struck a reported $60 million annual deal with Google for API access to its data shortly before its IPO. The company framed it as a business necessity. Critics framed it as monetizing 18 years of free human labor without compensating the people who produced it.
The Legal Picture Is a Mess
In the US, there is no comprehensive federal data privacy law. That's not a bug in the system — it's a feature that the data industry has spent decades lobbying to preserve.
What exists instead is a patchwork: California's CCPA gives residents some opt-out rights, Illinois's BIPA covers biometric data, and sector-specific laws like HIPAA protect medical information. But for the vast majority of Americans, there is no legal right to know whether your content is being used to train AI, no right to compensation if it is, and limited ability to demand deletion.
The EU's GDPR provides meaningfully stronger protections, which is part of why the most aggressive data-for-AI arrangements tend to originate in US markets first.
Class action litigation is starting to fill some of the gaps — lawsuits against OpenAI, Stability AI, and others are working through the courts — but legal resolution is years away at minimum.
What You Can Actually Do Right Now
You're probably not going to opt out of the internet. But there are practical moves worth making:
Read the terms when they change. Platforms are required to notify you. Most people ignore those emails. Start skimming them, especially for language about "improving services," "machine learning," or "third-party partners."
Use opt-out tools where they exist. Some platforms offer explicit AI training opt-outs buried in privacy settings. LinkedIn, for example, added one after user pressure. It's not always easy to find, but it's there.
Limit what you share on public-facing platforms. Private accounts, locked-down settings, and limiting the platforms you actively use all reduce your surface area.
Look into data broker opt-outs. Sites like DeleteMe and Privacy Bee (paid services) or the free opt-out portals maintained by companies like Acxiom can reduce your profile in the commercial data ecosystem.
Be skeptical of "free" productivity tools. If a tool is free and it processes documents, emails, or voice, the data economics deserve scrutiny. Check whether the company has AI products and whether your inputs might be feeding them.
The Bigger Picture
The AI training data gold rush isn't slowing down. If anything, the competition for high-quality human-generated content is intensifying as labs hit the ceiling on easily available web data. Synthetic data is being explored as an alternative, but for now, the most valuable training material is still the stuff real people produced without knowing it had a price tag.
The uncomfortable truth is that the AI systems people are marveling at were built, in large part, on a foundation of uncompensated human creativity and communication. That's not a conspiracy — it's just how the economics shook out in the absence of regulation.
Knowing that doesn't give you back control. But it does mean you can stop being quite so surprised when the next terms-of-service update lands in your inbox.
Read it this time.