Every frontier model you have used was built, in large part, on data that already existed. The text had been written, the photographs taken, the code committed — for human reasons, long before anyone thought to train on any of it. Common Crawl did not create a single page of the web; it collected what was already there. The great AI argument of the past four years has really been about payment for that inheritance: who owes what for the found data of the internet.
The next era of AI cannot be built that way — because the data it needs mostly does not exist.
You cannot scrape what was never recorded
Consider what an autonomous system actually has to learn. A humanoid robot needs to know what it feels like — in forces, torques, and joint angles — to walk on uneven ground, to open a stubborn jar, to carry an awkward load up stairs. A surgical robot needs thousands of hours of annotated procedures, with instrument telemetry aligned to video, including the rare complications. A warehouse robot needs teleoperation logs from human operators doing the job well and doing it badly. A drone needs sensor streams from weather it will only encounter once a year. An industrial inspection system needs vibration, temperature, and pressure signatures of machines in the act of failing — data every plant owner has and no crawler will ever see.
Roboticists have a name for the underlying problem: the things easiest for humans are hardest for machines. But note what the examples have in common. Almost none of this data is lying around on the public internet. The web was a byproduct of human life; robot training data has to be a product — deliberately recorded, annotated, and maintained, at real cost. The research community already knows this. The most useful datasets in robotics and embodied AI exist because someone paid camera wearers, teleoperators, and labs to produce them. Nobody has ever needed to pay the web to exist. Everyone will need to pay for this.
Markets are how you manufacture supply
If the data must be deliberately produced, then someone has to fund its production — and the only mechanism that funds production reliably, at scale, across millions of independent producers, is a market. That is the reframe I want to offer the AI industry: compensation is not a tax on model training. It is the supply chain of physical AI.
A working data market needs four pieces of plumbing. Registration, so it is knowable who owns a given asset and on what terms. Fingerprinting, so the asset is identifiable wherever it travels. Attribution signals, so usage can be connected to sources at whatever confidence the technology supports. And prompt settlement, so the money actually reaches the producer — not eighteen months later through a lawsuit, but as a routine consequence of use. I have seen this machinery work before: two decades ago we licensed music and game content from the major publishers and settled the royalties on phone bills across an entire continent. The technology was primitive. The market worked anyway, because every party could see that usage became payment.
What the AI industry actually gains
Supply that would not otherwise exist. When registration and settlement are routine, data production becomes a business. Hospitals license annotated procedure archives. Plant operators license failure signatures. Motion-capture studios, sensor cooperatives, and yes — people wearing cameras while walking on uneven ground — produce exactly the data autonomy needs, because producing it pays. The alternative is not free data. The alternative is no data.
Provenance as a safety property. A language model trained on an unverifiable corpus embarrasses its maker. A surgical robot trained on an unverifiable corpus is a courtroom exhibit. As AI moves into machines that touch people, chain-of-title stops being paperwork and becomes an engineering requirement: you cannot certify what you cannot trace. Registered, fingerprinted data with a clean audit trail is also the strongest practical defense against data poisoning — you know what went in, and who attested to it.
Legal certainty that unlocks investment. Today, training-data exposure is an unquantifiable contingent liability, which is the kind boards hate most. Licensing converts it into a predictable cost of goods — and documented licensing produces, as a byproduct, precisely the evidence that the EU’s transparency obligations and India’s consent-first regime now demand. Companies do not build decade-long roadmaps on top of open legal questions. They build them on top of invoices.
Coverage of the tail, where autonomy fails. Scraped data over-represents the common case; autonomous systems fail on the rare one — the unusual gait, the odd lighting, the one-in-a-thousand complication. Nobody uploads their failures to the open internet for free. Pay for edge cases, and the edge cases appear. This is the quiet, unglamorous reason a compensated data economy produces better robots than a scraped one ever could.
There is a fifth benefit that accrues to everyone: the end of the adversarial web. Right now the equilibrium is blocking by default, paywalls, poisoning tools, and litigation — an arms race in which AI companies and creators burn resources making each other’s lives harder. The moment the meter runs and payment flows, the incentive flips. Rights holders stop fighting to keep their work out of AI and start competing to get it in. That is what a functioning market looks like from the inside.
One more thing, stated carefully. Some of the most valuable data of the coming decade will also be the most personal — conversational recordings, interpersonal interaction, the texture of ordinary human life. The rule there is simple: the more personal the data, the more indispensable auditable consent and compensation infrastructure becomes. Registration is what makes consent provable. Payment is what makes it durable.
The honest objection
All of this costs something. Registration is friction, licensing is overhead, and attribution at scale remains an unsolved research problem in its strongest form. True — and none of it is disqualifying. Collective licensing has cleared music royalties for a century using pooled allocation, not per-play forensic certainty. Settlement can consume attribution signals at whatever confidence exists today — deterministic where retrieval is logged, attested where usage is reported under contract, allocated where only pools are practical — and improve as the science does. Markets do not wait for perfect measurement. They start with an accepted standard and refine it. Every commodity market in history worked this way.
The registered economy
So: would the AI industry benefit if all content were fingerprinted, attribution worked at scale, and every knowledge producer — individual or organization — were promptly paid? The generative era asked who owes what for the past. The physical era asks a better question: who will produce the future, and why would they?
Registration begets attribution, attribution begets payment, payment begets supply, and supply begets capability. The companies that win the next decade of AI will not be the ones with the biggest crawlers. They will be the ones with the best-paid suppliers.
Get paid when AI uses your knowledge. That is not a slogan against the AI industry. It is the business model of its next decade.