For three thousand years, an idea came with a source. Now the answer arrives without it.
Whether chiseled into stone, printed on paper, or served through a Google link, an idea could not be separated from the object that carried it. The idea always arrived with its provenance intact. Nobody had to build it in because nobody could remove it. Generative AI is the first medium without that property. Now the answer arrives without a source, and nobody, not even the people who built the system, can say which works produced it.
Act I: Physical Knowledge
For most of recorded history, the carrier, location, and source were inseparable, so provenance did not need to be engineered as a separate layer. A tablet in Nineveh or a Bible chained to a lectern in Wittenberg could only be consulted by reaching it. The same object that carried the knowledge also carried evidence of its origin. When monks wrote books by hand, trust came from the institution, the monastery or the university that housed the manuscript.
When the printing press arrived, anyone with capital could set up a shop and print anything. People viewed early printed books with deep suspicion. Scholars feared books were full of typos, pirated, or demonic. The printed book was a terrifying new medium that threatened to flood Europe with untrustworthy text.
The industry saved itself and earned public trust by inventing the title page to establish provenance. Ratdolt, Maler, and Löslein gave the title, place, date, and printer's name a leaf of their own in Venice in 1476, and within two generations the practice was standard across the trade. Readers learned to trust a book because it bore the mark of a reputable printer known for auditing manuscripts for accuracy.
By the time Martin Luther translated the Bible into vernacular German and published it in 1534, most everyday Germans could not read or write. Luther's Bible carried distinct woodcut illustrations by Lucas Cranach's workshop and the mark of Hans Lufft, the Wittenberg printer. Whether you agreed with Luther or wanted to burn the book, a real human was visibly answerable for every word.
Over time, print gave way to magnetic tape, optical discs, and, with the internet, the hyperlink. Whatever the medium, the work always arrived with its claimant attached, and the provenance never moved. Trust was retained.
Act II: The Warning
In 1969 and 1970, Johns Hopkins and the Brookings Institution brought computer scientists, social scientists, industry executives, and federal policymakers together in Washington to examine what rapidly advancing computing and communications technologies would mean for government, institutions, and the public interest.
Emilio Daddario, chairman of the House Subcommittee on Science, Research, and Development and later the architect of the Office of Technology Assessment, moderated one session. He observed that knowledge was increasingly exchanged through magnetic tapes, remote consoles, information networks, and large data banks. The expansion was so rapid it was hard to document what was happening.
Daddario posited that powerful computerized information systems, unless precautionary steps were taken, might spawn further systems in Parkinsonian abandon, producing poor-quality scientific and technical information. He added that science flourishes only when it is untrammeled and open-ended. Information systems must not be institutionalized in ways that inhibit that freedom.
Herbert Simon then gave the talk everyone still quotes: information consumes the attention of whoever receives it, so a wealth of information creates a poverty of attention.
Simon held the title of Richard King Mellon University Professor of Computer Science and Psychology at Carnegie Mellon University and was a founder of artificial intelligence. He would take the Nobel Prize in economics eight years later. By the end of the session, the two men had converged, Simon saying his own answers were nearly identical to the congressman's.
At this time, a proposed National Data Center was already a major federal controversy. The plan to consolidate federal statistical records in one place had been through two years of privacy hearings in the House and Senate.
Act III: The Answer Without a Source
When Congress rejected the National Data Center in the late 1960s, fearing that a single institution would hold an uninspectable "dossier" on every citizen, it drew a boundary around government, assuming that if the state was forbidden from consolidating all that information in one place, individual privacy and human agency would be preserved.
What Congress failed to anticipate was that private firms would do what the state was barred from doing. Capital, answerable to no committee, built the infrastructure instead.
They built it at a scale nobody in that room could have pictured. Not statistical records, but everything: published work, personal data, forums, archives, whatever could be reached. Unlike the data center Congress refused, this one does not store records that can be audited. It absorbs them and dissolves them into weights.
Ask a question now, and you receive a synthesis rather than links. The labor Simon described, reading ten sources and deciding among them, is gone. So is everything that made the deciding possible. The answer is generated word by word from what the model learned across billions of parameters, and nothing in it records which works produced it.
Daddario's warning arrived on schedule. A 2026 Graphite study of web-crawl data shows that roughly half of what now gets published on the web was written by machines rather than by people, and nobody can look at a page and say with certainty which it is. Each generation learns from the last one's mistakes as though they were facts.
Act IV: Density Is Not Evidence
Quality is not a property you can read off a text. It is a judgment about who produced it and on what basis, which is why the source was never decoration. It was the entire basis for evaluation.
Without a source, you cannot tell whether an answer rests on evidence or repetition. A model's output tracks the density of its training data, and density measures how much was written on a subject, not how much of it was right. Nothing in the architecture distinguishes weight of evidence from weight of text.
Under the old arrangement, when a source was known, you could evaluate credibility immediately because the claim arrived with a name attached. You could tell a laboratory's peer-reviewed research from an influencer's blog before reading a word. People have always coped with abundance by choosing whom to believe. That is the filtering Simon described, and everyone does it without noticing. When an answer names no source, that choice is not overruled; it is removed.
This is Daddario's fear arriving intact, by a route he did not anticipate. He expected the danger to lie in accumulation, in too much held in systems too closed to examine. What exists is worse. No database to open, no record to subpoena, no file to correct. The source is not hidden. It is gone, and the decisions get made anyway.
Act V: The Record
The question is not how to see inside the model. Nothing will make that possible, and no regulation will reconstruct what has already been dissolved into weights. The question is how a work keeps its identity once the link that carried it is gone.
Print answered a version of this question once before. When the printing press broke provenance, the trade did not repair the medium. It added a page naming the answerable parties, and trust followed. This time the record cannot ride inside the work, because the medium dissolves whatever it carries. It must sit outside.
It begins with registration, before the work enters circulation: what the work is, who owns it, the terms of its use, and when the claim was made. The work is fixed as a fingerprint that survives republication, and the claim resolves to a persistent identifier any system can read. None of this requires the cooperation of the model that will eventually ingest the work.
Registration is only the first step. A work that can be identified can be recognized wherever it surfaces, and any system that retrieves it can be asked, before it uses the work, who owns it and on what terms. That loop is already running. Models talk to models, agents query agents, and machines generate, scrape, re-ingest, and re-synthesize content.
No one can retroactively recover which training data produced a given sentence. But every interaction from here forward that passes through a lookup or leaves a copy behind can be logged, and it must be logged now.
None of this is new arithmetic. The telecommunications industry priced the flow of global networks for forty years on the Call Detail Record, a line written every time traffic crossed from one operator's network to another's, and the operators settled against those lines without ever trusting each other. Registration, then attribution, then settlement is that same machinery, pointed at knowledge instead of minutes. That is what we are building at Alltio.
For three thousand years, nobody had to engineer this, because the material carried its own proof. The medium is weightless now. Trust must be engineered back into existence: every answer came from someone, and the record exists so we can still say whose work it was.