alltio.
Blog · Regulation

The world’s largest internet market just rewrote the rules for data — and put an end to free AI training

India's DPDP Act, its new AI governance framework, and a copyright proposal that rejects free training data.

By Carl Gunell, Founder · July 19, 2026 · 7 min read

Ask an American AI executive about the EU AI Act and you will usually get an answer. Ask about India's Digital Personal Data Protection Act, and you will usually get a blank look. That is a mistake. India is the world's largest internet market by user count, one of the largest producers of the text, code, and media that AI systems learn from — and over the past nine months it has assembled, piece by piece, a data and AI governance framework that any company training or deploying models at global scale will have to answer to.

Three developments matter, and they matter together.

1. The DPDP Act is now operational — and it is consent-first

India passed the Digital Personal Data Protection Act in 2023, but the machinery to enforce it arrived in November 2025, when the Ministry of Electronics and Information Technology notified the final DPDP Rules and stood up the Data Protection Board. Implementation is phased: the Board is already operational and taking complaints, the consent-manager framework opens in November 2026, and full compliance is due by May 2027 — with government consultations in early 2026 exploring whether to compress that timeline further. Penalties run up to ₹250 crore, roughly $30 million, per violation.

What makes the DPDP regime distinctive — and demanding — for AI is its architecture. Unlike Europe's GDPR, which offers several lawful bases for processing personal data, including “legitimate interests,” India's law is built almost entirely on consent and a narrow list of legitimate uses. Two consequences follow directly for AI development:

Purpose limitation bites hard. Data collected to deliver a service — an e-commerce transaction, a messaging app — cannot simply be repurposed to train a foundation model. That secondary use requires its own basis. Any company whose training pipeline draws on Indian user data inherits this constraint.

Erasure meets machine learning. The Act gives individuals the right to withdraw consent and have their data erased. For a trained model, this raises a question the industry has not solved: if a model cannot “unlearn” a data point without retraining, what does compliance look like? India is forcing that question onto the engineering roadmap.

Web-scraped training corpora sit uneasily inside this framework. Personal data is embedded in text, images, and metadata across the open web, and obtaining valid consent from each individual in a scraped dataset is not feasible. Companies training on Indian-originated content will need documented provenance for what they used and on what basis — not assurances, records.

2. India's AI Governance Guidelines chose a “techno-legal” path

In November 2025, MeitY released the India AI Governance Guidelines — the country's first comprehensive AI governance framework, built from a multi-year process that drew more than 2,500 public submissions. India deliberately declined to copy Europe's approach. Rather than a standalone AI Act, the Guidelines apply existing law through a risk-based lens and lean on what the drafters call “techno-legal” regulation: embedding legal principles directly into technical systems.

The clearest example is content provenance. Confronting deepfakes and synthetic media, the Guidelines recommend adopting global provenance standards — explicitly naming C2PA — and watermarking, placing governance at the protocol level rather than relying on after-the-fact enforcement. The same instinct runs through the framework's treatment of data: verifiable records, machine-readable terms, compliance by design.

For anyone who has watched infrastructure markets form, this is a familiar signal. When a regulator says “the record itself must carry the proof,” it is describing registries, attestation, and audit trails — not policy documents.

3. On copyright, India is heading toward licensing — not a free pass

The most consequential development is the least reported in the United States. The Guidelines flagged that India's Copyright Act may not cover the use of copyrighted works in AI training at all, and a committee under the Department for Promotion of Industry and Internal Trade was tasked with closing the gap. Its first Working Paper, published in December 2025 for consultation, examined the regulatory options — and rejected a blanket text-and-data-mining exception of the kind Japan and Singapore adopted, on the ground that it amounts to a zero-price license that leaves Indian creators uncompensated. It equally recognized that voluntary bilateral licensing cannot work at the scale of modern training corpora. What it proposed instead is a hybrid model: training may proceed, but compensation to rights holders is built into the framework.

Meanwhile, the Delhi High Court is weighing ANI Media's case against OpenAI — judgment was reserved in March 2026 and is awaited — which may produce India's first substantive judicial ruling on whether training on publicly available copyrighted works is fair dealing under Indian law.

Read the direction of travel. The world’s largest internet market is converging on the position that AI training on creative and knowledge work is a licensable event — permitted, priced, and paid — rather than a free harvest or a flat prohibition.

Why Americans should care

First, exposure: any model trained on the open web has ingested Indian-originated content, and any AI product serving Indian users falls under the DPDP regime as it phases in. Second, precedent: with the EU mandating training-data transparency and India designing compensated-training frameworks, two of the three largest digital markets are aligned on the principle that provenance must be documented and use must be settled. Jurisdictions do not usually walk that kind of consensus back. American creators who assume the question of AI compensation will be resolved solely in US courts are watching the wrong theater.

Third — and this is the part I find most striking — India's proposed answer has no infrastructure to run on yet. A hybrid licensing model requires exactly the machinery that bespoke deals cannot provide: a registry that knows whose work is whose, machine-readable terms, usage reporting both sides accept as evidence, and settlement that reaches individual rights holders in their own currency. That is not a policy problem. It is a clearing problem — the same one telecom solved for international settlement and music solved for collective licensing.

Alltio was built as that layer: a neutral registry of works, standardized licensing contracts, cryptographically signed usage reports, and a settlement engine that pays every registered rights holder. When frameworks like India's move from working paper to law, the record will already need to exist. The time to put your work on it is before the regulators ask — not after.

Registered with Alltio — authorship and licensing terms on record, SHA-256 fingerprint

Registered with Alltio · record c67906e3…bee2eec4 · registered 2026-07-22

Put your knowledge on the record.

Free during the founding period — fingerprinted in your browser, never uploaded.

Register an asset

The Clearing Layer

Sign up for The Clearing Layer

Alltio’s newsletter on the AI knowledge economy — licensing deals, regulation, physical-AI data, and what each one means for the people who own the inputs. No promotion, no noise, unsubscribe any time.