In 2012, an engineer at the National Informatics Centre named Neeta Verma led the design and launch of a portal called data.gov.in. It was built on a simple idea, formalized that same year in the National Data Sharing and Accessibility Policy: government data belongs to the public, and the state should push it out proactively rather than wait for citizens to ask for it one Right to Information application at a time. Verma would go on to become Director General of the NIC, the technical architect behind some of India’s most-used digital platforms, and eventually Chief Advisor for Information Technology to the Election Commission of India. The open data portal was one line in a much longer career. But it carried a specific promise: that a farmer, a journalist, a researcher, or a member of parliament would be able to go to one website and find government data in a usable, machine-readable form, published because publishing it was the law, not because someone had to fight for it.
Thirteen years later, in 2025, the government quietly rebuilt data.gov.in from scratch. The new version, called OGD 2.0, runs on a modern microservices architecture, allows municipal bodies to upload data directly through role-based access, and promises real-time availability at scale. Officials describe it as a leap forward. Read differently, it is an admission. You do not re-architect a platform after a decade if the original one was working. The rebuild is the government’s own verdict on the first attempt: the data was there, technically, but almost nobody could trust it, find it, or use it for anything that mattered.
This is the story of what happens when a government commits to openness in principle and then never builds the plumbing to make openness real. It is also the essay that closes a loop this series opened in its very first installment, the one about farmers trading crop advice over WhatsApp forwards because no official channel gave them anything better. That essay asked why citizens build informal data networks around a state that claims to be modernizing. This one answers: because the formal channel exists on paper, publishes something, and still fails the basic test of being usable.
The promise on paper
The National Data Sharing and Accessibility Policy of 2012 was, by the standards of the time, a genuinely progressive document. It asked every government department to catalog its datasets and make them available in open, non-proprietary formats. It sat alongside the Right to Information Act of 2005 as a second pillar of transparency: where RTI let a citizen pull information out of the state one request at a time, NDSAP was supposed to make the state push information out continuously, without being asked. Together they were meant to close the gap between “technically public” and “actually accessible.”
Data.gov.in became the visible face of that ambition. Over the following decade it accumulated a large catalog: by various counts it has held more than 8,300 dataset catalogs and in excess of 390,000 individual data resources, drawn from more than 165 government departments across 33 sectors. On paper, that is one of the larger open government data repositories in the world. It should have been a foundation good enough for private industry, researchers, journalists, and citizens to build on.
There is also the question of license. Verma’s team formulated the Government Open Data License, or GODL, the legal instrument that is supposed to tell a user what they are allowed to do with a downloaded dataset: whether it can be redistributed, modified, used commercially, cited a certain way. A license only matters if the dataset behind it is current and complete. GODL solved a real problem, permission, without solving the deeper one, quality. It is possible to have a perfectly clear legal right to use a dataset that is three years out of date and missing half its columns.
It is worth placing India’s experience against the global benchmark. The Open Data Inventory, an index that assesses roughly 190 countries on the completeness and openness of their official statistics, ranked Malaysia, Singapore, Finland, Denmark, and Poland at the top of its 2024-25 edition, each scoring in the 80s and 90s out of 100 on criteria like machine-readable formats, bulk download options, open licensing, and metadata quality, the same dimensions where India’s own reviewers say data.gov.in falls short. The gap is not a difference in ambition. India’s 2012 policy was, if anything, ahead of its time on paper. The gap is a difference in whether anyone was ever made responsible for keeping the promise current after the launch photo was taken.
The National Data Governance Framework Policy, still in draft form as it moves through the Ministry of Electronics and Information Technology, effectively concedes the same point in different language. It describes government data as being managed in “differing and inconsistent ways” across ministries and proposes a new body, the India Data Management Office, to impose a uniform national standard. You do not need a new office to standardize something that is already standardized. The proposal is itself evidence of the fragmentation it is meant to fix.
Why this looks familiar
Readers of this series will recognize the shape of the problem, because it is the shape of nearly every essay in it. This is the same structural failure examined in the piece on data feudalism: the state generates enormous volumes of data as a byproduct of running the country, but treats the finished product as something to be hoarded, mismanaged, or released in a form nobody outside the department that created it can actually use. It is the same failure examined in the piece on federal data sharing, where states generate the overwhelming majority of the data flowing into central government portals but frequently cannot get clean access to their own numbers back. And it is the same failure that opened this entire series: when the formal system will not give people usable, current, trustworthy data, they build informal ones instead. Farmers turn to WhatsApp groups. Journalists turn to RTI applications and private data vendors. Researchers turn to international agencies and multilateral databases that sometimes hold better, more current Indian statistics than the Indian government’s own portal does, because those organizations have the resources to clean and standardize what New Delhi published once and never touched again.
There is a particular irony in the fact that Indian public data is often easier to access reliably through the World Bank, the International Monetary Fund, or an academic dataset built by a foreign university than through the website the Indian government built specifically to distribute it. That is not a failure of ambition. NDSAP 2012 was ambitious. It is a failure of maintenance, incentives, and accountability, the unglamorous plumbing that never gets a ribbon-cutting ceremony but is the entire difference between a portal that works and one that merely exists.
The new audience for openness
The National Data Governance Framework Policy’s centerpiece, the proposed India Datasets Program, is worth examining closely, because it reveals who this second wave of “open data” is actually being built for. The program is designed to give anonymized, non-personal government and private-sector data to India-based artificial intelligence researchers and startups, coordinated through a platform that already exists in early form as AIKosh, hosted on the government’s AI hub. This is a legitimate and important goal. India’s AI ecosystem does need large, clean, representative training datasets, and government data is a natural source.
But notice the shift. The 2012 vision of open data was civic: a citizen should be able to see how the state is performing. The 2025 vision, as currently drafted, is industrial: a curated, non-personal dataset feed for AI developers and startups, managed through a new centralized office. Both goals are worth pursuing. They are not the same goal, and a policy optimized for one will not automatically deliver the other. A dataset useful for training a language model is not the same as a dataset that lets an opposition MP calculate how many promised toilets were never built, or lets a district journalist verify whether a irrigation scheme’s beneficiary numbers match what officials claimed at a public meeting. If the country’s data governance energy and budget flow toward the AI-training use case because it has more political and commercial momentum, the civic use case, the one this whole open data project was originally justified on, risks being quietly deprioritized a second time.
What it actually costs
None of this is abstract. NITI Aayog’s own report, India’s Data Imperative: The Pivot Towards Quality, released in June 2025, put a number on the cost of exactly the data quality failures described above. It found that errors and duplicate entries in government data systems can inflate welfare budgets by four to seven percent annually, money spent on records that are wrong, doubled, or no longer accurate, rather than on the people those programs are meant to serve. The same report noted that Aadhaar processed more than 27 billion authentications in the 2024-25 financial year, that UPI now clears roughly 23.9 trillion rupees in transactions every month, and that the Ayushman Bharat Digital Health ID program has issued more than 369 million IDs. India has built the plumbing to move data at extraordinary scale. What NITI Aayog’s report says, in effect, is that scale was never the problem the country needed to solve next. Quality, ownership, and interoperability are, and the state has been slower to admit that than its engineers have been to build new pipes.
What a real open data policy would require
A credible reform agenda here does not need a new slogan. India already had the correct instinct in 2012. What it needs is the unglamorous infrastructure of accountability that never got built the first time.
First, a statutory duty to maintain, not just publish. NDSAP created an obligation to release datasets. It created no obligation to keep them current, and departments have behaved exactly as an ungraded system predicts: many uploaded once and stopped. Any successor policy needs a maintenance clock attached to every dataset, with public visibility into how overdue an update is.
Second, an independent audit function. Comptroller and Auditor General reports already examine financial accounts department by department. There is no equivalent institution checking whether a ministry’s published dataset matches its internal records, or whether a Chief Data Officer’s release schedule is real or fictional. Without an auditor, “open” data quietly becomes theater.
Third, a working feedback channel. Citizens, journalists, and researchers who find an error in a government dataset today have no formal way to flag it and get a response. A platform that cannot be corrected by the people using it is not actually open, it is broadcast-only.
Fourth, deliberate investment in the civic use case alongside the AI-training use case. The India Datasets Program and AIKosh are worth building. So is a version of open data that helps an ordinary citizen check whether a scheme reached the village it was supposed to reach. Both deserve a budget line. Only one currently has political momentum.
Fifth, tie Chief Data Officer performance to usage and quality metrics, not upload counts. A ministry that publishes a dataset once and never updates it should not score the same as one that keeps its data current, documented, and correctable. Right now, both get credit for having “released” something.
The unwritten number
Every essay in this series has ended on the same structural point: somewhere in the gap between what the state claims to know and what it can actually produce, there is a number nobody has ever calculated, the true cost of not knowing. For open data, that number is not hidden in a database the government refuses to release. It is hidden in the difference between a portal that exists and a portal that works, between data that is technically public and data that is genuinely usable. Neeta Verma’s team built the first version of that promise in 2012. The country is now, fourteen years later, rebuilding it a second time. The question this reform essay leaves open is whether the second version will finally be held to the standard the first one never was, or whether in another decade a third generation of engineers will be asked to rebuild it again.

