Header Sticky Banner
OPINION
#Opinion

The Zero-Cost Infrastructure for the Nepali AI Ecosystem

In public policy terms, every noncompliant document published today raises the future cost of remediation—an accumulating negative externality.
alt=
Representative Photo
By Nischal Dhungel , Sanjib Lamichhane

Nepal's policymakers keep discussing artificial intelligence as a strategy problem: which ministry should own it, which foreign partner should fund it, which pilot project should launch first. This framing misses the actual constraint. The problem is not strategy. It is a raw material. Before any AI system can read, summarize, or reason about Nepal's public documents, those documents must be usable text. Right now, most of them do not.



From an economics perspective, machine-readable government data is a pure public good. It is non-rival (one researcher’s use does not diminish another’s) and non-excludable once published. Each municipality that releases clean Unicode documents generates positive externalities for the entire ecosystem—startups, auditors, journalists, citizens—yet captures almost none of the social return. The predictable result is under-provision: a classic market failure that persists inside the government itself. A single, enforceable publishing standard can correct this government failure at near-zero fiscal cost.


We analyzed a self-curated dataset of roughly 48,000 Nepali local government PDFs drawn from all 753 municipalities. Of 908 documents examined in detail, 75 percent required Optical Character Recognition (OCR), software that scans and extracts text from images, just to produce readable text. Only 25 percent were genuinely machine-readable Unicode documents. Three-quarters of what the state already publishes digitally does not qualify as digital text. This is not a shortage of documents. It is a shortage of documents that a machine can actually read.


The economic cost is already measurable in wasted labor. Every hour spent by researchers or civil society groups running OCR and reconstructing tables is an hour not spent analyzing budgets or building tools—an opportunity cost that compounds across 753 municipalities. Once templates and habits change, the marginal cost of producing the next document correctly falls essentially to zero, while the social return continues indefinitely. The current equilibrium, therefore, generates a growing “data debt”: each non-compliant PDF raises the future rescue costs, a textbook stock-flow problem familiar from infrastructure and environmental economics.


Related story

Local governments allowing mining in Chure to generate income


Two institutional failures drive most of the damage. The first is path dependence. Decades of legacy fonts (Preeti, Kantipur, PCS) locked the bureaucracy into encodings that look like Nepali text but are extracted as Latin gibberish. These fonts create technological lock-in: the private cost of switching appears high to individual offices, even though the social cost of remaining locked in is far higher. The second is the substitution of scans for born-digital text. Budgets and council minutes are routinely published as image-only PDFs, destroying both searchable text and table structure. Flat OCR severs the relationship between numbers and their row-column meaning—the highest-value data in the corpus becomes the least accessible.


The second failure is scans instead of text. Many of the documents that matter most, budgets, council minutes, and development plans, are photographs of printed pages with no text layer at all. Every word has to be reconstructed via OCR, and Devanagari OCR remains weaker than Latin-script tooling. Conjuncts, matras, and the halant are frequently misread. Worse, flat OCR destroys table structure. Budget documents are dense multi-column tables where a number only means something in relation to its row and column. Standard OCR reads the words but loses the grid, so an allocation amount detaches from the project it was allocated to. This is the highest-value data in the entire corpus, and it is the least accessible.


Compounding both problems: Nepali digits with unit declarations that live nowhere near the numbers they modify, zero standardization of layout or filenames across 753 municipalities, and no central repository, only 753 separate websites of loose PDF downloads. None of these problems is hard in isolation. Together, they make producing one usable Nepali training example orders of magnitude more expensive than producing one in English, not because the language is under-resourced in principle, but because the pipeline to it is broken in practice.


The remedy does not require a new ministry, a national AI strategy, or additional budgetary outlays; it requires only a change in institutional rules. Six simple publishing standards would resolve nearly all of the problem: mandate Unicode Devanagari only for all new government publications; prohibit image-only PDFs whenever a text original exists; Adopt global digital accessibility standards (specifically WCAG 2.1 AA) as the baseline, since documents accessible to humans with disabilities are by design machine-readable; publish every table as a structured data file (CSV or JSON) alongside the PDF at zero extra cost from the same spreadsheet; enforce a common file-naming convention and document schema; and establish one central repository while preserving local flexibility in presentation.


These six items are not a technology upgrade. They are a change in habit, enforceable through a single circular issued by the Ministry of Federal Affairs and General Administration or the National Information Commission, at no fiscal cost, with no new hiring, and with no dependency on outside vendors.


Individual offices need not wait. With software already on their computers, they can begin tomorrow: type in Unicode, never scan a born-digital document, export tables as separate files, adopt consistent naming, declare units in metadata, tag headings properly, and designate one staff member as the pre-publication standards checkpoint. The short-run opportunity cost—staff time spent mastering new habits—is real, but it is front-loaded and finite. Once templates are rewritten, marginal compliance cost drops to zero. The distribution of gains and losses is favorable: the few offices locked into legacy workflows are far outnumbered by the beneficiaries—central auditors, civil-society monitors, private language-technology firms, and citizens who gain transparent local budgets.


Policy instincts that treat data standards as a “nice-to-have” to be phased in later reverse the correct sequencing. In public policy terms, every noncompliant document published today raises the future cost of remediation—an accumulating negative externality. Treating standards as optional guarantees they will be ignored precisely when deadline pressure is highest, the moment most government documents are produced.


Nepal cannot import a ready-made AI infrastructure for its language; Devanagari-specific tokenization and normalization must be done domestically. What it can do is stop destroying its own raw material. Clean public text lowers the fixed costs of subsequent Nepali-language AI applications, raises the expected private return on investment in language technology, and improves the efficiency of public expenditure monitoring. Because usable data is a stock that grows only when the flow is fixed, the delay itself is costly. The social rate of return on this institutional reform is therefore unusually high: the private cost of the remedy approaches zero while the long-run social benefits compound.


Stopping the destruction of the country’s linguistic raw material is the cheapest piece of digital infrastructure Nepal can build. It requires no new money, no new organizations, and no foreign partners—only a change in the rules that govern how the state already publishes what it already produces. The alternative is permanent high marginal costs for every future attempt to make artificial intelligence work in Nepali.


(Nischal Dhungel is a fellow at the Nepal Institute for Policy Research (NIPoRe) and has an MSc in Economic Theory and Policy from Bard College, New York. Sanjib Lamichhane is a graduate-level AI research scholar at East Texas A&M University, USA.)

See more on: Nepali AI Ecosystem
Related Stories
ECONOMY

Rs 740 million allocated for digital infrastructur...

qNrbXIo4dCJ2bQUpjYaUM1nYmogGDV0AlH8t5oeV.jpg
SOCIETY

NICCI AGM highlights roadblocks to investment in N...

DSC09239-1766368179.webp
SOCIETY

Nepal’s performance in information ecosystem satis...

VIBEreport_20240626183029.JPG
ECONOMY

NRB and IFC join hands to strengthen Nepal’s finte...

WhatsAppImage2023-09-01at12_20230901160206.11
ECONOMY

NAJCC to collaborate with Nepal in ecosystem, wate...

NAJCC to collaborate with Nepal in ecosystem, water management