Nigeria’s AI Ambition Needs a Data Foundation
We should invest aggressively in AI. But not every problem is an AI problem, and AI can't make up for digital infrastructure we haven't built.
contents (11)
- AI is not an oracle
- Sometimes a database is the innovation
- Drug verification is another good example
- The trade-data project raises the same question
- Nigeria’s own AI strategy recognizes the problem
- Build the boring things
- Imagine a government API layer
- Where AI clearly belongs
- A better test for government AI projects
- Data infrastructure is AI infrastructure
- Build the infrastructure, then let Nigerians surprise us
Nigeria should be investing in artificial intelligence. I want to say that first, so the argument that follows isn’t mistaken for resistance to AI.
The technology is moving too fast, and its economic impact is too large, for Nigeria to sit on the sidelines. Building local expertise matters. So does supporting Nigerian researchers, developing models that understand Yoruba, Hausa, Igbo and Nigerian English, and giving Nigerian developers access to models, compute and funding.
I’m particularly interested in what happens when AI is applied to public services. But the more I look at some of our AI initiatives, the more I think we sometimes start at the wrong end of the problem. We ask:
How can we use AI here?
when the first question should be:
What is the problem, and what is the simplest reliable technology that solves it?
Those are different questions, and they lead to different systems.
AI is not an oracle
People often talk about modern AI systems as if we’ve built machines that simply know things. We haven’t.
Large language models are remarkable. They can work with language, find patterns, summarize large amounts of information, translate, classify, reason across sources and put a natural-language interface on complicated systems. They can also be wrong.
More importantly, a model can’t create authoritative information that doesn’t exist. Suppose a government agency has incomplete records, inconsistent naming, inaccessible databases, outdated websites, information buried in PDFs and no reliable API. Putting an AI interface in front of it fixes none of that. You just have AI sitting on top of bad information infrastructure.
That matters most when the system belongs to a government.
Sometimes a database is the innovation
I recently went through the projects listed by Nigeria’s National Centre for Artificial Intelligence and Robotics (NCAIR).1 One caught my attention immediately: using machine learning to identify top AI researchers of Nigerian descent.
NCAIR describes gathering data such as academic publications, conference presentations, patents and other indicators of research impact, then using machine learning to identify the researchers who have contributed most to the field.
There may be a place for machine learning here at scale, for example in entity resolution (working out that three differently spelled names belong to one person) or in classifying research areas across millions of records. But the description raises a more basic engineering question: what part of this problem actually requires AI?
Research databases already hold structured information about authors, publications, institutions, citations and topics. Much of the problem can be handled with data aggregation, deterministic ranking, search and human verification.
There’s nothing wrong with that. If a database, a search index and a well-designed algorithm solve a problem more reliably and transparently, use them. Government technology should optimize for outcomes, not novelty.
Drug verification is another good example
Another NCAIR project, with the Design Lab at CcHUB, proposes using Meta’s Llama 3 to build a platform for verifying and validating pharmaceutical products.1 Counterfeit medicine is a serious problem, and AI could contribute a lot here.
A citizen could photograph a pack of medicine and have computer vision read the product details. Someone could ask questions in Yoruba or Hausa. A pharmacist could query complicated regulatory information in plain language. AI could flag suspicious patterns across supply-chain data.
But the core question, is this product legitimate?, should be answered from authoritative data:
- Who manufactured it?
- Is it registered, and under what number?
- Which batch does it belong to, and has that batch been recalled?
- When does it expire?
- Who imported or distributed it?
Those are database questions. AI can be an excellent interface to that system, but it shouldn’t become the system’s source of truth. The difference sounds small, yet it changes the whole architecture.
The trade-data project raises the same question
NCAIR also describes fine-tuning Llama 2 on 20 years of product-level Nigerian trade data, alongside Meta’s Social Connectedness data.1 There may be interesting research questions in that.
But consider a researcher or policymaker asking:
What were Nigeria’s five largest non-oil exports to Ghana last year, and how have they changed over five years?
I don’t want a language model to remember those numbers. I want a system that retrieves authoritative trade records, does the calculation, and then lets the language model explain what the numbers mean:
question → structured data → computation → AI explanation → sourcesrather than:
question → AI model → answerThe first design gives us something government badly needs: provenance. We can ask where a number came from.
Nigeria’s own AI strategy recognizes the problem
This argument isn’t at odds with Nigeria’s National Artificial Intelligence Strategy. The strategy states that “access to quality data is fundamental to developing robust and reliable AI systems.” It notes that many Nigerian datasets “suffer from inaccuracies, incompleteness, and a lack of standardisation,” and that different sectors and organizations “maintain their data repositories with varying standards and formats.”2
Its answer is to adopt globally recognized data-quality standards and to create an Open Data Initiative that brings public and private data together.2
So the policy understands the problem. My concern is emphasis and implementation. If data is the foundation for useful AI, building that foundation should be one of the most visible and aggressive parts of our AI programme, because its benefits reach well beyond AI.
Build the boring things
Nigeria needs more boring technology. I mean that as a compliment.
- Give government entities stable identifiers.
- Give public datasets consistent schemas.
- Publish APIs, document them properly, and version them instead of silently changing them.
- Turn important PDF-only datasets into machine-readable records.
- Standardize how ministries and agencies publish public information.
- Make government websites usable on inexpensive phones.
- Build shared components for digital public services.
- Define interoperability standards so one government system can reliably talk to another.
- Record when information was collected, who produced it and when it was last updated.
- Make non-sensitive public data genuinely easy for developers, researchers, journalists and businesses to use.
None of this is as exciting as announcing a new large language model. But it creates something more powerful: an ecosystem that thousands of applications can be built on.
Imagine a government API layer
Imagine developers could reach Nigerian public information through well-documented interfaces, with resources like these:
GET /agenciesGET /budgetsGET /contractsGET /legislationGET /representativesGET /schoolsGET /health-facilitiesGET /tradeGET /public-servicesThe URL structure isn’t the point. What matters is what sits behind it: common identifiers, predictable schemas, documentation, historical records, provenance and dependable access.
That’s when AI gets genuinely interesting. A journalist could ask:
Show me federal road contracts awarded in Oyo State since 2023, compare reported completion against disbursements, flag unusual cases and give me the original records.
The AI doesn’t need to invent anything. It translates the question, queries authoritative systems, runs the analysis, explains the result and links back to the evidence.
A citizen could ask the same infrastructure:
Who represents my constituency, what bills have they sponsored, and how did they vote on the last education bill? Explain it to me in Yoruba.
That uses what AI is extraordinarily good at while keeping factual authority where it belongs. It’s a far more interesting version of AI-enabled government.
Where AI clearly belongs
None of this means government should wait for perfect databases before deploying AI. That would be just as misguided. Some Nigerian problems need capabilities that conventional software can’t provide economically.
Language is the obvious one. NCAIR’s N-ATLaS is an open-source, multilingual model covering Yoruba, Hausa, Igbo and Nigerian-accented English,3 with automatic speech recognition for those languages.4
That could matter a great deal for public services. Picture a citizen who can’t comfortably navigate a complicated government website but can send a voice message in Hausa. Speech recognition understands the request, an AI system works out what they need, government APIs return the authoritative information, and the model explains the answer in a language they understand.
That isn’t AI looking for a problem. It’s AI removing a real barrier.
Document extraction is another strong case. So are translation, fraud and anomaly detection, satellite imagery analysis, classifying huge unstructured archives, and natural-language interfaces to complicated public systems. The question isn’t whether something uses AI. It’s what AI contributes to the solution.
A better test for government AI projects
Before approving an AI project, we should be able to answer one question clearly:
What becomes possible, or substantially cheaper, faster or better, because this system uses AI?
If the answer is unclear, try solving the problem without AI. Maybe the answer is PostgreSQL. Maybe it’s Elasticsearch, an API or a deterministic rules engine. Maybe it’s cleaning a dataset nobody has maintained properly for fifteen years.
That isn’t a failure of ambition. It’s engineering. And when AI is the right tool, use it aggressively.
Data infrastructure is AI infrastructure
The most important shift I’d like to see is in how we think about this. Open data programmes, government APIs, interoperability standards, data-quality work and digital public infrastructure shouldn’t be treated as separate from Nigeria’s AI ambitions. They are part of the AI strategy.
The national strategy already points this way by naming data quality, standardization and open data as priorities. The next step is making that visible in what we build. A useful national technology stack might look like this:
Authoritative government systems ↓Clean, structured and versioned data ↓Common identifiers and standards ↓Interoperability and documented APIs ↓Privacy, permissions and provenance ↓Public and private developer ecosystems ↓AI models, agents and intelligent applications ↓Citizen servicesAI sits near the top for a reason: every layer underneath makes it more capable. And those layers keep producing value even when no AI is involved.
Build the infrastructure, then let Nigerians surprise us
Government doesn’t need to predict every useful application developers will build. It needs to create the conditions that make them possible.
Nigeria’s fintech success shows what happens when developers can build on accessible digital rails. A government-data equivalent could unlock applications nobody has thought of yet. Researchers, startups, journalists, civic organizations, universities and government agencies would all use it. So would AI companies, probably in ways none of us can predict.
Nigeria should pursue its AI ambitions. NCAIR’s stated mission is to promote the research, development and adoption of AI “for economic growth, improved quality of life, and global competitiveness,”3 and it is already building local models and funding AI research.1 Those are worthwhile goals.
But ambition also needs discipline about where a technology belongs. Sometimes the right answer is a language model. Sometimes it’s machine learning. And sometimes the revolutionary technology Nigeria needs is a clean table, a documented API and a server that reliably returns 200 OK.
We should get very good at knowing the difference.
Footnotes
-
National Centre for Artificial Intelligence and Robotics (NCAIR), “Work Done So Far.” Describes the researcher-identification, pharmaceutical-verification (Llama 3) and trade-data (Llama 2) projects, and NCAIR’s AI research grants. ↩ ↩2 ↩3 ↩4
-
NCAIR, NITDA and the Federal Ministry of Communications, Innovation and Digital Economy, National Artificial Intelligence Strategy (September 2025), Objective 3.2, p. 40. ↩ ↩2
-
NCAIR, homepage. Mission statement and the N-ATLaS language model. ↩ ↩2
-
Open Source For You, “Nigeria Unveils Open Source N-ATLAS AI Model To Champion African Languages At UNGA80,” 22 September 2025. ↩