STAT & STATE STAT & STATE Complexity, Clarified

From Aadhaar to AI: The Stat Behind India's New Sovereign State

How a 2-billion-parameter model, a 4,096-GPU grant, and two DPI veterans are building India's answer to the global AI race.

By Stat & State Desk
Updated August 2026

SNEAK PEEK

Sarvam raised $234M at a $1.5B valuation. Its Indic tokenizer is 4x more efficient than Western models.

Sarvam AI has closed a $234 million Series B led by HCLTech to build India's sovereign AI stack. By training models like Sarvam-1 (2B) and Sarvam-105B ('Indus') from scratch on Indic languages, the startup achieves a token efficiency of 1.4 to 2.1 tokens per word for Indian languages, compared to 4 to 8 tokens for global models. Backed by a ₹246.72 crore IndiaAI Mission grant, Sarvam is attempting to build the digital infrastructure that India's intelligence stack will run on.

STORY

THE STAT

A country of 1.4 billion people, speaking across 22 scheduled languages, is increasingly relying on artificial intelligence to interact with the digital economy. Yet there are zero major global large language models that were originally trained with Indic languages as their core priority. The frontier models building the global technology wave are English-first, Silicon Valley-shaped, and priced for a dollar economy.

For a decade, India solved this kind of foundational infrastructure problem by building its own open-access digital layers rather than importing someone else’s closed systems. The biometric system of Aadhaar provided digital identity for 1.4 billion citizens, while the Unified Payments Interface (UPI) built a payments rail that clears over 20 billion transactions monthly. In terms of sheer transaction count, this domestic network now rivals Visa’s global processing volume of roughly 21.4 billion transactions per month.

This is India’s Digital Public Infrastructure (DPI) playbook: open, unified digital layers backed by the state that the private market plugs into. Artificial Intelligence represents the next layer of that stack, and Sarvam AI is positioning itself as the builder of India’s sovereign intelligence layer.

UPI MONTHLY TRANSACTIONS
20B+

The volume of payments cleared monthly by India's UPI, surpassing or rivaling the global transaction counts of major credit card networks.


THE STATE

Sarvam AI was founded in July 2023 by Dr. Vivek Raghavan and Dr. Pratyush Kumar, two veterans who helped shape India’s digital public infrastructure landscape over the last decade. Dr. Raghavan spent more than ten years at the UIDAI as a chief biometric architect, designing the deduplication engines that allow Aadhaar to verify billions of records. Dr. Kumar co-founded AI4Bharat at IIT Madras, building open Indic datasets and linguistic benchmarks.

Their core product pitch is Sovereign AI: the belief that a nation’s most vital intelligence infrastructure should be built, hosted, and governed domestically. In June 2026, the capital markets backed this vision when Sarvam closed a $234 million Series B funding round at a $1.5 billion post-money valuation, led by HCLTech with a $150 million commitment alongside Bessemer Venture Partners, Khosla Ventures, and Peak XV Partners.

The Indian state also holds active skin in the game. Under the Innovation Centre pillar of the IndiaAI Mission, Sarvam received a ₹246.72 crore compute allocation, granting it subsidized access to 4,096 NVIDIA H100 GPUs hosted on local cloud infrastructure. This support was structured as Compulsorily Convertible Debentures (CCDs), giving the Government of India a convertible option to acquire a 1–2% equity stake in Sarvam during its financing round.

Sarvam AI Cumulative Capital Raised ($ Millions)

Loading interactive graphic...
Figure 1: Sarvam AI cumulative capital raised over time, including its Series A and the first close of its Series B. Data source: Peak XV and HCLTech regulatory filings (2026).

THE TECHNOLOGY

The key technical battleground is tokenization, the process by which a language model breaks down text into numerical fragments. Because global models are optimized for English and European characters, they fragment Indic words heavily. A single word in Hindi, Tamil, or Bengali often costs four to eight tokens in Western models, whereas Sarvam’s custom tokenizer achieves an efficiency of 1.4 to 2.1 tokens per word.

This tokenizer performance reduces inference costs directly, as API charges scale with token usage. It also improves latency by reducing the generation steps required, while preserving the model’s effective context window.

The Tokenization Tax (Tokens per 1,000 Indic Words)

Loading interactive graphic...
Figure 2: Tokens required to process 1,000 words in Indian languages. Fewer tokens translate directly to faster responses and lower inference costs. Data source: Sarvam AI Tokenizer Technical Whitepaper (2026).

Beyond foundational text processing, Sarvam has developed a full product suite. This includes the compact Sarvam-1 (2B) model, the larger Sarvam-30B, and the flagship Sarvam-105B Mixture-of-Experts model, which operates under the open-source Apache 2.0 license. The suite extends to speech-to-text with Saaras, speech synthesis with Bulbul, and the Kaze smart glasses, which run Indic models locally on the edge.


THE COMPETITIVE STATE

Every sovereign AI pitch must answer how it compares to global frontier laboratories. Rather than competing directly on general English benchmarks, Sarvam has focused on cost-efficient local deployment. By using a Mixture-of-Experts (MoE) architecture for the flagship Sarvam-105B model, they ensure that only about 10.3 billion active parameters are triggered per token, allowing the model to run on less expensive local servers.

This architectural focus creates clear specialization trade-offs. Sarvam wins on regional language fluency, colloquial code-mixed inputs like Hinglish, local data residency compliance, and overall cost per query. Western frontier models like GPT-4o, Gemini 1.5, and Claude 3.5 maintain a significant lead in complex coding, mathematical reasoning, and general English logic.

This division of labor shows up in benchmark evaluations. While Sarvam-105B scores competitively on general reasoning tests, it is optimized to execute agentic workflows, database queries, and payment integrations in regional Indian languages where Western models struggle.

Sarvam-105B vs. Peer Models (Benchmark Scores %)

Loading interactive graphic...
Figure 3: Sarvam-105B’s reasoning and mathematics performance compared to standard 70B-100B parameter class peers. Data source: Sarvam AI Technical Evaluation Report (March 2026).
THE WIN

Local Fluency vs. Global Scale

Sarvam achieves a 90% pairwise win rate against global frontier models on fluency and naturalness in major Indic languages, prioritizing domestic utility over universal benchmarks.


THE SOVEREIGN STACKS EXPLORER

Compare the sovereignty stacks, specialization vectors, and benchmark outcomes using the interactive dashboard below.

The Sovereignty Stack

Click on any layer (Applications, Models, or Compute) to compare the domestic stack with the imported alternative.

🇮🇳 SOVEREIGN STACK
🌐 IMPORTED STACK
Foundational Models Layer
Sovereign Approach

Built from scratch on Indic datasets (like Sarvam-1 and Sarvam-105B) and housed locally. They are optimized for regional syntax and tokenization speed, and are open-sourced under Apache 2.0 licenses.

Imported Approach

Proprietary weights and weights-as-a-service APIs. Models are English-first; regional languages are treated as second-class inputs, causing high fragmentation and token charges.


THE TRIGGER

India has already established its capability to build population-scale digital rails. The push into AI represents the same playbook applied to the intelligence layer: local founders with public sector experience, capital from domestic giants, and a technical tokenizer advantage where global hyperscalers face natural limitations.

While a $1.5 billion valuation represents a major investment, independent validation will remain critical as government and enterprise integrations scale. The direction, however, is clear: the country is actively building its own computing infrastructure rather than relying entirely on external developer ecosystems.


Data Sources: IndiaAI Mission (MeitY); Sarvam AI Technical Evaluation Report (March 2026); HCLTech Regulatory Filings (June 2026); NPCI Monthly UPI Statistics (2026); Visa Inc. Annual Reports (2025).
Note: Cumulative funding includes the initial Series A and the first close of the Series B.
Last Updated: August 2026