Metaverse
Inside India’s two-track strategy to become an AI powerhouse – Crypto News
With 22 official languages and hundreds of spoken dialects, India faces a monumental challenge in building AI systems that can work across this multilingual landscape.
In the demo area of the event, this challenge was front and centre, with startups showcasing how they’re tackling it. Among those were Sarvam AI, demonstrating Sarvam-Translate, a multilingual model fine-tuned on Google’s open-source large language model (LLM), Gemma.
Next to it, CoRover demonstrated BharatGPT, a chatbot for public services such as the one used by the Indian Railway Catering and Tourism Corporation (IRCTC).
At the event, Google announced that AI startups Sarvam, Soket AI and Gnani are building the next generation of India AI models, fine-tuning them on Gemma.
At first glance, this might seem contradictory. Three of these startups are among the four selected to build India’s sovereign large language models under the ₹10,300 crore IndiaAI Mission, a government initiative to develop home-grown foundational models from scratch, trained on Indian data, languages and values. So, why Gemma?
Building competitive models from scratch is a resource-heavy task involving multiple challenges and India does not have the luxury of building from scratch, in isolation. With limited high-quality training datasets, an evolving compute infrastructure and urgent market demand, the more pragmatic path is to start with what is available.
These startups are therefore taking a layered approach, fine-tuning open-source models to solve real-world problems today, while simultaneously building the data pipelines, user feedback loops and domain-specific expertise needed to train more indigenous and independent models over time.
Fine-tuning involves taking an existing large language model already trained on vast amounts of general data and teaching it to specialize further on focused and often local data, so that it can perform better in those contexts.
Build and bootstrap
Project EKA, an open-source community driven initiative led by Soket, is a sovereign LLM effort, being developed in partnership with IIT Gandhinagar, IIT Roorkee and IISc Bangalore. It is being designed from scratch by training code, infrastructure and data pipelines, all sourced within India. A 7 billion-parameter model is expected in the next four-five months, with a 120 billion-parameter model planned over a 10-month cycle.
“We’ve mapped four key domains: agriculture, law, education and defence,” says Abhishek Upperwal, co-founder of Soket AI. “Each has a clear dataset strategy, whether from government advisory bodies or public-sector use cases.”
A key feature of the EKA pipeline is that it is entirely decoupled from foreign infrastructure. Training happens on India’s GPU cloud and the resulting models will be open-sourced for public use.
The team, however, has taken a pragmatic approach, using Gemma to run initial deployments. “The idea is not to depend on Gemma forever,” Upperwal clarifies. “It’s to use what’s there today to bootstrap and switch to sovereign stacks when ready.”

View Full Image
CoRover’s BharatGPT is another example of this dual strategy in action. It currently runs on a fine-tuned model, offering conversational agentic AI services in multiple Indian languages to various government clients, including IRCTC, Bharat Electronics Ltd, and Life Insurance Corporation.
“For applications in public health, railways and space, we needed a base model that could be fine-tuned quickly,” says Ankush Sabharwal, CoRover’s founder. “But we have also built our own foundational LLM with Indian datasets.”
Like Soket, CoRover treats the current deployments as both service delivery and dataset creation. By pre-training and fine-tuning Gemma to handle domain-specific inputs, it is trying to improve accessibility today while building a bridge to future sovereign deployments.
“You begin with an open-source model. Then you fine-tune it, add language understanding, lower latency and expand domain relevance,” Sabharwal explains.
“Eventually, you’ll swap out the core once your own sovereign model is ready,” he adds.

View Full Image
Amlan Mohanty, a technology policy expert, calls India’s approach an experiment in trade-offs, betting on models such as Gemma to enable rapid deployment without giving up the long-term goal of autonomy. “It’s an experiment in reducing dependency on adversarial countries, ensuring cultural representation and seeing whether firms from allies like the US will uphold those expectations,” he says.
Mint reached out to Sarvam and Gnani with detailed queries regarding their use of Gemma and its relevance to their sovereign AI initiatives, but the companies did not respond.
Why local context is critical
For India, building its own AI capabilities is not just a matter of nationalistic pride or keeping up with global trends. It’s more about solving problems that no foreign model can adequately address today.
Think of a migrant from Bihar working in a cement factory in rural Maharashtra, who goes to a local clinic with a persistent cough. The doctor, who speaks Marathi, shows him a chest X-ray, while the AI tool assisting the doctor explains the findings in English, in a crisp Cupertino accent, using medical assumptions based on Western body types. The migrant understands only Hindi and much of the nuance is lost. Far from being just a language problem, it’s a mismatch in cultural, physiological and contextual grounding.
A rural frontline health worker in Bihar needs an AI tool that understands local medical terms in Maithili, just as a farmer in Maharashtra needs crop advisories that align with state-specific irrigation schedules. A government portal should be able to process citizen queries in 15 languages with regional variations.
These are high-impact and everyday use cases where errors can directly affect livelihoods, functioning of public services and health outcomes. Fine-tuning open models gives Indian developers a way to address these urgent and ground-level needs right now, while building the datasets, domain knowledge and infrastructure that can eventually support a truly sovereign AI stack.
This dual-track strategy is possibly one of the fastest ways forward, using open tools to bootstrap sovereign capacity from the ground up.
“We don’t want to lose the momentum. Fine-tuning models like Gemma lets us solve real-world problems today in applications such as agriculture or education, while we build sovereign models from scratch,” says Soket AI’s Upperwal. “These are parallel but separate threads,” says Upperwal. “One is about immediate utility, the other about long-term independence. Ultimately these threads will converge.”
A strategic priority
The IndiaAI Mission is a national response to a growing geopolitical issue. As AI systems become central to education, agriculture, defence and governance, over-reliance on foreign platforms raises the risks of data exposure and loss of control.
This was highlighted last month when Microsoft abruptly cut off cloud services to Nayara Energy after European Union sanctions on its Russian-linked operations. The disruption, which was reversed only after a court intervention, raised alarms on how foreign tech providers can become geopolitical pressure points.
Around the same time, US President Donald Trump doubled tariffs on Indian imports to 50%, showing how trade and tech are increasingly being used as leverage.

View Full Image
Besides reducing dependence, sovereign AI systems are also important for India’s critical sectors to accurately represent local values, regulatory frameworks and linguistic diversity.
Most global AI models are trained on English-dominant and Western datasets, which make them poorly equipped to handle the realities of India’s multilingual population or the domain-specific complexity of its systems.
This becomes a challenge when it comes to applications such as interpreting Indian legal judgments or accounting for local crop cycles and farming practices in agriculture.
Mohanty says that sovereignty in AI isn’t about isolation, but about who controls the infrastructure and who sets the terms. “Sovereignty is basically about choice and dependencies. The more choice you have, the more sovereignty you have.”
He adds that full-stack independence from chips to models is not feasible for any country, including India. Even global powers such as the US and China balance domestic development with strategic partnerships. “Nobody has complete sovereignty or control or self-sufficiency across the stack, so you either build it yourself or you partner with a trusted ally.”
Mohanty also points out that the Indian government has taken a pragmatic approach by staying agnostic to the foundational elements of its AI stack. This stance is shaped less by ideology and more by constraints such as lack of Indic data, compute capacity and ready-made open-source alternatives built for India.
India’s data lacunae
Despite the momentum behind India’s sovereign AI push, the lack of high-quality training data, particularly in Indian languages, continues to be one of its most fundamental roadblocks. While the country is rich in linguistic diversity, that diversity has not translated into digital data that AI systems can learn from.
Manish Gupta, director of engineering at Google DeepMind India, cited internal assessments that found that 72 of India’s spoken languages, which had over 100,000 speakers, had virtually no digital presence. “Data is the fuel of AI and 72 out of those 125 languages had zero digital data,” he says.
To address this linguistic challenge for Google’s India market, the company launched Project Vaani in collaboration with the Indian Institute of Science (IISc).
This initiative aims to collect voice samples across hundreds of Indian districts. The first phase captured over 14,000 hours of speech data from 80 districts, representing 59 languages, 15 of which previously had no digital datasets. The second phase expanded coverage to 160 districts and future phases aim to reach all 773 districts in India.
“There’s a lot of work that goes into cleaning up the data, because sometimes the quality is not good,” Gupta says, referring to the challenges of transcription and audio consistency.
Google is also developing techniques to integrate these local language capabilities into its large models.
Gupta says that learnings from widely spoken languages such as English and Hindi are helping improve performance in lower-resource languages such as Gujarati and Tamil, largely due to cross-lingual transfer capabilities built into multilingual language models.
The company’s Gemma LLM incorporates Indian language capabilities derived from this body of work. Gemma ties into LLM efforts run by Indian startups through a combination of Google’s technical collaborations, infrastructure guidance and by making its collected datasets publicly available.
According to Gupta, the strategy is driven by both commercial and research imperatives. India is seen as a global testbed for multilingual and low-resource AI development. Supporting local language AI, especially through partnerships with startups such as Sarvam, Soket AI and Gnani.ai, allows Google to build inclusive tools that can scale beyond India to include other linguistically complex regions in Southeast Asia and Africa.
For India’s sovereign AI builders, the lack of readymade and high-quality Indic datasets means that model development and dataset creation must happen in parallel.
For the Global South
India’s layered strategy to use open models now, while concurrently building sovereign models, also offers a roadmap for other countries navigating similar constraints. It’s a blueprint for the Global South, where nations are wrestling with the same dilemma on how to build AI systems that reflect local languages, contexts and values without the luxury of vast compute budgets or mature data ecosystems. For these countries, fine-tuned open models offer a bridge to capability, inclusion, and control.
“Full-stack sovereignty in AI is a marathon, not a sprint,” Upperwal says. “You don’t build a 120 billion model in a vacuum. You get there by deploying fast, learning fast and shifting when ready.”
Full-stack sovereignty in AI is a marathon, not a sprint.
— Abhishek Upperwal
Singapore, Vietnam and Thailand are already exploring similar methods, using Gemma to kickstart their local LLM efforts.
By 2026, when India’s sovereign LLMs, including EKA, are expected to be production-ready, Upperwal says the dual track will likely converge, and bootstrapped models will fade while homegrown systems may take their place.
But even as these startups build on open tools such as Meta’s Llama or Google’s Gemma, which are engineered by global tech giants, the question of dependency continues to loom. Even for open-source models, control over architecture, training techniques and infrastructure support still leans heavily on Big Tech.
While Google has open-sourced speech datasets, including Project Vaani, and extended partnerships with IndiaAI Mission startups, the terms of such openness are not always symmetrical. India’s sovereign plans, therefore, depend not on shunning open models but on eventually outgrowing them.
“If Google is directed by the US government to close down its weights (model parameters), or increase API (application programming interface) prices or change transparency norms, what would the impact be on Sarvam or Soket?” questions Mohanty, adding that while the current India-US tech partnership is strong, future policies could shift and jeopardize India’s digital sovereignty.
In the years ahead, India and other nations in the Global South will face a critical question over whether they can convert this borrowed support into a complete, sovereign AI infrastructure, before the terms of access shift or the window to act closes.
Key Takeaways
- As artificial intelligence systems become central to education, agriculture, defence and governance, over-reliance on foreign platforms raises the risks of data exposure and loss of control.
- With 22 official languages and hundreds of spoken dialects, India faces a monumental challenge in building AI systems that can work across this multilingual landscape.
- For India, building its own AI capabilities is not just a matter of nationalistic pride—it’s more about solving problems that no foreign model can adequately address today.
- So, Indian startups are fine-tuning open-source models to solve real-world problems today.
- Simultaneously, they are building the data pipelines and domain-specific expertise needed to train more indigenous models.
- India’s layered strategy offers a roadmap for other countries navigating similar constraints.
-
Blockchain3 days agoThis Week in Stablecoins: TradFi Wants Blockchain – Crypto News
-
Technology3 days agoAmazon Exec Predicts Commercial Quantum Computers in 2031 – Crypto News
-
Technology1 week agoiOS 27 Public Beta 1 released with Siri AI: How to download and eligible devices – Crypto News
-
Blockchain1 week agoKraken Pro Launches API Partner Program Supporting Specialized Integrations – Crypto News
-
Blockchain1 week agoDogecoin Reclaims $0.073 As Meme Traders Look For A Cleaner Rebound – Crypto News
-
Blockchain1 week agoDogecoin Reclaims $0.073 As Meme Traders Look For A Cleaner Rebound – Crypto News
-
Technology2 days agoTech Layoffs Hit 2-Year High as Companies Embrace AI – Crypto News
-
De-fi1 week agoDTCC Starts Live Tokenized-Securities Trades With More Than Two Dozen Firms – Crypto News
-
others1 week agoThree Russians Accused of Facilitating $62,000,000 Cybercrime Scheme That Targeted Americans – Crypto News
-
Technology1 week agoCabinet approves ₹1.9 tn push for chips, mobiles to deepen local manufacturing – Crypto News
-
others1 week agoBybit Wins Excellence in Innovation and Strategic Leadership Awards at Peru Blockchain Conference 2026 – Crypto News
-
Business1 week ago
Blockchain Association Casts CLARITY Act as Crypto Crime-Fighting Bill – Crypto News
-
Cryptocurrency1 week agoJapan passes the crypto law traders wanted but its 20% tax could still wait until 2028 – Crypto News
-
others1 week agoNumerai Completes Third Strategic NMR Buyback, Bringing Total Repurchases to $3.2 Million – Crypto News
-
Blockchain6 days agoKaspersky Uncovers Malware Framework Targeting Crypto Investors – Crypto News
-
Blockchain6 days agoKaspersky Uncovers Malware Framework Targeting Crypto Investors – Crypto News
-
Technology1 week agoMotorola Edge 70 Max launched in India with Snapdragon 8 Gen 5, 7,100mAh battery: Check price, specs and launch offers – Crypto News
-
Business1 week ago
BlackRock Bitcoin ETF (IBIT) Gets US SEC Approval to Increase Options Contract Limit – Crypto News
-
Business1 week ago
BlackRock Bitcoin ETF (IBIT) Gets US SEC Approval to Increase Options Contract Limit – Crypto News
-
Business1 week ago
BlackRock Bitcoin ETF (IBIT) Gets US SEC Approval to Increase Options Contract Limit – Crypto News
-
Blockchain1 week agoKraken Pro Launches API Partner Program Supporting Specialized Integrations – Crypto News
-
Cryptocurrency1 week agoBitcoin miner CleanSpark signed a $6.6B AI lease before securing the $2.1B required to build it – Crypto News
-
others1 week ago
$10T Morgan Stanley’s E*TRADE Completes Rollout of Spot Bitcoin, Ethereum, Solana Trading – Crypto News
-
Blockchain1 week agoCardano Tests Support As ADA Traders Look For A Better Catalyst – Crypto News
-
Blockchain1 week agoCardano Tests Support As ADA Traders Look For A Better Catalyst – Crypto News
-
Blockchain1 week agoCardano Tests Support As ADA Traders Look For A Better Catalyst – Crypto News
-
Technology1 week ago
Anthropic IPO: Meta Platforms Eyes $10 Bln Deal With Anthropic As AI Race Heats Up – Crypto News
-
Technology1 week ago
Best Leveraged Prediction Market Platforms for Event Traders – Crypto News
-
others1 week agoNumerai Completes Third Strategic NMR Buyback, Bringing Total Repurchases to $3.2 Million – Crypto News
-
others7 days ago
US-Iran Update: Oil Prices Surge By 4% as Iran Escalates Fresh Attacks on Gulf Allies – Crypto News
-
Technology6 days agoIs Zepto down? Users hit by mass outage, face disruptions on app and website – Crypto News
-
Blockchain4 days agoBrazilian Securities Watchdog Takes Closer Look at Tokenization – Crypto News
-
Blockchain4 days agoBrazilian Securities Watchdog Takes Closer Look at Tokenization – Crypto News
-
Business1 week ago
U.S.-Iran War Update: Iran Denies Plans for Peace Talks Despite Trump’s Claims – Crypto News
-
Technology1 week ago
U.S.-Iran War Update: Iran Denies Plans for Peace Talks Despite Trump’s Claims – Crypto News
-
Technology1 week ago
U.S.-Iran War Update: Iran Denies Plans for Peace Talks Despite Trump’s Claims – Crypto News
-
Business1 week ago
Robinhood Chain’s $INDEX 150% Rally Turns Trading Fees Into Real Stock Rewards – Crypto News
-
Cryptocurrency1 week agoBitcoin miner CleanSpark signed a $6.6B AI lease before securing the $2.1B required to build it – Crypto News
-
others1 week agoBybit Wins Excellence in Innovation and Strategic Leadership Awards at Peru Blockchain Conference 2026 – Crypto News
-
Technology1 week agoVivo T5 Lite 44W 5G in India with 120Hz display, 6,500mAh battery: Check price and specs – Crypto News
-
Blockchain1 week agoTrump Aide Faces Scrutiny Over $100K in Kalshi Speech Bets: ABC – Crypto News
-
others1 week ago62,150 Americans Warned After ‘Unauthorized Party’ Breaches Healthcare Firm in Houston – Personal and Health Records Potentially Exposed – Crypto News
-
Cryptocurrency7 days agoXRP’s Price Health Is on the Line, Did Shiba Inu (SHIB) Finally Bottom? Ethereum’s (ETH) Mini-Golden Cross: Crypto Market Review – Crypto News
-
Technology7 days agoMicrosoft trains sales team to talk down OpenAI, Anthropic AI models, pitches itself as ‘full AI platform’: Report – Crypto News
-
De-fi6 days ago‘Coinbase Man’ Token Crashes as Armstrong Swaps His Avatar for a CryptoPunk – Crypto News
-
Blockchain6 days agoUniswap Founder Proposes v4 Protocol Fees Across Multiple Networks – Crypto News
-
Technology5 days agoInstagram, Facebook down? Thousands of users report mass outage; netizens say ‘no posts coming up on feed after refresh’ – Crypto News
-
others1 week ago
XRP Price & Open Interest Jump as Whales Scoop 70M Coins in a Week – Crypto News
-
Blockchain7 days agoChainlink Holds Support As CCIP Adoption Becomes A Longer-Term Test – Crypto News
-
Blockchain7 days agoChainlink Holds Support As CCIP Adoption Becomes A Longer-Term Test – Crypto News
