Search

English / Technology

Who Has Built a Language Model of Their Own in Southeast Asia?

Who Has Built a Language Model of Their Own in Southeast Asia?
Photo by cottonbro studio: https://www.pexels.com/photo/portrait-shot-of-a-woman-5473960/

Eight of the bloc's eleven member states now have at least one homegrown AI language model. A look at what has been built, who is funding it, and why Brunei, Laos, and Timor-Leste are still waiting.

Every large AI model carries the assumptions of the data it was trained on, and most of the data behind ChatGPT, Gemini, and Claude comes from English-language text scraped from the global internet. For a language like Filipino, Khmer, or Burmese, that means a chatbot answering in something close to the right words but missing the idiom, the local reference, or the cultural register that a native speaker would expect. It also means that when a government agency, a bank, or a hospital wants to deploy AI at scale, it is often adapting a foreign-built system rather than one shaped by its own institutions and its own language data.

That gap is why a growing number of ASEAN governments, universities, telecom operators, and private conglomerates have started building their own models over the past three years. As of September 2026, eight of the bloc's eleven member states have produced at least one homegrown AI language model or are actively partnering to build one, according to a review of company announcements, government statements, and academic papers. The other three, Brunei, Laos, and Timor-Leste, have no major model of their own yet and rely on regional projects such as Singapore's SEA-LION to fill the gap.

Why build a national model at all

The case that governments and companies make for a homegrown model is rarely about beating the biggest labs at their own game. It is about three narrower things: language coverage, data sovereignty, and control over infrastructure.

Language coverage is the most visible argument. Southeast Asia is home to hundreds of languages and dialects, most of which are what AI researchers call "low-resource," meaning there simply is not much text available online to train on. Thai, Vietnamese, Bahasa Indonesia, and Filipino all fall into this category relative to English or Mandarin, and smaller languages like Khmer or Burmese are further behind still. AI Singapore, which built the regional SEA-LION family of models, has said its explicit goal is correcting for the "significant underrepresentation of Southeast Asian language data" in mainstream AI training sets.

Data sovereignty is the second and increasingly the more political argument. Malaysia's Gamuda Technologies markets its Wira model on the basis that it is air-gapped and does not need any external network connection to run, a feature the company's leadership frames as protection against "foreign disruption" during a geopolitical crisis. Indonesia's Sahabat-AI, developed by Indosat Ooredoo Hutchison and GoTo with support from Nvidia and AI Singapore, was likewise framed at launch around "digital sovereignty" rather than raw performance benchmarks.

The third argument, infrastructure control, mostly applies to the handful of projects with serious capital behind them. YTL Power's Ilmu, for instance, is tied to a broader build-out that includes a 500-megawatt green data centre and an Nvidia-powered supercomputer in Malaysia, positioning the model as one output of a much larger sovereign compute strategy rather than a standalone research product.

Homegrown AI models and initiatives by ASEAN country

Counts include shipped models, published research models, and formally announced development partnerships, as of September 2026.

View the underlying data as a table
Country Models / initiatives
Indonesia 4
Malaysia 3
Myanmar 3
Vietnam 3
Singapore 2
Thailand 2
Cambodia 1
Philippines 1
Brunei 0
Laos 0
Timor-Leste 0
Source: company and government announcements, academic publications, 2020–2026

Indonesia has the most, but not all of it is a product

Indonesia's four entries are not all comparable. Sahabat-AI, launched in November 2024, is the most commercially visible: an open-source ecosystem built by Indosat and GoTo with contributions from four universities and support from Nvidia's NeMo platform, released initially as 8-billion and 9-billion parameter models tuned for Bahasa Indonesia and regional languages. Cendol, published by a research team including Indonesian and Southeast Asian NLP researchers at ACL 2024, is an instruction-tuned open model built specifically for Indonesian and its local languages rather than a commercial product. IndoBERT and IndoBART are older still, dating to 2020 academic work that produced the benchmark encoder and sequence-to-sequence models much of the country's later Indonesian-language NLP research has built on.

Malaysia, Myanmar, and Vietnam each have three. Malaysia's spread illustrates how differently these projects can be funded: Ilmu came from YTL Power, a listed utilities and infrastructure company, and was unveiled in August 2025 as what its developers describe as the top-scoring model for Malay on the MMLU benchmark; Wira came a year later from Gamuda Technologies and is aimed squarely at government and enterprise deployments rather than public chat use; MaLLaM, by contrast, grew out of a more community and academic effort to build an open Malaysian-language model. Myanmar's three, all released by a single independent developer, Min Si Thu, between December 2023 and early 2024, range from a lightweight 128-million-parameter model to a 1.42-billion-parameter version covering 61 languages, built without the corporate or state backing seen elsewhere in the region. Vietnam's three include PhoGPT from VinAI, unveiled at AI Day 2023 with a 7.5-billion-parameter instruct variant; VinaLLaMA, an independently developed model built on Meta's LLaMA-2 architecture with contributions from Nous Research, LAION, Google Cloud, and StabilityAI; and ViGPT from VinBigData, aimed at commercial chatbot deployment.

Singapore builds for the region, not just for itself

Singapore's two projects stand apart because neither was built to serve Singapore alone. SEA-LION, developed by AI Singapore under the government-backed National Multimodal LLM Project, was explicitly designed to cover multiple Southeast Asian languages at once, including Thai, Vietnamese, and Bahasa Indonesia, and by mid-2025 had reportedly gained adoption from regional firms including Indonesia's GoTo Group. MERaLiON, developed by A*STAR's Institute for Infocomm Research with backing from the Infocomm Media Development Authority, launched in December 2024 and was expanded in May 2025 to handle Malay, Tamil, Thai, Bahasa Indonesia, and Vietnamese alongside English, Mandarin, and Singlish, with added code-switching and emotion-recognition capabilities aimed at customer service and social-service applications.

That regional design is also why SEA-LION keeps showing up in projects outside Singapore's borders. Cambodia's KhmerLLM, formalized through a January 2025 memorandum of understanding between AI Singapore and AI Forum Cambodia, is being built as part of the SEA-LION family rather than as a fully separate national effort, with the initial phase focused on collecting and cleaning Khmer-language data before any advanced instruction-tuned version is attempted. It is the clearest example in the region of a country without the capital or research base to build alone choosing to plug into a neighbor's infrastructure instead.

Thailand and the Philippines: one flagship, one research paper

Thailand's two entries, OpenThaiGPT and Typhoon, both address the same underlying problem from different directions. Typhoon, built by SCB 10X, the venture and innovation arm of Thai banking group SCBX, is the more actively maintained of the two, with open model weights, a research roadmap that includes audio and vision variants, and an explicit goal of establishing Thailand as a serious Southeast Asian AI research base within three years. The developers describe Thai as a "low-resource" language in global AI terms, arguing that most existing models can process Thai text but lack any real grasp of Thai cultural context.

The Philippines, by contrast, has one entry so far, and it is a research model rather than a shipped product. FiLLM, described in a 2025 paper, is built on top of the regional SEALLM-7B model using a fine-tuning technique called LoRA, and is aimed at practical Filipino-language tasks such as named entity recognition, part-of-speech tagging, and text summarization rather than general-purpose chat. The country's other widely cited AI initiative, Pilipinas AI, is worth separating out here: it is an AI infrastructure and coordination platform, not a foundation model, so it is not counted among the eleven countries' model tallies above.

The three still waiting

Brunei, Laos, and Timor-Leste have no major homegrown foundation model publicly documented as of this writing. All three are among the smaller economies in the bloc, without the venture capital, telecom-operator scale, or university AI research infrastructure that funded most of the projects above. For now, users and institutions in these countries who want a model with any grasp of local language and context are left with regional options such as SEA-LION, or with the same global chatbots that started this whole conversation about representation in the first place.

Cambodia's experience with KhmerLLM suggests one route for the three without a project of their own: rather than raising the capital for an independent model, a smaller country can enter a foundation-model family like SEA-LION as a data and language partner, gaining a stake in a working system without carrying the full cost of building one from scratch.

What "homegrown" is doing a lot of work in this tally

It is worth being precise about what these nineteen entries actually are, because "homegrown AI model" covers a wide range of maturity. Some, like Sahabat-AI and Ilmu, are backed by billions of dollars in supporting infrastructure and are being pitched to government agencies as production systems. Others, like MyanmarGPT or FiLLM, are the work of one developer or one research team, openly published but not deployed at any commercial scale. A count of models by country says something about where activity is happening, not about which systems people are actually using day to day, and the two rankings would not necessarily match.

Country Model(s) Developer Status, as of Sept. 2026
Brunei None documented
Cambodia KhmerLLM AI Singapore, AI Forum Cambodia Partnership formalized Jan. 2025, early data-collection stage
Indonesia Sahabat-AI Indosat Ooredoo Hutchison, GoTo Launched Nov. 2024, 8B/9B parameters
Cendol Academic consortium (ACL 2024) Open research model
IndoBERT / IndoBART Academic (2020) Benchmark models, still widely used
Laos None documented
Malaysia Ilmu YTL AI Labs Launched Aug. 2025
Wira Gamuda Technologies Launched May 2026, government/enterprise focus
MaLLaM Community/academic effort Open research model
Myanmar MyanmarGPT family Min Si Thu (independent) Dec. 2023–early 2024, 128M–1.42B parameters
Philippines FiLLM Academic research (2025 paper) Research model, built on SEALLM-7B
Singapore SEA-LION AI Singapore Regional family model, adopted beyond Singapore
MERaLiON A*STAR I2R, IMDA Launched Dec. 2024, expanded May 2025
Thailand Typhoon SCB 10X Active open-source development
OpenThaiGPT Thai open-source community Open research model
Timor-Leste None documented
Vietnam PhoGPT VinAI (Vingroup) Launched Dec. 2023, 7.5B parameters
VinaLLaMA Independent researchers, with Nous Research, LAION, Google Cloud, StabilityAI Released Dec. 2023, 2.7B/7B parameters
ViGPT VinBigData Commercial chatbot deployment

What connects nearly every project on this list, whether it comes from a telecom operator, a bank's venture arm, a state research institute, or a single independent developer, is that it was built because a general-purpose model built elsewhere did not treat the local language as a priority. Whether that translates into models people actually use, rather than models that exist, is the question the region's next few years of AI deployment will answer.


Sources:

GoTo Company press release on Sahabat-AI (Nov. 2024); AI Singapore, SEA-LION documentation and model cards; Cahyawijaya et al., "Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages," ACL 2024; Koto et al., "IndoLEM and IndoBERT," COLING 2020; YTL AI Labs and The Edge Malaysia on ILMU (Aug. 2025); Gamuda Technologies on Wira (May 2026); MinSiThu, MyanmarGPT project documentation (GitHub, Hugging Face); arXiv 2505.18995, "FiLLM, a Filipino-optimized Large Language Model"; IMDA and A*STAR I2R on MERaLiON (Dec. 2024, May 2025 update); SCB 10X and OpenTyphoon.ai; VinAI on PhoGPT (Dec. 2023); Nguyen, Pham and Dao, "VinaLLaMA: LLaMA-based Vietnamese Foundation Model," arXiv 2312.11011 (Dec. 2023); The Better Cambodia and Kiripost on the AI Singapore and AI Forum Cambodia KhmerLLM partnership (Jan. 2025).

Tags:

Thank you for reading until here