$ 44.76 € 50.99 zł 11.66
+21° Kyiv +16° Warsaw +16° Washington

Taiwan expands its own AI training database — Taipei Times

Lev Shevtsov 24 September 2026 01:37
Taiwan expands its own AI training database — Taipei Times

Taiwan is developing its own database for training artificial intelligence systems so that they more accurately understand the country’s language, culture and social values. On September 15, Taiwan’s Ministry of Digital Affairs invited publishers, writers and e-book platforms to join the creation of Taiwan’s Sovereign AI Training Corpus. The database currently contains about 5,000 datasets and more than 2.2 billion text tokens, Taipei Times reports.

Minister of Digital Affairs Lin Yi-ching said that a language model’s understanding of democracy, checks and balances and other concepts is shaped by the data on which it was trained. According to her, local language, culture and values must be included in training datasets for a more accurate understanding of Taiwan.

Three components of sovereign AI

The article notes that countries are increasingly reconsidering their dependence on American and Chinese technology companies. Taiwan’s concept of sovereign AI includes three areas: knowledge sovereignty, under which a system correctly understands Taiwan; data sovereignty, intended to keep sensitive information under local control; and technological sovereignty, designed to reduce dependence on foreign model providers.

More current news is available on the UA.News Telegram channel Telegram.

In 2023, Taiwan launched the government-backed Trustworthy AI Dialogue Engine, or TAIDE. The project was created to work better with traditional Chinese and Taiwanese knowledge. TAIDE head Lee Yu-chieh noted at the time that the multilingual BLOOM model had been trained on 16% simplified Chinese data, while the share of traditional Chinese data was only 0.05%.

Localization improves accuracy

TAIDE adapts open models whose parameters can be downloaded and modified. Early versions used Meta’s Llama, while the current Gemma-3-TAIDE-12B is based on Google’s Gemma 3 12B. Taiwan’s TMMLU+ test, which contains more than 20,000 questions across 66 subjects, showed that TAIDE’s February model gave 58% correct answers compared with 54% for the original Gemma. In questions on Taiwan’s geography, the result was 70% versus 61%.

At the same time, the material stresses that local training does not eliminate the gap with the most powerful foreign models. The ministry’s corpus can also be used for RAG technology, under which a system searches for information in a verified document database before responding. Chang Yung-chun, deputy director of the AI Research Center in Medicine at Taipei Medical University, believes that a stronger reasoning model supported by verified Taiwanese data can outperform a weaker locally adapted model in complex tasks.

Read us on
Download our app