English-dominant AI models leave Cantonese and other languages lagging behind
Top AI systems excel in English but struggle with languages like Cantonese, risking poorer translations and data contamination.
Current leading AI models perform best in English because roughly half of online content is in that language, providing abundant training material. Languages with fewer digital resources—including Cantonese, which is spoken by 85 million people—suffer from scarce, often noisy data, leading to mistranslations and nonsensical outputs. Start-ups like Votee and larger firms such as Naver and the Indosat-Goto partnership are attempting to create regional models, yet they lack the billions of parameters seen in GPT-4 or DeepSeek’s R1.
Researchers caution that relying on AI-generated translations to augment datasets can embed cultural biases and degrade language quality over time, a phenomenon dubbed “model collapse.” The issue highlights a broader need for more investment in multilingual AI to ensure equitable access to digital tools worldwide.
