Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

English-dominant AI models leave Cantonese and other languages lagging behind

Top AI systems excel in English but struggle with languages like Cantonese, risking poorer translations and data contamination.

Current leading AI models perform best in English because roughly half of online content is in that language, providing abundant training material. Languages with fewer digital resources—including Cantonese, which is spoken by 85 million people—suffer from scarce, often noisy data, leading to mistranslations and nonsensical outputs. Start-ups like Votee and larger firms such as Naver and the Indosat-Goto partnership are attempting to create regional models, yet they lack the billions of parameters seen in GPT-4 or DeepSeek’s R1.

Researchers caution that relying on AI-generated translations to augment datasets can embed cultural biases and degrade language quality over time, a phenomenon dubbed “model collapse.” The issue highlights a broader need for more investment in multilingual AI to ensure equitable access to digital tools worldwide.

Get the beta ↗