Your Singapore team is on a call. Someone says:
That sentence has English, Malay, and Singlish grammar patterns all mixed together. Try to transcribe it with Otter.ai or Google Meet's built-in captions. You'll get something like:
The words are garbled. The meaning is lost. And if you're using that transcript as your meeting notes, you've got a mess.
This is code-switching. And it's breaking almost every AI transcription tool on the market.
What Is Code-Switching?
Code-switching is when a bilingual or multilingual speaker mixes two or more languages in a single conversation — often within the same sentence.
It's not broken English. It's not bad grammar. It's actually a sign of linguistic sophistication. Bilingual speakers code-switch because it's the most efficient way to communicate with other bilingual people in their community.
It's normal everywhere in Southeast Asia:
Every Southeast Asian professional does this. It's how we communicate. But here's the problem: AI transcription models were trained on monolingual speech. They've never seen this before.
Why AI Breaks on Code-Switching
Modern speech-to-text models work by running audio through a pipeline:
Tokenization
Breaking audio into phonemes — individual soundsLanguage identification
Figuring out which language is being spoken — this is where code-switching breaks everythingPattern matching
Matching sound patterns to known words in the identified languageGrammar correction
Using a language model to clean up errors and fill gapsCode-switching breaks at Step 2. When you say a sentence in Singlish, the AI hears English words and Malay grammar patterns simultaneously. The language identification model gets confused. Is this English or Malay? The model has to pick one. It picks wrong. Everything downstream breaks.
Here's a real example:
The meaning was there. The model didn't preserve it.
The Training Data Problem
Why does this happen? Because the datasets used to train these models don't include code-switched speech.
Google, OpenAI, and other major AI labs trained their speech models on enormous pools of audio — but almost all of it is monolingual:
- English audio (billions of hours)
- Mandarin audio (billions of hours)
- Spanish, French, German, etc. (hundreds of millions of hours each)
But Singlish? Taglish? Bahasa Campur? There's almost no training data. Why? Because code-switched speech is:
- Hard to label — Is this English or Malay? Both? The labeler has to make a judgment call
- Not standardized — Singlish spoken in Singapore sounds different from Singlish in Malaysia
- Seen as "low prestige" — Academic datasets focus on formal, monolingual speech
- Computationally expensive — Mixed-language models are harder and more costly to train
So the models ignore it. And when they encounter it in the wild, they fail.
Why This Matters for SE Asian Teams
Imagine you're a regional PM. You record a standup with your Singapore, Bangkok, and Manila teams. Everyone code-switches naturally — it's how they communicate best. You use Otter.ai to transcribe. You get roughly:
So you spend 20 minutes manually fixing the transcript. Or you just don't use it. Either way, you've lost the main benefit of transcription: saving time.
For a team of 10 people meeting 3× a week, bad transcription wastes roughly 500 hours a year. For a 50-person regional team, that's 2,500 hours. That's real money sitting on the floor.
The Current Solutions (And Why They Don't Work)
How We're Solving This at BYSIK
When we started building BYSIK, we noticed this problem immediately. Our first user was a Singapore startup with a team across SG, MY, and ID. They said: "Every transcription tool fails on our meetings because we speak Singlish. We just stopped using transcription."
That's when we realized: the existing solutions aren't built for Southeast Asia. So we did something different.
We built a dataset of actual code-switched audio from SE Asian professionals — Singlish, Taglish, Bahasa Campur, mixed Thai-English, all of it. Carefully labeled and used to fine-tune our speech-to-text model on this specific data.
Instead of deciding "Is this English or Malay?" upfront, we use embedding models that represent words in a shared semantic space. "Lah" and "already" both carry meaning in context — the model learns this relationship instead of forcing a language classification.
Singlish from Singapore sounds different from Singlish in Malaysia. Bangkok Thai-English is different from Northern Thai-English. We built models that handle these regional variations rather than treating each dialect as the same input.
Instead of "correcting" code-switched speech into monolingual grammar, we keep it as spoken. If you said "cannot lah," the transcript says "cannot lah." The meaning — and the voice — is preserved.
The result:
Is it perfect? No. Code-switching is inherently ambiguous — even humans sometimes disagree on what was said. But 85% is accurate enough to be genuinely useful for meeting notes, which is the whole point.
The Bigger Picture: Why SE Asia Keeps Getting Left Behind
This code-switching problem is a microcosm of a bigger issue. AI is built in the US and China. The training data is English, Mandarin, and a few other "high-resource" languages. Everything else — including all of Southeast Asia's languages and dialects — gets treated as an edge case.
The tools work great for monolingual English speakers in San Francisco. They're fine for Mandarin speakers in Shanghai. But for a multilingual team in Singapore? For developers in Manila who code-switch naturally?
The tools fail.
And the assumption is: "That's an edge case. Most of the world speaks monolingual English or Mandarin anyway."
Southeast Asia is 650 million people. It's not an edge case. It's a massive market that's been ignored because building for multilingual, code-switched speech is harder than building for monolingual English.
What Needs to Change
The Practical Takeaway
If you're running a team in Southeast Asia and you've been frustrated with transcription accuracy, now you know why.
It's not your audio quality. It's not your accent. It's not that you're speaking "wrong."
It's that the tools were built for a different market. They were trained on monolingual speech. Code-switching breaks them.
Tools are getting better at handling multilingual and code-switched speech. If you've given up on transcription because it didn't work, it might be worth trying again. And if you find a tool that actually understands how your team talks — stick with it. You've found something rare.
Built for how Southeast Asia actually speaks.
BYSIK AI was trained on real code-switched speech from SE Asian teams — Singlish, Taglish, Bahasa Campur, and more. Try it on your next meeting.
Try BYSIK AI Free →And if you're building tools for Southeast Asia, feel free to reach out — support@bysik.app. I'm always interested in talking to founders who are solving regional problems instead of just copying the US.