Your Singapore team is on a call. Someone says:

🇸🇬 Meeting participant
"Eh, so we need to deploy this feature lah, but the database query very slow lor. How we optimize? Boleh ask the backend team?"

That sentence has English, Malay, and Singlish grammar patterns all mixed together. Try to transcribe it with Otter.ai or Google Meet's built-in captions. You'll get something like:

What was said
"Eh, so we need to deploy this feature lah, but the database query very slow lor. How we optimize? Boleh ask the backend team?"
AI heard
"Eh, so we need to deploy this feature la, but the database query very slow or how we optimize..."

The words are garbled. The meaning is lost. And if you're using that transcript as your meeting notes, you've got a mess.

This is code-switching. And it's breaking almost every AI transcription tool on the market.

What Is Code-Switching?

Code-switching is when a bilingual or multilingual speaker mixes two or more languages in a single conversation — often within the same sentence.

It's not broken English. It's not bad grammar. It's actually a sign of linguistic sophistication. Bilingual speakers code-switch because it's the most efficient way to communicate with other bilingual people in their community.

It's normal everywhere in Southeast Asia:

🇸🇬 Singapore Singlish
🇲🇾 Malaysia Bahasa Rojak
🇵🇭 Philippines Taglish
🇹🇭 Thailand Thai-English
🇮🇩 Indonesia Bahasa Campur
🇻🇳 Vietnam Vietglish

Every Southeast Asian professional does this. It's how we communicate. But here's the problem: AI transcription models were trained on monolingual speech. They've never seen this before.

Why AI Breaks on Code-Switching

Modern speech-to-text models work by running audio through a pipeline:

Step 1

Tokenization

Breaking audio into phonemes — individual sounds
Step 2

Language identification

Figuring out which language is being spoken — this is where code-switching breaks everything
Step 3

Pattern matching

Matching sound patterns to known words in the identified language
Step 4

Grammar correction

Using a language model to clean up errors and fill gaps

Code-switching breaks at Step 2. When you say a sentence in Singlish, the AI hears English words and Malay grammar patterns simultaneously. The language identification model gets confused. Is this English or Malay? The model has to pick one. It picks wrong. Everything downstream breaks.

Here's a real example:

What was said
"Eh, cannot lah, server go down already."
AI's guess
English (the root words are English)
AI output
"Eh cannot la server go down already" (strips Singlish particles — they don't fit English grammar)
Actual meaning
"No, we can't do that right now, because the server has crashed."

The meaning was there. The model didn't preserve it.

The Training Data Problem

Why does this happen? Because the datasets used to train these models don't include code-switched speech.

Google, OpenAI, and other major AI labs trained their speech models on enormous pools of audio — but almost all of it is monolingual:

But Singlish? Taglish? Bahasa Campur? There's almost no training data. Why? Because code-switched speech is:

So the models ignore it. And when they encounter it in the wild, they fail.

Why This Matters for SE Asian Teams

Imagine you're a regional PM. You record a standup with your Singapore, Bangkok, and Manila teams. Everyone code-switches naturally — it's how they communicate best. You use Otter.ai to transcribe. You get roughly:

~60%
Accuracy on the pure English parts
~40%
Accuracy on the code-switched parts — the parts that matter most

So you spend 20 minutes manually fixing the transcript. Or you just don't use it. Either way, you've lost the main benefit of transcription: saving time.

⚠

For a team of 10 people meeting 3× a week, bad transcription wastes roughly 500 hours a year. For a 50-person regional team, that's 2,500 hours. That's real money sitting on the floor.

The Current Solutions (And Why They Don't Work)

Option 1 Doesn't scale
Use a tool built for your specific language
Some tools handle Singlish or Taglish specifically. But they only work for one code-switched dialect. If your team spans Singapore and Manila, you're still out of luck.
Option 2 Absurd
Record separate videos in each language
Some teams actually do this — one recording in English, another in the local language. It doesn't reflect how people actually talk, and nobody does it twice.
Option 3 Getting better
Use Google Meet or Zoom's built-in captions
Still 50–60% accurate on code-switched speech. Fine for quick reference during a meeting. Not usable for meeting notes you'd actually rely on.
Option 4 Works, but expensive
Hire someone to manually transcribe
Humans understand code-switching. A good transcriber gets the meaning right even when grammar is mixed. Slow and expensive at scale, but the gold standard for accuracy.
Option 5 Most common
Just don't transcribe
Most regional teams do this. They record meetings but never transcribe because the tools are so bad at code-switching. The recordings sit in a Drive folder, unwatched.

How We're Solving This at BYSIK

When we started building BYSIK, we noticed this problem immediately. Our first user was a Singapore startup with a team across SG, MY, and ID. They said: "Every transcription tool fails on our meetings because we speak Singlish. We just stopped using transcription."

That's when we realized: the existing solutions aren't built for Southeast Asia. So we did something different.

1 — Training Data
We trained on code-switched speech

We built a dataset of actual code-switched audio from SE Asian professionals — Singlish, Taglish, Bahasa Campur, mixed Thai-English, all of it. Carefully labeled and used to fine-tune our speech-to-text model on this specific data.

2 — Architecture
We use language-agnostic embedding models

Instead of deciding "Is this English or Malay?" upfront, we use embedding models that represent words in a shared semantic space. "Lah" and "already" both carry meaning in context — the model learns this relationship instead of forcing a language classification.

3 — Variation
We handle dialect and accent variation

Singlish from Singapore sounds different from Singlish in Malaysia. Bangkok Thai-English is different from Northern Thai-English. We built models that handle these regional variations rather than treating each dialect as the same input.

4 — Fidelity
We preserve the original speech patterns

Instead of "correcting" code-switched speech into monolingual grammar, we keep it as spoken. If you said "cannot lah," the transcript says "cannot lah." The meaning — and the voice — is preserved.

The result:

Other tools
40–50%
accuracy on code-switched speech
BYSIK AI
85%+
accuracy on code-switched speech
ℹ

Is it perfect? No. Code-switching is inherently ambiguous — even humans sometimes disagree on what was said. But 85% is accurate enough to be genuinely useful for meeting notes, which is the whole point.

The Bigger Picture: Why SE Asia Keeps Getting Left Behind

This code-switching problem is a microcosm of a bigger issue. AI is built in the US and China. The training data is English, Mandarin, and a few other "high-resource" languages. Everything else — including all of Southeast Asia's languages and dialects — gets treated as an edge case.

The tools work great for monolingual English speakers in San Francisco. They're fine for Mandarin speakers in Shanghai. But for a multilingual team in Singapore? For developers in Manila who code-switch naturally?

The tools fail.

And the assumption is: "That's an edge case. Most of the world speaks monolingual English or Mandarin anyway."

⚠

Southeast Asia is 650 million people. It's not an edge case. It's a massive market that's been ignored because building for multilingual, code-switched speech is harder than building for monolingual English.

What Needs to Change

For researchers
Start collecting and publishing datasets of code-switched speech. It's harder than monolingual data — label each language segment, account for regional variation, handle ambiguous cases — but it's important. Southeast Asia's languages matter.
For AI companies
Stop treating code-switching as an edge case. Train your models on it. Users in SE Asia deserve tools that work for how they actually talk, not how you think they should talk.
For regional companies
If existing tools don't work for you, you don't have to accept it. Build your own, or support tools built for your market. Demand better from the vendors you're already paying.
For SE Asian founders
This is an opportunity. The entire region is using transcription tools that don't work for how we actually speak. That's a real problem worth solving — and the local market context is an advantage no SF startup can replicate.

The Practical Takeaway

If you're running a team in Southeast Asia and you've been frustrated with transcription accuracy, now you know why.

It's not your audio quality. It's not your accent. It's not that you're speaking "wrong."

It's that the tools were built for a different market. They were trained on monolingual speech. Code-switching breaks them.

✓

Tools are getting better at handling multilingual and code-switched speech. If you've given up on transcription because it didn't work, it might be worth trying again. And if you find a tool that actually understands how your team talks — stick with it. You've found something rare.

Built for how Southeast Asia actually speaks.

BYSIK AI was trained on real code-switched speech from SE Asian teams — Singlish, Taglish, Bahasa Campur, and more. Try it on your next meeting.

Try BYSIK AI Free →
P.S. — If you want to geek out about the linguistics of code-switching, there's a whole field of research on it. Start here: Poplack's "The Bilingual's Linguistic System: Evidence for Asymmetric Competence." It's genuinely fascinating stuff.

And if you're building tools for Southeast Asia, feel free to reach out — support@bysik.app. I'm always interested in talking to founders who are solving regional problems instead of just copying the US.
Full disclosure: I founded BYSIK AI because of this exact problem. We're solving code-switching for transcription in Southeast Asia. But even if you use a different tool, I hope this helped you understand why transcription has been hard in SE Asia — and why it's getting better.