Have you ever had that moment where you finish something, then realize you never needed to build it
That was me that day — serious work in the morning writing 37 Chinese pronunciation rules, all dead code by afternoon
If you are an engineer, this is a real case of LLMs flipping old rule-based systems If you are a PM or founder building Taiwan-Chinese products, this tells you something practical: your team might cut half of TTS maintenance cost and still improve audio quality
How it started
I have a Chinese reading product where students listen to AI read textbook content aloud
One day PM messaged me: "It pronounced garbage as lā jī, that's mainland pronunciation. We need Taiwan pronunciation lè sè"
Wrong pronunciation in a Chinese learning product is an education incident, not a "let's observe a bit more" bug
I opened TTS and listened once, yes, mainland accent. That is where this story starts
Four PM requirements
Before coding, align requirements. PM's list was clean:
- Taiwan accent (for example, "attack" should be
gōng jí, notgōng jī) - Natural, not robotic
- Polyphonic characters must be correct (in "cheer", the character should be pronounced
hè, nothē) - Names and places must be correct (Tai Tzu-ying, Chen Yen-po, week 214, year 2021)
There are more than three usable TTS providers, but we tested these first:
| Provider | Taiwan Accent | Cost | Quality |
|---|---|---|---|
| Azure Speech (zh-TW HsiaoChen) | ✅ Native | Expensive | Good, slightly rigid |
| Google Chirp3-HD | ❌ Mainland accent | Medium | Very good |
| Gemini Flash TTS (preview) | 🟡 Prompt-controllable | Cheap | Good |
We tried all three, and each had tradeoffs
Version 1: Azure + SSML <phoneme>
Our earliest production version used Azure plus SSML <phoneme> tags to manually fix polyphonic characters:
<phoneme alphabet="x-microsoft-zhuyin" ph="ㄏㄜˋ">he</phoneme>cai
One look at this XML tells the story: each polyphonic character needs manual zhuyin tagging
Azure's upside is native Taiwan pronunciation. Downsides:
- Expensive, bill keeps climbing every month
- Every new polyphonic case needs manual engineering rules
- Slightly rigid voice quality, less human
We ran this for a while, and I kept thinking there had to be a cheaper and more natural way
Version 2: Gemini + 37-rule homophone replacement table
When Google released Gemini Flash TTS preview, it caught my attention
Ultra-low cost (over two thousand sentences for only about $0.30), great quality, except
Default pronunciation leans mainland
So what did I do. I came up with a method that now feels pretty wild in hindsight
TTS cares about pronunciation more than literal meaning, so what if I swap characters with Taiwan-pronounced homophones
Take the word "attack" as an example:
- Mainland accent: final character read as
jī(tone 1) -> whole word becomesgōng jī - Taiwan accent: final character should be
jí(tone 2) -> whole word should begōng jí - My hack: find another character pronounced
jíin Taiwan speech - Rewrite input and feed TTS, betting it reads by phonetics so output sounds like correct
gōng jí
And this is not an edge case. Taiwan vs mainland has many same-character different-pronunciation words:
| Word | Taiwan | Mainland | Difference |
|---|---|---|---|
| garbage | lè sè | lā jī | both characters differ |
| attack | gōng jí | gōng jī | final syllable tone 2 vs tone 1 |
| enterprise | qì yè | qǐ yè | first syllable tone 4 vs tone 3 |
| research | yán jiù | yán jiū | final syllable tone 4 vs tone 1 |
| danger | wéi xiǎn | wēi xiǎn | first syllable tone 2 vs tone 1 |
| smile | wéi xiào | wēi xiào | first syllable tone 2 vs tone 1 |
| expectation | qí dài | qī dài | first syllable tone 2 vs tone 1 |
| quality | zhí liàng | zhì liàng | first syllable tone 2 vs tone 4 |
| France | fà guó | fǎ guó | first syllable tone 4 vs tone 3 |
| cheer | hè cǎi | hē cǎi | first syllable tone 4 (shout) vs tone 1 (drink) |
| recognize | rèn shì | rèn shi | final syllable full tone vs neutral tone |
Same traditional Chinese surface form, noticeably different pronunciation. This has always been a pain point for Chinese TTS in Taiwan products
The references are public and official:
- MOE Revised Mandarin Dictionary — official authority with Taiwan zhuyin per character
- moedict.tw — open-source site version, fastest for lookup
- Wikipedia: Taiwanese Mandarin and Standard Mandarin in ROC — systematic cross-strait differences
The problem is: we had the lists, but could not modify the model
Rule-style TTS like Azure stays rigid; high-quality Chirp3-HD stays mainland-accented; training a Taiwan-accent model ourselves was beyond budget and data
That left two paths:
- Character-swap workaround: trick TTS with homophone replacement — this was Version B
- Prompt-level control: use a model that understands natural-language instructions — later we found Gemini 3.1 preview fits this
The rest of this post is about how these two paths played out, and why path #1 looked smart but still broke other things
The hidden assumption behind my swap strategy: TTS reads text as phonetic symbols and does not care whether the word is real
_TAIWAN_TTS_REPLACEMENTS = [
("garbage", "music-color"), # TW le se vs CN la ji
("research", "study-old"), # final syllable: TW jiu vs CN jiu(alt tone)
("danger", "surround-risk"), # initial syllable: TW wei vs CN wei(alt tone)
("attack", "attack-urgent"), # final syllable tone fix requested by PM
# ... 37 rules
]
Every line followed the same pattern: find a Taiwan-pronounced homophone and force it in, using TTS like a phonetic machine
I also added Arabic-number-to-Chinese conversion: 214 -> two hundred fourteen, so Gemini would not read it in English
Built it, ran batch, regenerated 2417 audio files, pushed to staging for PM acceptance
Then I hit a pitfall
Pitfall: one function was never called
PM came back: "'attack' still sounds mainland"
What
I checked logs and found the issue: during batch execution, the code only removed punctuation and never applied my 37 replacement rules
Gemini received original text, so the whole table did nothing and the entire batch run was wasted
One-line fix, rerun batch, another $0.30 burned, then finally correct
That round alone cost half a day
To trust the output, I built an audit system
At that point PM raised a requirement that later felt very insightful:
"I need to validate every single audio sentence"
Reasonable concern — rules might misfire, Gemini might misread, I might push an unchecked version
So I built:
- Append-only JSONL log (29 fields per record): original sentence sent to Gemini, replaced text, triggered rules, audio SHA, storage path, generation time, generator identity, and which old version got replaced
- Back-office audit page: PM can browse 2000+ lines in a browser, filter by "replacement applied", "contains numbers", "Taiwan-specific terms", inspect full metadata, and play audio directly
Later I realized the biggest value is not debugging, it is trust
Engineers can live with rough logs for debugging, but product teams accepting AI-generated output need a sense of accountability and traceability
I had underestimated that layer
Version 3: one prompt was enough
At noon I sat down and stared at that 37-rule table
Coincidentally, another engineer in the same industry wrote a post (the Gemini 3.1 TTS hands-on from evanlin.com) showing decent results by directly prompting Taiwan accent and Taiwan wording
I listened again and noticed something
Gemini default was not pure mainland accent. It was a mixed accent with some Taiwan friendliness
Only certain words (garbage, attack, research) leaned mainland. Overall it felt "mostly right, locally wrong"
So what if I only said "please read in Taiwan accent". Would that also fix those local wrong spots
I did not know, but a batch run was only 20 minutes + $0.30. Cost was too low not to test
So I built a variant system:
- Variant A: no replacements, only prompt prefix
- Variant B: replacement-table version
prompt prefix = "Please read the following in traditional Chinese used in Taiwan, with a warm and natural tone:"
Ran Variant A batch: 2417 sentences / ~20 minutes / ~$0.30
A/B blind listening: 44 samples, 6 categories, PM made the call
To turn "feels better" into a verifiable conclusion, I built an A/B blind-listening page
44 samples covering:
- Taiwan pronunciation high-frequency words (garbage / research / danger / attack / enterprise / score)
- Other Taiwan pronunciation cases (smile / as much as possible / rest / quality / recognize / knowledge)
- Personal names with hard characters (Chen Yen-po / Yang Chun-han / Tai Tzu-ying / Chen Yu-fei)
- Places (Tokyo / Taiwan)
- Numbers (214 / 2021 / 100)
- Chinese-English mixed reading (NASA / Hemsworth / PEACE / Frankenstein)
Left in green was A (prompt-only), right in orange was B (replacement-table), each pair shown with side-by-side audio elements
I should also explain who listened
Our team has two high-school engineering interns. Both can independently write React, use git, and handle PR review. Their coding ability is honestly on par with many fresh graduates entering software jobs. They are also very willing to use AI as an amplifier, not trapped by old "handcrafted code only" beliefs
For this project, they were unexpectedly well matched — high-schoolers just came out of years of short-video and YouTube consumption, and their sensitivity to mainland vs Taiwan accent differences is better than mine. For me, checking whether "attack" was tone 1 or tone 2 took multiple replays. For them, one play was enough. I have been in a tech echo chamber too long, and my accent perception got dulled
They also helped me tune prompt wording — phrases like "warm tone" and "natural narration" were iterated together. The final prompt line for A included their contributions
So this blind test was not only me. It was me + two interns + PM, and the result was consistent:
- A sounded more natural: smoother sentence flow, warmer tone
- B was slightly stiff: replacement characters disturbed semantic rhythm and pause patterns
- Hard names: A surprisingly pronounced uncommon names correctly via prompt alone
- Numbers: A made Gemini read
week 214as "week two hundred fourteen" directly, so my number-conversion rules became unnecessary
Decision: use A from now on
One detail that stayed in my head: why did "attack-urgent" sound weird
After blind listening, one intern told me: "In version B, the rewritten word sounds uncomfortable, not sure why"
It took me a while to understand
Recall why we did this replacement — mainland accent reads one syllable as tone 1, so we swapped with a Taiwan tone-2 homophone to force pronunciation
That strategy assumes: TTS only reads phonetics and ignores lexical meaning
But Gemini did not play by that rule
The rewritten token is not a real word
Old-era TTS (like Azure) roughly follows text -> phoneme -> acoustic features -> waveform pipelines (Tacotron / FastSpeech style, explained in Microsoft neural TTS survey). So for it, two strings can be mostly just two phonetic sequences — homophone substitution can work
But new LLM-based TTS (VALL-E, NaturalSpeech 3, Gemini TTS) reframes TTS as a language-model task. In Google's official blog, Gemini TTS "not only knows what to say, but how to say it," deciding delivery based on transcript context
So the point that semantic context affects prosody is backed by official docs and papers, not just my guess
As for the more specific claim "fake words damage sentence-level prosody" — that is my reasonable architecture-level inference, but I have not seen a direct benchmark paper yet. Happy to be corrected
In practice, though, the observation was clear: feed that fake rewritten word into Gemini, and prosody gets strange. The syllable is right, the breath is off. If this inference is right, the implication is simple — we fixed pronunciation by interfering with what the model is actually good at
| Old TTS (Tacotron / FastSpeech family) | LLM TTS (VALL-E / Gemini family) |
|---|---|
| String -> phoneme -> acoustic features -> waveform | String -> context-aware interpretation -> prosody -> waveform |
| OOV often misread with G2P fallback | Fake words do not crash, but context prosody drifts |
| Better with richer rules/dictionaries | Better with more precise prompts |
We also noticed one subtle detail — mainland-style "mo" can carry a slight curled sound, while Taiwan style is cleaner and flatter
Variant A read it in the Taiwan way. Variant B lost that subtle friendliness after replacement
This is exactly the kind of "mouthfeel" engineers rarely discuss, but users feel immediately
Deleting one day's work
After merging the decision PR, I sat there looking at code I had written that morning
_TAIWAN_TTS_REPLACEMENTStable -> dead code_apply_taiwan_pronunciation-> dead code_numbers_to_chinese_tw-> dead code_clean_for_gemini-> dead code- Polyphonic-audit docs built for Variant B -> PR closed, not merged
- 5 research issues (polyphonic chars / mixed language / pausing / names / places) -> all closed with "Variant A OK"
One day earlier I thought these were core architecture. One day later they became museum pieces kept only for rollback
I did not delete them, because deletion cost was higher than keeping them — if Variant A gets complaints later, reverting one PR + switching storage path brings B back immediately
Three lessons I will carry forward
1. Try prompt first, build rules second
I spent half a day on 37 rules, and ten minutes on the prompt
Honestly, I only understood this after doing it, not before. My old habit was "rules first, prompt as fallback" — veteran engineer inertia
Now I reverse it: prompt first, rules only if prompt is not enough
LLMs push experiment cost so low that this should change development order
2. The value of audit trail is trust, not debugging
I thought provenance logs were for my own debugging. After shipping, I realized
The real user is PM, who needs the right to validate line by line
Engineers getting bitten by bugs is routine, but product teams have real anxiety signing off AI-generated content they did not verify with their own eyes
In AI products, audit trail is less a debug tool and more a sleep-at-night mechanism for non-engineers
3. Long-term value of the variant system
tts-variants.yaml + --variant A|B was built for one A/B experiment, but after using it I saw it can stay long-term
If Gemini ships new voice options, we can add Variant C If we need lower-grade-specific tone, add Variant D Storage split by folder avoids cache contamination
A well-designed one-time experiment can become infrastructure by itself
This is a principle I often underuse — spending 30 extra minutes on a light abstraction now can remove a full architecture refactor later
4. High-school interns who can code are more useful than you think
Across this whole story, several key judgments on pronunciation quality were made by two high-school engineering interns — accent differences, prompt wording, blind-test conclusions
My old mental model of "intern mentoring" was to assign low-risk chores, teach git, and call it good if the pipeline runs
Reality was the opposite — their coding ability was on par with many junior professionals, and age was an advantage:
- Their ears were not dulled by the tech echo chamber, so accent intuition was sharper
- Their LLM mindset was more native than ours — they were not stuck in "rules are engineering, prompts are cheating"
- They tried wild experiments I would not try, and several became final product decisions
If you are in industry and hesitating on high-school interns, my take is: worth it. Condition: treat them like engineers, not helpers
What happened in the next few weeks (bonus)
After deleting 37 rules, where did that bandwidth go
I thought things would calm down, but more hidden problems surfaced
1. Sentence splitting also moved from regex to LLM
We originally split sentences by regex over punctuation (period, exclamation, question mark, ellipsis), maintaining many edge cases
Since prompt-style thinking won on pronunciation, I copied it to splitting — moved to semantic splitting with Opus 4.7
Same pattern as TTS: rule era needs hundreds of corner cases, LLM reads semantics and outputs 2301 clean chunks
2. Audit trail evolved into a merge gate
PM's "line-by-line acceptance" request turned into stricter engineering discipline
I wrote a verifier with --strict. After each batch regeneration, JSONL records must match cloud audio objects 1:1 (no extra, no missing, exact keys), or PR cannot merge
It evolved from "PM acceptance dashboard" into "engineering merge gate"
3. Elegant fallback for safety filter
Out of 2000+ lines, Gemini safety filter refused two lines (unclear why, content looked normal)
Fix was elegant: keep the same cache key and call Chirp3-HD (the mainland-accent provider) for those two lines
Runtime does not care about provider source, it only plays cache hits. Because variant/provider abstraction was already clean, this fallback added only ~30 lines
4. Previously hidden UX problems surfaced
Once pronunciation was fixed, deeper issues became visible:
- Loudness jumps between paragraphs -> added loudnorm to standardize at -16 LUFS
- Subtitle highlights drifted from audio timing -> switched to character-weighted progress + prefetch
- TTS button had only idle/playing states, loading and error were confusing -> redesigned into 4-state UX
Those issues were always there, but hidden by the big fire of wrong accent
After putting out the big fire, small fires become visible — this is normal engineering reality
Cost comparison (final bill)
| Solution | Human Maintenance | Monthly API Cost | Audio Quality | Taiwan Accent |
|---|---|---|---|---|
| Azure + SSML phoneme | High (manual per character) | Expensive | Good, slightly rigid | ✅ Native |
| Chirp3-HD (mainland accent) | 0 | Medium | Very good | ❌ |
| Gemini Variant B (replacement table) | Medium (37 rules + upkeep) | ~$0.30 per 2k lines | Slightly rigid | ✅ via rules |
| Gemini Variant A (prompt-only) | Zero | ~$0.30 per 2k lines | Best | ✅ via prompt |
A wins across the board: zero maintenance, best quality, lowest cost
My judgment
At least in this project, it felt like my old work — phoneme tagging, sentence rules, polyphonic dictionaries, homophone replacement, SSML tags — now has much lower ROI on LLM TTS
The higher-output time now goes to prompt design, audit trail, variant system, cost monitoring, and trust mechanisms
I am not saying old skills are useless. If LLM prosody becomes unstable, old phoneme-level control still works as fallback. But the main path clearly moved
Have you also built something and watched it become obsolete the same day? These moments are some of the most valuable learning artifacts in this era
If you are building Chinese education or edtech products and are stuck on AI implementation, feel free to book 30 minutes — there is a good chance the pitfall I already hit is exactly the one you are hitting now
