
You can make TTS sound human in most production systems without retraining, fine-tuning or even swapping the voice model. Modern neural voices already produce clean, natural-sounding audio. What usually gives a voice agent away is not the sound of the voice but what the engine was asked to say: an unexpanded abbreviation, a date read digit by digit, a question delivered like a statement, a 60-word sentence with no place to breathe.
This guide covers the layer that fixes those problems — the text and control pipeline between your application and the TTS engine. It is written for engineers building voice agents, assistants and multilingual products, with extra attention to Persian and mixed Persian–English speech.
Why TTS Can Sound Robotic Even With a Good Model
A neural TTS model learns to map text (or phonemes) to speech from recorded examples. It is very good at the things it saw constantly during training: common words, well-punctuated sentences, standard spellings. It is much weaker at inputs that are rare or ambiguous in written text:
- Symbols and digits —
3/4,$1.2M,10:30,v2.1,№ - Abbreviations and acronyms —
Dr.,St.,NASAvs.FBIvs.SQL - Homographs — read, lead, record; in Persian, words like «کرد» or «شیر» whose vowels are not written
- Structure the model can’t hear — Markdown, bullet points, URLs, emoji, table fragments
- Missing or wrong punctuation — the main signal most engines use for phrasing and intonation
When one of these appears, the model does its best guess. The guess is often intelligible but wrong in a way listeners notice immediately. A single mispronounced product name in an otherwise perfect sentence is enough to make the whole voice feel synthetic.
The practical conclusion: before you blame the model, look at the text you are sending it.
The Difference Between Voice Quality and Speech Naturalness
It helps to separate two properties that people often lump together as “sounds robotic”.
Voice quality (acoustic) is about the audio itself: timbre, clarity, absence of artifacts, breathiness, consistency between sentences. This is mostly determined by the model and the voice you chose. You fix it by changing models, voices or sample rates.
Speech naturalness (linguistic) is about whether the delivery matches how a person would say that content: correct words, sensible phrasing, stress on the right word, pauses in the right places, questions that sound like questions. This is heavily influenced by the input.
The discussion in Voice AI Mastery’s article on making TTS sound more human highlights an important distinction between the acoustic quality of a voice and the linguistic quality of what it says (How to Make Your TTS Sound Human — Without Touching the Model File). That distinction is the premise of everything below: once acoustic quality is “good enough”, most remaining complaints are linguistic, and linguistic problems can be fixed in your own code.
Text Preprocessing: The Most Underrated TTS Layer
Most TTS engines include a text front end that performs some normalization. Relying on it alone is risky for three reasons:
- It doesn’t know your domain. It can’t know that
ZX-9is a product that should be read “Z X nine” or thatMinIOis “min-eye-oh”. - It doesn’t know your locale conventions. Is
03/04/2026March 4 or April 3? Is1.500one and a half or one thousand five hundred? - It varies between engines and versions. If you switch providers, behavior changes silently.
A preprocessing layer you own makes the behavior explicit, testable and portable. Its job is to turn written text into speakable text: text that has exactly one reasonable way to be read aloud.
| Problem | Cause | Solution | Example |
|---|---|---|---|
| Digits read one by one | Engine treats number as an identifier | Expand to words by context | 2026 → “twenty twenty-six” (year) |
| Wrong date | Ambiguous day/month order | Normalize using the user’s locale | 03/04 → “April third” (en-GB) |
| Product name mangled | Out-of-vocabulary word | Pronunciation lexicon | Galatea → “gal-uh-TEE-uh” |
| Flat, run-on delivery | Very long sentence, few commas | Segment at clause boundaries | Split into 2–3 sentences |
| “Asterisk asterisk” read aloud | Markdown in LLM output | Strip formatting before TTS | **Note:** → “Note:” |
| Statement-like question | Missing question mark | Restore punctuation | Want to continue → Want to continue? |
Punctuation and Pause Control
Punctuation is the cheapest prosody control you have. Most engines use it to decide where phrases end, how long to pause and whether pitch should fall or rise. Treat it as markup for the voice, not as copy-editing.
Periods create a full phrase boundary with a falling contour. Replacing a long comma chain with two or three sentences usually does more for naturalness than any other single change.
Commas create shorter breaks. They are most useful after introductory phrases (“If the payment fails, …”) and between list items. Too many commas produce a choppy, over-careful delivery.
Dashes and colons often produce a slightly longer or more “announcing” pause, but behavior varies by engine — test before relying on them.
Ellipses can produce hesitation in some engines and nothing in others. Avoid them unless you have verified the effect.
When punctuation isn’t precise enough and your engine supports SSML, use explicit breaks:
<speak>
Your verification code is <say-as interpret-as="characters">4 8 2</say-as>
<break time="250ms"/> <say-as interpret-as="characters">9 1 7</say-as>.
</speak>
A practical rule: use punctuation for normal speech rhythm, and reserve <break> for places where a listener needs time to process or write something down — codes, phone numbers, step boundaries in instructions.
Prosody, Stress, Pitch and Speaking Rhythm
Prosody is the melody and rhythm of speech: pitch movement, loudness, duration and pauses. Listeners use it to find the important word in a sentence and to tell questions from statements.
Put important information where stress naturally falls
In English, the default sentence stress tends to fall near the end of a phrase. You can use that instead of fighting it:
- Weak: “The meeting that you asked about, which was moved, is now on Thursday, not Friday.” (the key contrast is buried mid-sentence)
- Better: “Your meeting has moved. It’s now on Thursday, not Friday.”
Rewriting for natural focus works on every engine. SSML <emphasis> can help too, but support and strength vary — some neural engines ignore it or apply it subtly.
Questions
English yes/no questions typically rise at the end (“Is that correct?”), while wh-questions usually fall (“What time works for you?”). Engines infer this mainly from the question mark and word order. Common failures:
- The question mark was dropped by an upstream step (for example, an LLM that returned a sentence fragment).
- A question is embedded in a longer statement: “I can book it for you so do you want the morning slot.” Split it: “I can book it for you. Do you want the morning slot?”
Speaking rate and pitch
Adjust rate for content, not globally. Dense information — numbers, addresses, confirmation codes — benefits from a slightly slower rate. Small changes (around 5–10%) are usually enough; large changes introduce artifacts in many voices. Pitch changes via <prosody pitch> are rarely needed for naturalness and are easy to overdo.
Pronunciation Engineering
Pronunciation errors are the most noticeable naturalness defect because listeners immediately know the word is wrong. Handle them with a pronunciation lexicon: a maintained mapping from written forms to speakable forms.
There are three levels of control, from most portable to most precise:
- Respelling — replace the word with a spelling the engine reads correctly (
Nginx→ “engine x”). Works everywhere, but is engine-dependent in its results. - Alias substitution — SSML
<sub alias="…">keeps the original text in logs while speaking the alias. - Phonemes — SSML
<phoneme alphabet="ipa" ph="…">or a W3C PLS lexicon, where supported. Most precise; not all engines or voices honor it.
import re
LEXICON = {
"Nginx": "engine x",
"MinIO": "min eye oh",
"PostgreSQL": "postgres Q L",
"SQL": "sequel",
"Galatea": "gala tea uh",
}
_pattern = re.compile(
r"(?<!\w)(" + "|".join(map(re.escape, sorted(LEXICON, key=len, reverse=True))) + r")(?!\w)"
)
def apply_lexicon(text: str) -> str:
return _pattern.sub(lambda m: LEXICON[m.group(1)], text)
Sorting keys by length matters: it makes PostgreSQL match before SQL. Keep the lexicon in version control, add an entry every time a listener reports a mispronunciation, and add that sentence to your regression tests.
Context-aware pronunciation. Homographs need context: “I read the report yesterday” vs. “Please read the report.” If a homograph matters in your domain, resolve it with a rule based on nearby words, a part-of-speech tagger, or by rewriting the template so it is unambiguous (“I went through the report yesterday”).
Acronyms fall into three groups — spelled out (FBI), read as a word (NASA), or conventional (SQL as “sequel” in many teams, “S Q L” in others). Decide per acronym and encode it in the lexicon; don’t let the engine guess.
Numbers, Dates, Abbreviations and Special Characters
Written numbers are compressed; spoken numbers depend on what the number is.
| Written | Meaning | Spoken form |
|---|---|---|
2026 |
year | twenty twenty-six |
2026 |
quantity | two thousand twenty-six |
1st |
ordinal | first |
$1.2M |
currency | one point two million dollars |
3.5% |
percentage | three point five percent |
10:30 |
time | ten thirty |
v2.3.1 |
version | version two point three point one |
+1 415 555 0132 |
phone | read in digit groups with short pauses |
https://example.com/docs |
URL | usually don’t read it; say “the link in the message” |
A few rules that prevent most errors:
- Classify before expanding. Use surrounding words (“in 2026”, “costs”, “at”, “version”) and patterns to decide what a number is. When in doubt, choose the reading that is least wrong if mistaken.
- Apply the user’s locale, not the server’s, for dates, decimals and currency.
- Protect known abbreviations before sentence splitting, or “Dr. Smith” becomes two sentences.
- Normalize exactly once. If your engine also normalizes, feeding it already-expanded text is fine; feeding it half-normalized text (some digits expanded, some not) can produce doubled or inconsistent readings.
- Remove what shouldn’t be spoken: Markdown, emoji, HTML tags, reference markers like
[1].
A compact example for currency and percentages:
import re
def _num_words(value: str) -> str:
# Plug in your number-to-words library here (e.g. num2words).
from num2words import num2words
return num2words(float(value)) if "." in value else num2words(int(value))
SCALE = {"K": "thousand", "M": "million", "B": "billion"}
def expand_money(text: str) -> str:
def repl(m):
amount, scale = m.group(1).replace(",", ""), m.group(2)
words = _num_words(amount)
return f"{words} {SCALE[scale]} dollars" if scale else f"{words} dollars"
return re.sub(r"\$(\d[\d,]*(?:\.\d+)?)([KMB])?\b", repl, text)
def expand_percent(text: str) -> str:
return re.sub(r"(\d+(?:\.\d+)?)\s?%", lambda m: f"{_num_words(m.group(1))} percent", text)
Making Persian TTS Sound More Natural
Persian (Farsi) adds problems that English pipelines rarely handle:
- Short vowels are not written. «کرد» can be kard (did) or kord (Kurd); «شیر» can be milk, lion or tap. The engine must infer vowels from context, and it will sometimes be wrong.
- The ezafe is usually invisible. «کتاب من» is read ketāb-e man (“my book”), but the linking -e is not written. Missing or extra ezafe is a common source of unnatural Persian TTS.
- Zero-width non-joiner (ZWNJ, نیمفاصله). «میروم» written as «می روم» or «میروم» can change tokenization and, depending on the front end, pronunciation.
- Mixed character sets. Text copied from Arabic sources may contain Arabic «ي» and «ك» instead of Persian «ی» and «ک», plus Arabic-Indic digits.
- Formal vs. spoken register. Written Persian («میخواهم بروم») differs strongly from spoken Persian («میخوام برم»). A voice agent reading formal written text sounds like a news anchor in a casual conversation.
The Persian section of this article (below) covers these in depth with Persian examples and code.
Handling Persian–English Mixed Text
Persian technical speech is full of English: product names, acronyms, commands. A Persian voice reading English words with Persian letter-to-sound rules produces errors; an English voice reading Persian is worse. Options, in order of preference:
- Transliterate in the lexicon — write the English term in Persian script as Persian speakers say it:
Kubernetes→ «کوبرنتیز»,API→ «اِیپیآی». This keeps a single voice and a single language, which usually sounds most natural for code-switching speakers. - Use SSML
<lang>to mark an English span, if your engine supports switching language (or voice) mid-sentence. Test carefully: some engines switch voice timbre abruptly. - Translate when a common Persian term exists and your audience uses it («هوش مصنوعی» instead of “AI”).
SSML and Other Control Mechanisms
SSML (Speech Synthesis Markup Language, a W3C standard) is the most widely supported way to control speech beyond plain text. Useful elements:
| Element | Purpose | Typical use |
|---|---|---|
<break> |
explicit pause | codes, step-by-step instructions |
<prosody> |
rate, pitch, volume | slow down dense information |
<emphasis> |
stress | contrastive focus |
<say-as> |
interpretation hint | dates, characters, cardinal/ordinal numbers |
<sub> |
spoken alias | abbreviations, brand names |
<phoneme> |
exact pronunciation | names, homographs |
<lang> |
language switch | foreign terms |
Two caveats matter more than the element list:
- Support differs by engine, and sometimes by voice. Some neural voices ignore
<emphasis>or<phoneme>, or support only a subset of<say-as>formats. Check your provider’s documentation and verify by listening. - SSML is not a substitute for good text. Heavily tagged text is hard to maintain and can sound mechanical. Get the text right first; add tags where text alone can’t express the intent.
Some engines also offer non-SSML controls — custom lexicon uploads, style or emotion parameters, or phoneme input. These are engine-specific; wrap them behind your own interface so the rest of your pipeline doesn’t depend on one provider.
Building a TTS Preprocessing Pipeline
A maintainable pipeline is a sequence of small, testable steps:
source text (template, CMS, LLM output)
→ sanitize remove Markdown, HTML, emoji, citations, URLs
→ language tagging detect spans in other languages / scripts
→ normalize numbers, dates, times, currency, units, symbols
→ lexicon brand names, acronyms, known homographs
→ punctuation restore missing ?/., split long sentences
→ markup (optional) SSML breaks, prosody, say-as
→ chunk sentence/clause-sized pieces for streaming
→ TTS engine
→ post-process join chunks, trim silence, consistent loudness
| Technique | Implementation layer | Expected effect | Cost / complexity |
|---|---|---|---|
| Sanitizing LLM output | Preprocessing | Removes read-aloud formatting noise | Low |
| Number/date normalization | Preprocessing | Correct readings of dense info | Medium (locale rules) |
| Pronunciation lexicon | Preprocessing / engine lexicon | Fixes names and acronyms | Low to start, ongoing curation |
| Sentence segmentation | Preprocessing | Better phrasing and pauses | Low |
| SSML breaks/prosody | Markup | Precise control where needed | Medium; engine-dependent |
| Conversational rewriting | Content / LLM prompt | Less “read-aloud document” feel | Medium |
| Streaming chunking | Orchestration | Lower first-audio latency | Medium; affects prosody |
| Loudness/silence post-processing | Audio | Consistent listening experience | Low |
Write for the ear at the source
If an LLM generates the text, the cheapest fix is upstream: instruct it to write speakable output — short sentences, no lists or Markdown, no URLs, numbers the way they should be said, one question at a time. Then keep the preprocessing layer as a safety net, because models don’t follow formatting instructions perfectly.
Latency vs. naturalness
Streaming voice agents start synthesizing before the full response exists. The trade-off:
- Smaller chunks → earlier first audio, but each chunk is synthesized with less context, so intonation across chunk boundaries can reset or sound disconnected.
- Larger chunks → better prosody, but a longer wait before the user hears anything.
A common compromise: emit the first chunk at the first sentence or clause boundary, then send full sentences. Never cut in the middle of a clause.
import re
SENT_END = re.compile(r"(?<=[.!?؟…])\s+")
def sentences(text: str) -> list[str]:
return [s.strip() for s in SENT_END.split(text) if s.strip()]
def chunks(text: str, max_chars: int = 220):
"""First sentence goes out alone (fast start); later ones are grouped."""
parts = sentences(text)
if not parts:
return
yield parts[0]
buf = ""
for s in parts[1:]:
if buf and len(buf) + len(s) + 1 > max_chars:
yield buf
buf = s
else:
buf = f"{buf} {s}".strip()
if buf:
yield buf
Practical Before/After Examples
Long sentence → segmented speech
– Before: “Your order which you placed on Monday and which includes three items has been shipped and should arrive by Thursday but if it doesn’t you can contact support”
– After: “Your order from Monday has shipped. It has three items. It should arrive by Thursday. If it doesn’t, just contact support.”
– Difference: clear phrase boundaries, natural pitch reset per sentence, the key fact (Thursday) lands at a sentence end.
Ambiguous abbreviation → safe text
– Before: “Meet Dr. Lee at 5 St. James St.”
– After: “Meet Doctor Lee at five Saint James Street.”
– Difference: no guessing between “Saint” and “Street”; no false sentence break after “Dr.”
Question → correct intonation
– Before: “Would you like me to reschedule”
– After: “Would you like me to reschedule?”
– Difference: rising final contour instead of a flat statement.
Numbers → spoken form
– Before: “Revenue grew 12.5% to $3.4M in Q3 2026.”
– After: “Revenue grew twelve point five percent, to three point four million dollars, in the third quarter of twenty twenty-six.”
– Difference: every number read the way a person would say it; the commas give the listener time.
Mixed Persian–English → pronunciation-safe
– Before: «فایل را روی MinIO آپلود کنید و از API استفاده کنید.»
– After: «فایل را روی مینآیاو آپلود کنید و از اِیپیآی استفاده کنید.»
– Difference: English terms pronounced the way Persian speakers say them, in the same voice.
How to Evaluate TTS Naturalness
“It sounds better to me” doesn’t scale. Combine subjective and objective checks:
- Listening tests with native speakers. Mean Opinion Score (MOS, per ITU-T P.800) for absolute ratings; paired A/B preference tests for comparing two pipeline versions, which are usually more sensitive for small changes. MUSHRA (ITU-R BS.1534) is another option for multi-system comparisons.
- A regression set of hard sentences. Collect real failures: product names, dates, prices, questions, mixed-language sentences. Re-synthesize them on every pipeline change.
- ASR round-trip. Transcribe the TTS output with a speech recognizer and compare with the intended text. A rising word error rate flags pronunciation regressions automatically. It measures intelligibility, not naturalness, so treat it as a smoke test.
- Unit tests for the text layer. The preprocessing layer is deterministic — test it like any other code: input string in, expected speakable string out.
- Production signals. Repeat requests (“sorry?”), barge-ins and drop-offs at specific prompts often point to a phrase that sounds wrong.
Common Mistakes
- Sending raw LLM output to TTS — Markdown symbols, lists and links get read aloud or cause odd pauses.
- Evaluating only on demo sentences — the real failures are in prices, dates and names.
- Normalizing twice or partially — mixed expanded/unexpanded numbers confuse the engine’s own front end.
- Using the server locale instead of the user’s for dates and numbers.
- Over-tagging with SSML — dozens of breaks and prosody changes sound mechanical and are hard to maintain.
- Chunking mid-clause for latency — saves milliseconds, costs naturalness at every boundary.
- No owner for the lexicon — mispronunciations get reported and never fixed.
- Treating Persian like English with a different alphabet — ignoring ZWNJ, Arabic characters and register differences.
Production Checklist to Make TTS Sound Human
- Text sanitizer strips Markdown, HTML, emoji, citations and URLs
- Numbers, dates, times, currency and units normalized per user locale
- Pronunciation lexicon in version control, with an owner
- Abbreviations protected before sentence segmentation
- Missing question marks restored; long sentences split
- SSML used only where text alone is insufficient, and verified per voice
- Mixed-language spans handled by lexicon or
<lang> - Streaming chunks cut only at sentence/clause boundaries
- Consistent loudness and trimmed silence between chunks
- Regression set of hard sentences re-synthesized on every change
- ASR round-trip check in CI; periodic native-speaker listening tests
Conclusion
To make TTS sound human, start with the text, not the model. A modern neural voice will usually say well-prepared text naturally; it struggles with digits, abbreviations, missing punctuation, unmarked questions and words it has never seen. A preprocessing layer you own — sanitization, normalization, a pronunciation lexicon, sentence segmentation and targeted SSML — fixes most of these problems, works across engines, and can be tested like any other code. For Persian and code-switched speech, the same layer is where ZWNJ, character normalization, register and English terms get handled. Build it once, measure it continuously, and your voice agent will sound more human with the model you already have.
Frequently Asked Questions
How can I make TTS sound more human without retraining the model?
Improve the text you send: normalize numbers and dates, expand abbreviations, fix punctuation, split long sentences, add a pronunciation lexicon for names and acronyms, and use SSML where your engine supports it.
Does SSML work with every TTS engine?
No. Most major engines support core elements such as <break> and <prosody>, but support for <emphasis>, <phoneme> and some <say-as> formats varies by engine and even by voice. Always verify by listening.
Why does my TTS read numbers incorrectly?
Written numbers are ambiguous: 2026 can be a year or a quantity, 03/04 can be March or April. Classify each number by context and expand it to words in the user’s locale before synthesis.
How do I fix mispronounced brand names?
Maintain a pronunciation lexicon that maps each name to a respelling, an SSML alias or phonemes, apply it before synthesis, and add every fixed case to a regression test set.
Does splitting text for streaming hurt naturalness?
It can. Each chunk is synthesized with limited context, so cut only at sentence or clause boundaries. Sending the first sentence alone keeps latency low without cutting clauses.
How do I measure TTS naturalness?
Use native-speaker listening tests (MOS or A/B preference), a regression set of difficult sentences, and an ASR round-trip check to catch pronunciation regressions automatically.
References
- Mahimai Raja J, How to Make Your TTS Sound Human — Without Touching the Model File, Voice AI Mastery, 2026
- W3C, Speech Synthesis Markup Language (SSML) 1.1
- W3C, Pronunciation Lexicon Specification (PLS) 1.0
- ITU-T, Recommendation P.800 (MOS listening tests)
