Skip to main content
Voice AI

How to Make TTS Sound Human Without Touching the Model

Make TTS sound human with text normalization, pauses, prosody, pronunciation lexicons and SSML — a practical pipeline guide for voice AI engineers.

Diagram contrasting acoustic voice quality with linguistic naturalness in text-to-speech

You can make TTS sound human in most production systems without retraining, fine-tuning or even swapping the voice model. Modern neural voices already produce clean, natural-sounding audio. What usually gives a voice agent away is not the sound of the voice but what the engine was asked to say: an unexpanded abbreviation, a date read digit by digit, a question delivered like a statement, a 60-word sentence with no place to breathe.

This guide covers the layer that fixes those problems — the text and control pipeline between your application and the TTS engine. It is written for engineers building voice agents, assistants and multilingual products, with extra attention to Persian and mixed Persian–English speech.

Why TTS Can Sound Robotic Even With a Good Model

A neural TTS model learns to map text (or phonemes) to speech from recorded examples. It is very good at the things it saw constantly during training: common words, well-punctuated sentences, standard spellings. It is much weaker at inputs that are rare or ambiguous in written text:

  • Symbols and digits — 3/4, $1.2M, 10:30, v2.1, №
  • Abbreviations and acronyms — Dr., St., NASA vs. FBI vs. SQL
  • Homographs — read, lead, record; in Persian, words like «کرد» or «شیر» whose vowels are not written
  • Structure the model can’t hear — Markdown, bullet points, URLs, emoji, table fragments
  • Missing or wrong punctuation — the main signal most engines use for phrasing and intonation

When one of these appears, the model does its best guess. The guess is often intelligible but wrong in a way listeners notice immediately. A single mispronounced product name in an otherwise perfect sentence is enough to make the whole voice feel synthetic.

The practical conclusion: before you blame the model, look at the text you are sending it.

The Difference Between Voice Quality and Speech Naturalness

It helps to separate two properties that people often lump together as “sounds robotic”.

Voice quality (acoustic) is about the audio itself: timbre, clarity, absence of artifacts, breathiness, consistency between sentences. This is mostly determined by the model and the voice you chose. You fix it by changing models, voices or sample rates.

Speech naturalness (linguistic) is about whether the delivery matches how a person would say that content: correct words, sensible phrasing, stress on the right word, pauses in the right places, questions that sound like questions. This is heavily influenced by the input.

The discussion in Voice AI Mastery’s article on making TTS sound more human highlights an important distinction between the acoustic quality of a voice and the linguistic quality of what it says (How to Make Your TTS Sound Human — Without Touching the Model File). That distinction is the premise of everything below: once acoustic quality is “good enough”, most remaining complaints are linguistic, and linguistic problems can be fixed in your own code.

Text Preprocessing: The Most Underrated TTS Layer

Most TTS engines include a text front end that performs some normalization. Relying on it alone is risky for three reasons:

  1. It doesn’t know your domain. It can’t know that ZX-9 is a product that should be read “Z X nine” or that MinIO is “min-eye-oh”.
  2. It doesn’t know your locale conventions. Is 03/04/2026 March 4 or April 3? Is 1.500 one and a half or one thousand five hundred?
  3. It varies between engines and versions. If you switch providers, behavior changes silently.

A preprocessing layer you own makes the behavior explicit, testable and portable. Its job is to turn written text into speakable text: text that has exactly one reasonable way to be read aloud.

Problem Cause Solution Example
Digits read one by one Engine treats number as an identifier Expand to words by context 2026 → “twenty twenty-six” (year)
Wrong date Ambiguous day/month order Normalize using the user’s locale 03/04 → “April third” (en-GB)
Product name mangled Out-of-vocabulary word Pronunciation lexicon Galatea → “gal-uh-TEE-uh”
Flat, run-on delivery Very long sentence, few commas Segment at clause boundaries Split into 2–3 sentences
“Asterisk asterisk” read aloud Markdown in LLM output Strip formatting before TTS **Note:** → “Note:”
Statement-like question Missing question mark Restore punctuation Want to continue → Want to continue?

Punctuation and Pause Control

Punctuation is the cheapest prosody control you have. Most engines use it to decide where phrases end, how long to pause and whether pitch should fall or rise. Treat it as markup for the voice, not as copy-editing.

Periods create a full phrase boundary with a falling contour. Replacing a long comma chain with two or three sentences usually does more for naturalness than any other single change.

Commas create shorter breaks. They are most useful after introductory phrases (“If the payment fails, …”) and between list items. Too many commas produce a choppy, over-careful delivery.

Dashes and colons often produce a slightly longer or more “announcing” pause, but behavior varies by engine — test before relying on them.

Ellipses can produce hesitation in some engines and nothing in others. Avoid them unless you have verified the effect.

When punctuation isn’t precise enough and your engine supports SSML, use explicit breaks:

<speak>
  Your verification code is <say-as interpret-as="characters">4 8 2</say-as>
  <break time="250ms"/> <say-as interpret-as="characters">9 1 7</say-as>.
</speak>

A practical rule: use punctuation for normal speech rhythm, and reserve <break> for places where a listener needs time to process or write something down — codes, phone numbers, step boundaries in instructions.

Prosody, Stress, Pitch and Speaking Rhythm

Prosody is the melody and rhythm of speech: pitch movement, loudness, duration and pauses. Listeners use it to find the important word in a sentence and to tell questions from statements.

Put important information where stress naturally falls

In English, the default sentence stress tends to fall near the end of a phrase. You can use that instead of fighting it:

  • Weak: “The meeting that you asked about, which was moved, is now on Thursday, not Friday.” (the key contrast is buried mid-sentence)
  • Better: “Your meeting has moved. It’s now on Thursday, not Friday.”

Rewriting for natural focus works on every engine. SSML <emphasis> can help too, but support and strength vary — some neural engines ignore it or apply it subtly.

Questions

English yes/no questions typically rise at the end (“Is that correct?”), while wh-questions usually fall (“What time works for you?”). Engines infer this mainly from the question mark and word order. Common failures:

  • The question mark was dropped by an upstream step (for example, an LLM that returned a sentence fragment).
  • A question is embedded in a longer statement: “I can book it for you so do you want the morning slot.” Split it: “I can book it for you. Do you want the morning slot?”

Speaking rate and pitch

Adjust rate for content, not globally. Dense information — numbers, addresses, confirmation codes — benefits from a slightly slower rate. Small changes (around 5–10%) are usually enough; large changes introduce artifacts in many voices. Pitch changes via <prosody pitch> are rarely needed for naturalness and are easy to overdo.

Pronunciation Engineering

Pronunciation errors are the most noticeable naturalness defect because listeners immediately know the word is wrong. Handle them with a pronunciation lexicon: a maintained mapping from written forms to speakable forms.

There are three levels of control, from most portable to most precise:

  1. Respelling — replace the word with a spelling the engine reads correctly (Nginx → “engine x”). Works everywhere, but is engine-dependent in its results.
  2. Alias substitution — SSML <sub alias="…"> keeps the original text in logs while speaking the alias.
  3. Phonemes — SSML <phoneme alphabet="ipa" ph="…"> or a W3C PLS lexicon, where supported. Most precise; not all engines or voices honor it.
import re

LEXICON = {
    "Nginx": "engine x",
    "MinIO": "min eye oh",
    "PostgreSQL": "postgres Q L",
    "SQL": "sequel",
    "Galatea": "gala tea uh",
}

_pattern = re.compile(
    r"(?<!\w)(" + "|".join(map(re.escape, sorted(LEXICON, key=len, reverse=True))) + r")(?!\w)"
)

def apply_lexicon(text: str) -> str:
    return _pattern.sub(lambda m: LEXICON[m.group(1)], text)

Sorting keys by length matters: it makes PostgreSQL match before SQL. Keep the lexicon in version control, add an entry every time a listener reports a mispronunciation, and add that sentence to your regression tests.

Context-aware pronunciation. Homographs need context: “I read the report yesterday” vs. “Please read the report.” If a homograph matters in your domain, resolve it with a rule based on nearby words, a part-of-speech tagger, or by rewriting the template so it is unambiguous (“I went through the report yesterday”).

Acronyms fall into three groups — spelled out (FBI), read as a word (NASA), or conventional (SQL as “sequel” in many teams, “S Q L” in others). Decide per acronym and encode it in the lexicon; don’t let the engine guess.

Numbers, Dates, Abbreviations and Special Characters

Written numbers are compressed; spoken numbers depend on what the number is.

Written Meaning Spoken form
2026 year twenty twenty-six
2026 quantity two thousand twenty-six
1st ordinal first
$1.2M currency one point two million dollars
3.5% percentage three point five percent
10:30 time ten thirty
v2.3.1 version version two point three point one
+1 415 555 0132 phone read in digit groups with short pauses
https://example.com/docs URL usually don’t read it; say “the link in the message”

A few rules that prevent most errors:

  • Classify before expanding. Use surrounding words (“in 2026”, “costs”, “at”, “version”) and patterns to decide what a number is. When in doubt, choose the reading that is least wrong if mistaken.
  • Apply the user’s locale, not the server’s, for dates, decimals and currency.
  • Protect known abbreviations before sentence splitting, or “Dr. Smith” becomes two sentences.
  • Normalize exactly once. If your engine also normalizes, feeding it already-expanded text is fine; feeding it half-normalized text (some digits expanded, some not) can produce doubled or inconsistent readings.
  • Remove what shouldn’t be spoken: Markdown, emoji, HTML tags, reference markers like [1].

A compact example for currency and percentages:

import re

def _num_words(value: str) -> str:
    # Plug in your number-to-words library here (e.g. num2words).
    from num2words import num2words
    return num2words(float(value)) if "." in value else num2words(int(value))

SCALE = {"K": "thousand", "M": "million", "B": "billion"}

def expand_money(text: str) -> str:
    def repl(m):
        amount, scale = m.group(1).replace(",", ""), m.group(2)
        words = _num_words(amount)
        return f"{words} {SCALE[scale]} dollars" if scale else f"{words} dollars"
    return re.sub(r"\$(\d[\d,]*(?:\.\d+)?)([KMB])?\b", repl, text)

def expand_percent(text: str) -> str:
    return re.sub(r"(\d+(?:\.\d+)?)\s?%", lambda m: f"{_num_words(m.group(1))} percent", text)

Making Persian TTS Sound More Natural

Persian (Farsi) adds problems that English pipelines rarely handle:

  • Short vowels are not written. «کرد» can be kard (did) or kord (Kurd); «شیر» can be milk, lion or tap. The engine must infer vowels from context, and it will sometimes be wrong.
  • The ezafe is usually invisible. «کتاب من» is read ketāb-e man (“my book”), but the linking -e is not written. Missing or extra ezafe is a common source of unnatural Persian TTS.
  • Zero-width non-joiner (ZWNJ, نیم‌فاصله). «می‌روم» written as «می روم» or «میروم» can change tokenization and, depending on the front end, pronunciation.
  • Mixed character sets. Text copied from Arabic sources may contain Arabic «ي» and «ك» instead of Persian «ی» and «ک», plus Arabic-Indic digits.
  • Formal vs. spoken register. Written Persian («می‌خواهم بروم») differs strongly from spoken Persian («می‌خوام برم»). A voice agent reading formal written text sounds like a news anchor in a casual conversation.

The Persian section of this article (below) covers these in depth with Persian examples and code.

Handling Persian–English Mixed Text

Persian technical speech is full of English: product names, acronyms, commands. A Persian voice reading English words with Persian letter-to-sound rules produces errors; an English voice reading Persian is worse. Options, in order of preference:

  1. Transliterate in the lexicon — write the English term in Persian script as Persian speakers say it: Kubernetes → «کوبرنتیز», API → «اِی‌پی‌آی». This keeps a single voice and a single language, which usually sounds most natural for code-switching speakers.
  2. Use SSML <lang> to mark an English span, if your engine supports switching language (or voice) mid-sentence. Test carefully: some engines switch voice timbre abruptly.
  3. Translate when a common Persian term exists and your audience uses it («هوش مصنوعی» instead of “AI”).

SSML and Other Control Mechanisms

SSML (Speech Synthesis Markup Language, a W3C standard) is the most widely supported way to control speech beyond plain text. Useful elements:

Element Purpose Typical use
<break> explicit pause codes, step-by-step instructions
<prosody> rate, pitch, volume slow down dense information
<emphasis> stress contrastive focus
<say-as> interpretation hint dates, characters, cardinal/ordinal numbers
<sub> spoken alias abbreviations, brand names
<phoneme> exact pronunciation names, homographs
<lang> language switch foreign terms

Two caveats matter more than the element list:

  • Support differs by engine, and sometimes by voice. Some neural voices ignore <emphasis> or <phoneme>, or support only a subset of <say-as> formats. Check your provider’s documentation and verify by listening.
  • SSML is not a substitute for good text. Heavily tagged text is hard to maintain and can sound mechanical. Get the text right first; add tags where text alone can’t express the intent.

Some engines also offer non-SSML controls — custom lexicon uploads, style or emotion parameters, or phoneme input. These are engine-specific; wrap them behind your own interface so the rest of your pipeline doesn’t depend on one provider.

Building a TTS Preprocessing Pipeline

A maintainable pipeline is a sequence of small, testable steps:

source text (template, CMS, LLM output)
  → sanitize          remove Markdown, HTML, emoji, citations, URLs
  → language tagging  detect spans in other languages / scripts
  → normalize         numbers, dates, times, currency, units, symbols
  → lexicon           brand names, acronyms, known homographs
  → punctuation       restore missing ?/., split long sentences
  → markup (optional) SSML breaks, prosody, say-as
  → chunk             sentence/clause-sized pieces for streaming
  → TTS engine
  → post-process      join chunks, trim silence, consistent loudness
Technique Implementation layer Expected effect Cost / complexity
Sanitizing LLM output Preprocessing Removes read-aloud formatting noise Low
Number/date normalization Preprocessing Correct readings of dense info Medium (locale rules)
Pronunciation lexicon Preprocessing / engine lexicon Fixes names and acronyms Low to start, ongoing curation
Sentence segmentation Preprocessing Better phrasing and pauses Low
SSML breaks/prosody Markup Precise control where needed Medium; engine-dependent
Conversational rewriting Content / LLM prompt Less “read-aloud document” feel Medium
Streaming chunking Orchestration Lower first-audio latency Medium; affects prosody
Loudness/silence post-processing Audio Consistent listening experience Low

Write for the ear at the source

If an LLM generates the text, the cheapest fix is upstream: instruct it to write speakable output — short sentences, no lists or Markdown, no URLs, numbers the way they should be said, one question at a time. Then keep the preprocessing layer as a safety net, because models don’t follow formatting instructions perfectly.

Latency vs. naturalness

Streaming voice agents start synthesizing before the full response exists. The trade-off:

  • Smaller chunks → earlier first audio, but each chunk is synthesized with less context, so intonation across chunk boundaries can reset or sound disconnected.
  • Larger chunks → better prosody, but a longer wait before the user hears anything.

A common compromise: emit the first chunk at the first sentence or clause boundary, then send full sentences. Never cut in the middle of a clause.

import re

SENT_END = re.compile(r"(?<=[.!?؟…])\s+")

def sentences(text: str) -> list[str]:
    return [s.strip() for s in SENT_END.split(text) if s.strip()]

def chunks(text: str, max_chars: int = 220):
    """First sentence goes out alone (fast start); later ones are grouped."""
    parts = sentences(text)
    if not parts:
        return
    yield parts[0]
    buf = ""
    for s in parts[1:]:
        if buf and len(buf) + len(s) + 1 > max_chars:
            yield buf
            buf = s
        else:
            buf = f"{buf} {s}".strip()
    if buf:
        yield buf

Practical Before/After Examples

Long sentence → segmented speech
– Before: “Your order which you placed on Monday and which includes three items has been shipped and should arrive by Thursday but if it doesn’t you can contact support”
– After: “Your order from Monday has shipped. It has three items. It should arrive by Thursday. If it doesn’t, just contact support.”
– Difference: clear phrase boundaries, natural pitch reset per sentence, the key fact (Thursday) lands at a sentence end.

Ambiguous abbreviation → safe text
– Before: “Meet Dr. Lee at 5 St. James St.”
– After: “Meet Doctor Lee at five Saint James Street.”
– Difference: no guessing between “Saint” and “Street”; no false sentence break after “Dr.”

Question → correct intonation
– Before: “Would you like me to reschedule”
– After: “Would you like me to reschedule?”
– Difference: rising final contour instead of a flat statement.

Numbers → spoken form
– Before: “Revenue grew 12.5% to $3.4M in Q3 2026.”
– After: “Revenue grew twelve point five percent, to three point four million dollars, in the third quarter of twenty twenty-six.”
– Difference: every number read the way a person would say it; the commas give the listener time.

Mixed Persian–English → pronunciation-safe
– Before: «فایل را روی MinIO آپلود کنید و از API استفاده کنید.»
– After: «فایل را روی مین‌آی‌او آپلود کنید و از اِی‌پی‌آی استفاده کنید.»
– Difference: English terms pronounced the way Persian speakers say them, in the same voice.

How to Evaluate TTS Naturalness

“It sounds better to me” doesn’t scale. Combine subjective and objective checks:

  • Listening tests with native speakers. Mean Opinion Score (MOS, per ITU-T P.800) for absolute ratings; paired A/B preference tests for comparing two pipeline versions, which are usually more sensitive for small changes. MUSHRA (ITU-R BS.1534) is another option for multi-system comparisons.
  • A regression set of hard sentences. Collect real failures: product names, dates, prices, questions, mixed-language sentences. Re-synthesize them on every pipeline change.
  • ASR round-trip. Transcribe the TTS output with a speech recognizer and compare with the intended text. A rising word error rate flags pronunciation regressions automatically. It measures intelligibility, not naturalness, so treat it as a smoke test.
  • Unit tests for the text layer. The preprocessing layer is deterministic — test it like any other code: input string in, expected speakable string out.
  • Production signals. Repeat requests (“sorry?”), barge-ins and drop-offs at specific prompts often point to a phrase that sounds wrong.

Common Mistakes

  1. Sending raw LLM output to TTS — Markdown symbols, lists and links get read aloud or cause odd pauses.
  2. Evaluating only on demo sentences — the real failures are in prices, dates and names.
  3. Normalizing twice or partially — mixed expanded/unexpanded numbers confuse the engine’s own front end.
  4. Using the server locale instead of the user’s for dates and numbers.
  5. Over-tagging with SSML — dozens of breaks and prosody changes sound mechanical and are hard to maintain.
  6. Chunking mid-clause for latency — saves milliseconds, costs naturalness at every boundary.
  7. No owner for the lexicon — mispronunciations get reported and never fixed.
  8. Treating Persian like English with a different alphabet — ignoring ZWNJ, Arabic characters and register differences.

Production Checklist to Make TTS Sound Human

  • Text sanitizer strips Markdown, HTML, emoji, citations and URLs
  • Numbers, dates, times, currency and units normalized per user locale
  • Pronunciation lexicon in version control, with an owner
  • Abbreviations protected before sentence segmentation
  • Missing question marks restored; long sentences split
  • SSML used only where text alone is insufficient, and verified per voice
  • Mixed-language spans handled by lexicon or <lang>
  • Streaming chunks cut only at sentence/clause boundaries
  • Consistent loudness and trimmed silence between chunks
  • Regression set of hard sentences re-synthesized on every change
  • ASR round-trip check in CI; periodic native-speaker listening tests

Conclusion

To make TTS sound human, start with the text, not the model. A modern neural voice will usually say well-prepared text naturally; it struggles with digits, abbreviations, missing punctuation, unmarked questions and words it has never seen. A preprocessing layer you own — sanitization, normalization, a pronunciation lexicon, sentence segmentation and targeted SSML — fixes most of these problems, works across engines, and can be tested like any other code. For Persian and code-switched speech, the same layer is where ZWNJ, character normalization, register and English terms get handled. Build it once, measure it continuously, and your voice agent will sound more human with the model you already have.

Frequently Asked Questions

How can I make TTS sound more human without retraining the model?

Improve the text you send: normalize numbers and dates, expand abbreviations, fix punctuation, split long sentences, add a pronunciation lexicon for names and acronyms, and use SSML where your engine supports it.

Does SSML work with every TTS engine?

No. Most major engines support core elements such as <break> and <prosody>, but support for <emphasis>, <phoneme> and some <say-as> formats varies by engine and even by voice. Always verify by listening.

Why does my TTS read numbers incorrectly?

Written numbers are ambiguous: 2026 can be a year or a quantity, 03/04 can be March or April. Classify each number by context and expand it to words in the user’s locale before synthesis.

How do I fix mispronounced brand names?

Maintain a pronunciation lexicon that maps each name to a respelling, an SSML alias or phonemes, apply it before synthesis, and add every fixed case to a regression test set.

Does splitting text for streaming hurt naturalness?

It can. Each chunk is synthesized with limited context, so cut only at sentence or clause boundaries. Sending the first sentence alone keeps latency low without cutting clauses.

How do I measure TTS naturalness?

Use native-speaker listening tests (MOS or A/B preference), a regression set of difficult sentences, and an ASR round-trip check to catch pronunciation regressions automatically.

References

Have a project in mind?

From idea to prototype and product, we're with you.