Japanese Text to Speech: A Practical Pitch Accent Guide

Japanese neural TTS now passes for human in short narration, but pitch accent still breaks homographs and names. Fix it with kana respelling, SSML phoneme tags, a custom lexicon, and a native-ear review pass.

Editor in headphones adjusting an audio interface while reviewing a Japanese text to speech take

Japanese Text to Speech in 2026: How Natural Does It Really Sound?

Three years ago a Japanese synthetic voice gave itself away inside one sentence. The tell today is narrower, and it is almost always pitch rather than timbre. A neural voice can get breath, sibilance and sentence rhythm right, then flatten 箸 into 橋 and lose the listener who cannot say why the line felt off.

Japanese assigns each word a pitch contour, and that contour carries meaning the way stress does in English. Engines guess the contour from a parser, and parsers guess wrong on names, homographs and rare compounds.

What modern Japanese engines usually handle well

  • ます-form business narration built from short declarative sentences
  • Common counters such as 一つ、二人、三回 inside ordinary numeric ranges
  • Mainstream katakana loanwords that already sit in the dictionary
  • Breath placement at 。 boundaries

What still breaks

  • Homographs like 行った (いった / おこなった) and 人気 (にんき / ひとけ)
  • Surnames and place names: 日暮里, 東海林, 神楽坂
  • Minimal pairs separated only by accent: 雨 / 飴, 神 / 髪
  • English product names dropped mid-sentence
  • Long sentences where the engine picks the wrong 文節 boundary

ElevenLabs says its free plan includes 10,000 characters per month, which the company describes as roughly 10 minutes of audio. That is enough to test a Japanese script properly before you commit to any vendor.

Step 1 — Prepare the Japanese Script So the Engine Parses It Correctly

Time: 20–30 minutes per 1,000 characters. You need: a UTF-8 plain text editor and a glossary sheet for names and terms.

Most Japanese TTS problems are authoring problems. The engine reads what you typed, not what you meant, and a comma in the wrong place changes the phrase boundary.

Split anything past roughly 60 characters into two sentences. Long nested clauses are where parsers lose the subject and drop pitch in the wrong spot.

Place 、 where a human narrator would take a small breath, not where the grammar textbook allows one. Delete decorative punctuation such as ・, (), ―― and 〜 because engines treat them as pause instructions with unpredictable length.

Write numbers in the form you want spoken. 2025年 reads cleanly; 2025 alone may come out as a digit string. Same for 3,000円 versus 三千円 — pick one and stay consistent across the script.

Build a glossary row for every proper noun, product name and acronym before you generate anything. That list becomes your lexicon in Step 3, and it is the single highest-return thing you can prepare.

On its subscription plans, ElevenLabs bills text-to-speech at 1 credit per character, so a tighter script is directly a cheaper script.

Step 2 — Choose a Japanese Voice and Model That Fits Your Content Type

Time: 30–45 minutes. You need: three test passages of 80–120 characters each, plus headphones, not laptop speakers.

Do not audition voices on the vendor’s demo sentence. Use your own worst sentence — the one with a surname, a number and an English term in it.

Content type Voice traits to look for What to reject
E-learning / compliance Even tempo, clean ます endings, low breathiness Voices that rush particle endings
Product explainer Slight warmth, stable pitch across clauses Over-acted intonation on every sentence
Ad / promo Wider pitch range, crisp attack Voices that smear katakana loanwords
Audiobook / narration Consistent accent across long passages Models that reset tone per chunk
IVR / in-app prompts Flat, predictable, short-phrase clarity Emotional models with variable pacing

Model tier matters as much as voice identity. Fast models trade accent accuracy for latency, which is fine for prompts and bad for a course module.

ElevenLabs lists API text-to-speech at $0.08 per 1,000 characters on multilingual models and $0.04 per 1,000 characters on its Flash and Turbo models, which is the usual quality-versus-speed split.

If your Japanese audio ships inside video, check library breadth on the platform you already use — AI Studios lists text-to-speech voiceovers in 150+ languages and accents on paid plans and 30+ languages on its Free plan, alongside more than 1,000 AI voices.

Step 3 — Control Pitch Accent, Readings and Homographs with SSML and Custom Lexicons

Time: 45–60 minutes per 1,000 characters on a first pass. You need: SSML support or a pronunciation dictionary, and a published Japanese pitch-accent dictionary.

Standard Tokyo Japanese has four accent patterns. Learn them once and you can describe any correction to a vendor or a reviewer without hand-waving.

Pattern Where pitch drops Example Note
頭高 atamadaka after mora 1 雨 (あめ) High-low from the start
中高 nakadaka mid-word 心 (こころ) Drop lands inside the word
尾高 odaka after the final mora 花 (はな) Drop only shows on the following particle
平板 heiban no drop 飴 (あめ) Stays level into the particle

Three fixes, in order of how often they work. Respell the word in katakana inside the script — クレジット instead of an ambiguous kanji compound. Add a lexicon entry so the correction applies everywhere, not just in one line. Use an SSML phoneme tag when the engine accepts kana or IPA input.

When none of those exist, generate the problem word as its own short clip and splice it into the master take. Ugly, effective, and faster than fighting a parser.

Budget for iteration, because accent work means regenerating the same lines many times. With annual billing, ElevenLabs quotes effective monthly prices of $5 for Starter, $18.33 for Creator, $82.50 for Pro, $249.17 for Scale and $825 for Business.

Language specialist gesturing a pitch curve in the air while an audio engineer listens at a mixing desk

Step 4 — Tune Pacing, Pauses and Emphasis for Natural Japanese Delivery

Time: 20–40 minutes per finished minute of audio. You need: SSML break tags or a per-line speed control.

Japanese emphasis is not loudness. A native narrator emphasizes by resetting pitch at the top of a phrase and inserting a short beat before it, so cranking an SSML emphasis tag usually flattens the accent you just fixed.

Start from these values and adjust by ear rather than by rule:

Content type Pause after 、 Pause after 。 Speed
E-learning ~250 ms ~600 ms Slightly below default
Product explainer ~180 ms ~450 ms Default
Ad / promo ~120 ms ~350 ms Slightly above default
IVR prompt ~200 ms ~500 ms Default

Questions without か need a rising final mora, and most engines will not produce it from punctuation alone. Add a question mark, or respell the final particle, then listen.

Watch the particles. は, が and を getting clipped is the most common sign that your global speed is too high for the voice you picked.

Keep one unedited reference take. When ten pause tweaks later the line sounds worse, you need something to compare against.

Step 5 — Run a Native-Ear Review, Then Export and Sync to Video

Time: about 30 minutes per five minutes of audio. You need: a native Japanese speaker, a timecoded player and a shared correction list.

Run the review in two passes. First pass: your reviewer listens with no script and marks every timecode where they had to think. Second pass: they listen with the script open and classify each mark.

Use three severity buckets and fix only the first two:

  1. Wrong meaning — homograph, name, or accent flip that changes the word. Always fix.
  2. Unnatural but understood — odd break, robotic particle, rushed ending. Fix if it repeats.
  3. Stylistic preference — reviewer would have phrased it differently. Log it, ship anyway.

Export WAV at 48 kHz for video work, keep MP3 for review copies only, and normalize to whatever loudness target your distribution platform specifies.

For syncing to picture, generate a timecoded transcript of your own audio instead of eyeballing waveforms. ElevenLabs prices its Scribe speech-to-text API at $0.22 per hour and $0.39 per hour for the realtime version.

Japanese lines usually run longer than their English source, so leave slack in the edit before you lock the cut.

Native reviewer listening with eyes closed on headphones while an editor works at a studio console

Pre-Publish Quality Checklist for Japanese TTS

Run this before anything reaches a learner, a customer or a client. Every item is something a native listener notices within seconds, and every one of them is cheap to fix while the project file is still open.

  • [ ] Every proper noun in the script has a lexicon entry or katakana respelling
  • [ ] All homographs identified and locked to one reading
  • [ ] Numbers, dates and currency read back in the intended form
  • [ ] English terms pronounced consistently across the whole file
  • [ ] Particles は / が / を audible, not clipped
  • [ ] Question intonation rises where required
  • [ ] Pause lengths consistent between sections recorded on different days
  • [ ] Same voice and model version used throughout
  • [ ] Native reviewer signed off on bucket 1 and bucket 2 items
  • [ ] Loudness matched to platform target
  • [ ] Master exported as WAV 48 kHz, archived with the final script
  • [ ] Glossary saved for the next episode or module

Common Mistakes with Japanese Text to Speech (and How to Fix Them)

The failure that ruins more Japanese projects than anything else: shipping proper nouns unchecked. A course on a client’s product gets the client’s own company name wrong in the first 15 seconds, and nothing after that matters.

Fix it by extracting every name into a glossary during Step 1, generating a 30-second name-only test clip, and having a native speaker approve that clip before you produce the full script.

Other patterns worth avoiding:

  • Translating English pacing directly. English pause rhythm applied to Japanese sounds hesitant. Re-time after translation, not before.
  • Using emphasis tags for stress. They flatten accent contours. Insert a short pre-pause instead.
  • Switching models mid-project. Voice identity drifts between model versions. Lock the version in your project notes.
  • Testing on clean sentences. Audition with your hardest line, not the vendor’s demo.
  • Skipping the native review because the output sounded fine. If you do not read Japanese, you cannot hear an accent flip — that is what the reviewer is for.
  • Letting the script balloon. Character-based billing means every unnecessary clause has a price attached.

One more habit that pays off: keep a running list of every correction you make, per voice. By the third project, that list catches most errors before you generate a single line.

Frequently asked questions

How natural is Japanese text-to-speech for long-form narration?

Good neural voices hold up across e-learning modules and explainers, but accent accuracy drifts on proper nouns and rare compounds. Expect to correct roughly a handful of words per thousand characters, and always budget a native review pass before publishing.

What exactly is pitch accent, and why does it break TTS?

Japanese words carry a pitch contour that distinguishes meaning, such as 雨 (rain) versus 飴 (candy). Engines infer that contour from a parser, and parsers mispredict on names, homographs and loanwords, producing audio that sounds fluent but means the wrong thing.

Can I fix readings without SSML support?

Yes. Respell the word in katakana directly in the script, add it to a pronunciation dictionary if the platform offers one, or generate the problem word as a separate clip and splice it into the master take.

What does Japanese text-to-speech cost?

Pricing is usually per character. ElevenLabs lists API text-to-speech at $0.08 per 1,000 characters on multilingual models and $0.04 per 1,000 characters on Flash and Turbo models, and its free plan includes 10,000 characters per month.

Do I really need a native Japanese reviewer?

If nobody on your team reads Japanese, yes. Accent errors are inaudible to non-speakers because the audio itself sounds clean, so a native listener working through two passes with timecodes is the only reliable catch mechanism.

Sources

This article was drafted with AI assistance from the sources listed above and checked against them before publication.