Japanese neural TTS now passes for human in short narration, but pitch accent still breaks homographs and names. Fix it with kana respelling, SSML phoneme tags, a custom lexicon, and a native-ear review pass.

Japanese Text to Speech in 2026: How Natural Does It Really Sound?
Three years ago a Japanese synthetic voice gave itself away inside one sentence. The tell today is narrower, and it is almost always pitch rather than timbre. A neural voice can get breath, sibilance and sentence rhythm right, then flatten 箸 into 橋 and lose the listener who cannot say why the line felt off.
Japanese assigns each word a pitch contour, and that contour carries meaning the way stress does in English. Engines guess the contour from a parser, and parsers guess wrong on names, homographs and rare compounds.
What modern Japanese engines usually handle well
- ます-form business narration built from short declarative sentences
- Common counters such as 一つ、二人、三回 inside ordinary numeric ranges
- Mainstream katakana loanwords that already sit in the dictionary
- Breath placement at 。 boundaries
What still breaks
- Homographs like 行った (いった / おこなった) and 人気 (にんき / ひとけ)
- Surnames and place names: 日暮里, 東海林, 神楽坂
- Minimal pairs separated only by accent: 雨 / 飴, 神 / 髪
- English product names dropped mid-sentence
- Long sentences where the engine picks the wrong 文節 boundary
ElevenLabs says its free plan includes 10,000 characters per month, which the company describes as roughly 10 minutes of audio. That is enough to test a Japanese script properly before you commit to any vendor.
Step 1 — Prepare the Japanese Script So the Engine Parses It Correctly
Time: 20–30 minutes per 1,000 characters. You need: a UTF-8 plain text editor and a glossary sheet for names and terms.
Most Japanese TTS problems are authoring problems. The engine reads what you typed, not what you meant, and a comma in the wrong place changes the phrase boundary.
Split anything past roughly 60 characters into two sentences. Long nested clauses are where parsers lose the subject and drop pitch in the wrong spot.
Place 、 where a human narrator would take a small breath, not where the grammar textbook allows one. Delete decorative punctuation such as ・, (), ―― and 〜 because engines treat them as pause instructions with unpredictable length.
Write numbers in the form you want spoken. 2025年 reads cleanly; 2025 alone may come out as a digit string. Same for 3,000円 versus 三千円 — pick one and stay consistent across the script.
Build a glossary row for every proper noun, product name and acronym before you generate anything. That list becomes your lexicon in Step 3, and it is the single highest-return thing you can prepare.
On its subscription plans, ElevenLabs bills text-to-speech at 1 credit per character, so a tighter script is directly a cheaper script.
Step 2 — Choose a Japanese Voice and Model That Fits Your Content Type
Time: 30–45 minutes. You need: three test passages of 80–120 characters each, plus headphones, not laptop speakers.
Do not audition voices on the vendor’s demo sentence. Use your own worst sentence — the one with a surname, a number and an English term in it.
| Content type | Voice traits to look for | What to reject |
|---|---|---|
| E-learning / compliance | Even tempo, clean ます endings, low breathiness | Voices that rush particle endings |
| Product explainer | Slight warmth, stable pitch across clauses | Over-acted intonation on every sentence |
| Ad / promo | Wider pitch range, crisp attack | Voices that smear katakana loanwords |
| Audiobook / narration | Consistent accent across long passages | Models that reset tone per chunk |
| IVR / in-app prompts | Flat, predictable, short-phrase clarity | Emotional models with variable pacing |
Model tier matters as much as voice identity. Fast models trade accent accuracy for latency, which is fine for prompts and bad for a course module.
ElevenLabs lists API text-to-speech at $0.08 per 1,000 characters on multilingual models and $0.04 per 1,000 characters on its Flash and Turbo models, which is the usual quality-versus-speed split.
If your Japanese audio ships inside video, check library breadth on the platform you already use — AI Studios lists text-to-speech voiceovers in 150+ languages and accents on paid plans and 30+ languages on its Free plan, alongside more than 1,000 AI voices.
Step 3 — Control Pitch Accent, Readings and Homographs with SSML and Custom Lexicons
Time: 45–60 minutes per 1,000 characters on a first pass. You need: SSML support or a pronunciation dictionary, and a published Japanese pitch-accent dictionary.
Standard Tokyo Japanese has four accent patterns. Learn them once and you can describe any correction to a vendor or a reviewer without hand-waving.
| Pattern | Where pitch drops | Example | Note |
|---|---|---|---|
| 頭高 atamadaka | after mora 1 | 雨 (あめ) | High-low from the start |
| 中高 nakadaka | mid-word | 心 (こころ) | Drop lands inside the word |
| 尾高 odaka | after the final mora | 花 (はな) | Drop only shows on the following particle |
| 平板 heiban | no drop | 飴 (あめ) | Stays level into the particle |
Three fixes, in order of how often they work. Respell the word in katakana inside the script — クレジット instead of an ambiguous kanji compound. Add a lexicon entry so the correction applies everywhere, not just in one line. Use an SSML phoneme tag when the engine accepts kana or IPA input.
When none of those exist, generate the problem word as its own short clip and splice it into the master take. Ugly, effective, and faster than fighting a parser.
Budget for iteration, because accent work means regenerating the same lines many times. With annual billing, ElevenLabs quotes effective monthly prices of $5 for Starter, $18.33 for Creator, $82.50 for Pro, $249.17 for Scale and $825 for Business.

Step 4 — Tune Pacing, Pauses and Emphasis for Natural Japanese Delivery
Time: 20–40 minutes per finished minute of audio. You need: SSML break tags or a per-line speed control.
Japanese emphasis is not loudness. A native narrator emphasizes by resetting pitch at the top of a phrase and inserting a short beat before it, so cranking an SSML emphasis tag usually flattens the accent you just fixed.
Start from these values and adjust by ear rather than by rule:
| Content type | Pause after 、 | Pause after 。 | Speed |
|---|---|---|---|
| E-learning | ~250 ms | ~600 ms | Slightly below default |
| Product explainer | ~180 ms | ~450 ms | Default |
| Ad / promo | ~120 ms | ~350 ms | Slightly above default |
| IVR prompt | ~200 ms | ~500 ms | Default |
Questions without か need a rising final mora, and most engines will not produce it from punctuation alone. Add a question mark, or respell the final particle, then listen.
Watch the particles. は, が and を getting clipped is the most common sign that your global speed is too high for the voice you picked.
Keep one unedited reference take. When ten pause tweaks later the line sounds worse, you need something to compare against.
Step 5 — Run a Native-Ear Review, Then Export and Sync to Video
Time: about 30 minutes per five minutes of audio. You need: a native Japanese speaker, a timecoded player and a shared correction list.
Run the review in two passes. First pass: your reviewer listens with no script and marks every timecode where they had to think. Second pass: they listen with the script open and classify each mark.
Use three severity buckets and fix only the first two:
- Wrong meaning — homograph, name, or accent flip that changes the word. Always fix.
- Unnatural but understood — odd break, robotic particle, rushed ending. Fix if it repeats.
- Stylistic preference — reviewer would have phrased it differently. Log it, ship anyway.
Export WAV at 48 kHz for video work, keep MP3 for review copies only, and normalize to whatever loudness target your distribution platform specifies.
For syncing to picture, generate a timecoded transcript of your own audio instead of eyeballing waveforms. ElevenLabs prices its Scribe speech-to-text API at $0.22 per hour and $0.39 per hour for the realtime version.
Japanese lines usually run longer than their English source, so leave slack in the edit before you lock the cut.

Pre-Publish Quality Checklist for Japanese TTS
Run this before anything reaches a learner, a customer or a client. Every item is something a native listener notices within seconds, and every one of them is cheap to fix while the project file is still open.
- [ ] Every proper noun in the script has a lexicon entry or katakana respelling
- [ ] All homographs identified and locked to one reading
- [ ] Numbers, dates and currency read back in the intended form
- [ ] English terms pronounced consistently across the whole file
- [ ] Particles は / が / を audible, not clipped
- [ ] Question intonation rises where required
- [ ] Pause lengths consistent between sections recorded on different days
- [ ] Same voice and model version used throughout
- [ ] Native reviewer signed off on bucket 1 and bucket 2 items
- [ ] Loudness matched to platform target
- [ ] Master exported as WAV 48 kHz, archived with the final script
- [ ] Glossary saved for the next episode or module
Common Mistakes with Japanese Text to Speech (and How to Fix Them)
The failure that ruins more Japanese projects than anything else: shipping proper nouns unchecked. A course on a client’s product gets the client’s own company name wrong in the first 15 seconds, and nothing after that matters.
Fix it by extracting every name into a glossary during Step 1, generating a 30-second name-only test clip, and having a native speaker approve that clip before you produce the full script.
Other patterns worth avoiding:
- Translating English pacing directly. English pause rhythm applied to Japanese sounds hesitant. Re-time after translation, not before.
- Using emphasis tags for stress. They flatten accent contours. Insert a short pre-pause instead.
- Switching models mid-project. Voice identity drifts between model versions. Lock the version in your project notes.
- Testing on clean sentences. Audition with your hardest line, not the vendor’s demo.
- Skipping the native review because the output sounded fine. If you do not read Japanese, you cannot hear an accent flip — that is what the reviewer is for.
- Letting the script balloon. Character-based billing means every unnecessary clause has a price attached.
One more habit that pays off: keep a running list of every correction you make, per voice. By the third project, that list catches most errors before you generate a single line.
Frequently asked questions
How natural is Japanese text-to-speech for long-form narration?
Good neural voices hold up across e-learning modules and explainers, but accent accuracy drifts on proper nouns and rare compounds. Expect to correct roughly a handful of words per thousand characters, and always budget a native review pass before publishing.
What exactly is pitch accent, and why does it break TTS?
Japanese words carry a pitch contour that distinguishes meaning, such as 雨 (rain) versus 飴 (candy). Engines infer that contour from a parser, and parsers mispredict on names, homographs and loanwords, producing audio that sounds fluent but means the wrong thing.
Can I fix readings without SSML support?
Yes. Respell the word in katakana directly in the script, add it to a pronunciation dictionary if the platform offers one, or generate the problem word as a separate clip and splice it into the master take.
What does Japanese text-to-speech cost?
Pricing is usually per character. ElevenLabs lists API text-to-speech at $0.08 per 1,000 characters on multilingual models and $0.04 per 1,000 characters on Flash and Turbo models, and its free plan includes 10,000 characters per month.
Do I really need a native Japanese reviewer?
If nobody on your team reads Japanese, yes. Accent errors are inaudible to non-speakers because the audio itself sounds clean, so a native listener working through two passes with timecodes is the only reliable catch mechanism.
Sources
This article was drafted with AI assistance from the sources listed above and checked against them before publication.