How to write text that reads well
The same model reads the same text differently depending on how it is written. Below are the things that actually change the result — checked with a clock, not by feel.
1. Punctuation is the conductor
A full stop is not decoration — it is a pause with a measurable length. The same thirteen-word line, once with no punctuation at all and once broken into three sentences, comes out at 4.53 s versus 5.29 s. So each full stop is worth about 0.38 seconds of silence and, more importantly, it drops the melody instead of letting the words spill into each other.
✗ the meeting is at five we will talk about the budget then we go home
✓ The meeting is at five. We will talk about the budget. Then we go home.
A comma gives a shorter pause than a full stop, and an ellipsis the longest one. Question and exclamation marks change the melody rather than just the length; with the livelier voices (F4, F2) you hear it immediately, with the even ones (M5, F3) it is more discreet.
2. Numbers: write them the way you want to hear them
The model does read digits, but it decides on its own how to group them. “0888 123 456” takes 9.26 seconds; the same content written out in words takes 11.91 — the difference is precisely the grouping. If the listener is meant to write a phone number down, spelling it out in words gives you control over where the pauses fall.
for speed: 25 December 2026
for dictation: two five … December … twenty twenty six
for money: nineteen euro and ninety nine cents
The same goes for times, years and units. “14:30” works, but “half past two in the afternoon” is unambiguous in any language.
3. Abbreviations and acronyms
An acronym that should be read letter by letter is written with spaces between the letters: instead of “NASA”, write “N A S A”. One that is read as a word is left alone. A full stop inside an abbreviation (“Dr.”, “St.”) is sometimes taken for the end of a sentence and drops a pause in the middle of your phrase — if that bothers you, spell it out: “Doctor”, “Street”.
4. The <laugh>, <breath> and <sigh> tags
The three tags are not emoji — they are actual sound the model produces. Each adds roughly two to three tenths of a second and lands exactly where you put it. They work best between sentences rather than mid-phrase.
That was unexpected <laugh>. I did not see it coming.
Right <breath> let us start from the beginning.
Well <sigh> that is one way to do it.
Restraint matters: one tag every few sentences sounds like a person; three in a row sound like a parody.
5. The three sliders
Speed (0.5 – 2.0)
Speed divides the length almost exactly: the same sentence comes out at 4.60 s at 0.7, 3.27 s at 1.0 and 2.16 s at 1.5. So if you have 40 seconds of audio and a 30-second slot, you need 40 ÷ 30 ≈ 1.33. Above about 1.4 the consonants start to smear, though — cutting a sentence beats pushing to 1.6.
Quality steps (1 – 40)
Steps do not change the length of the audio — only how much work the model puts in. The same sentence comes out at 3.55 seconds at 4, 16 and 40 steps alike; what changes is the time to make it: 0.8 versus 2.9 versus 7.0 seconds. In other words, the highest quality is about nine times slower than the lowest for exactly the same text.
Hence the working habit: draft at 4–8 steps while you are still fixing the text, and render the final take at 24–40. There is no point waiting seven seconds for a version you are going to rewrite anyway.
Pause between sentences (0 – 2 s)
This slider adds exactly the silence it shows at every sentence boundary — at 1 second across three sentences the recording grows by 2 seconds. At 0 you get the model’s natural rhythm. It earns its keep for dictation, language drills, and video where the viewer has to read something on screen.
6. Long text
One run takes up to 1000 characters — roughly 150 words, or about a minute of speech. For anything longer, split by paragraph and render each separately: if one sentence comes out wrong you redo only that piece instead of waiting for the whole text again. Make the pieces end on a full stop rather than mid-sentence, or the melody breaks at the join.
7. A short checklist
- Every sentence ends with a mark.
- Numbers are written the way you want them heard.
- Acronyms are spaced out if they are read letter by letter.
- No tag sits in the middle of a phrase.
- Draft at low steps, final take at high steps.
- The text is under 1000 characters, or split by paragraph.
- The voice matches how long someone will be listening.