Text to speech
Type or paste your text, or drop in a PDF, a Word file or a photo of a page. Pick a voice, listen, and download it as an MP3. The speech is generated on your own device — your document never leaves your computer.
How it works
A speech model runs inside your browser and turns your writing into a voice — nothing is uploaded, so a contract, a diagnosis or a private letter stays on your machine. A PDF is read straight from its text layer, a Word file from its own contents, and a photo of a page goes through our OCR first, so a book or a printed letter works too. Numbers, amounts and symbols are spelled out before speaking, so "€25" is read as "twenty-five euro" rather than swallowed. The first time you press Speak, the voice engine downloads once (about 390 MB) and your browser keeps it after that.
Which languages can it speak?
English only, for now. The voice model we use is English, and every multilingual model we tested either forbids commercial use or carries a copyleft licence we will not put on a paid site. We would rather ship one language honestly than a dozen we have no right to. Other languages need their own licensing round.
Can I use the audio commercially?
Yes. The voice model is MIT-licensed and the voices come from a speech database that permits commercial use, so a video, a podcast or an audiobook is fine. The voices are synthetic and belong to nobody, so you are not impersonating a real person either.
Why does it sound flat on long texts?
The model reads sentence by sentence, without knowing the story, so it cannot build up drama the way a human narrator does. Short paragraphs with clear punctuation sound noticeably better than one long block — commas and full stops are what it uses to breathe.
Can I change the voice afterwards?
Yes — the pitch and speed sliders here cover most of it, and speed is a true time-stretch, so a slower voice does not turn into a deeper voice. For more, download the MP3 and run it through our voice changer, or drop it straight into the audio studio.
How do abbreviations and unusual names get pronounced?
Common abbreviations without vowels — PDF, XML, HTML, GDPR — are spelled out letter by letter automatically, and combinations like MP3 or 4K come out right on their own. Capital-letter words that contain vowels are deliberately left as words, so NASA and a sentence typed in capitals still sound normal. For anything that still comes out wrong, write it the way it should sound: the voice reads exactly what you type, so an unusual name or brand can simply be spelled phonetically — "Nike" as "Nykee", "Cicek" as "Cheechek" — until it sounds right.
