Documents to audio
Listen to your documents
Upload a document, check the extracted text, then send it to the studio and choose a voice. Useful for study notes, reports, articles and long-form reading you would rather hear than read.
How it works
- 1
Upload
Add one or more files. VocaTTS pulls out the text page by page.
- 2
Tidy the text
Remove headers, page numbers and anything you do not want read, then apply the text.
- 3
Generate audio
Open the text in Text to Speech, pick a voice and create the narration.
Good uses
Study material
Listen to lecture notes and readings while commuting.
Reports and articles
Review long documents away from the screen.
Drafts
Hear your own writing to catch awkward sentences.
Book chapters
Turn a chapter into audio, one section at a time.
Getting clean narration from a document
- Delete page headers, footers and footnote markers before generating; they otherwise get read aloud.
- Join lines broken in the middle of sentences so the voice does not pause mid-thought.
- For long documents, use multi-segment mode in Text to Speech to get one audio file per section.
- Scanned PDFs are images. Run them through an OCR tool first, then upload the text version.
Supported files
- Formats
- PDF, DOCX, EPUB, TXT, MD, CSV, JSON, RTF and LOG
- File size
- Up to 30 MB per file
- Scanned PDFs
- Need OCR first — there is no text layer to read
- Voices
- Any stock voice, plus your cloned and designed voices
- Length
- Each generation is limited by your plan's characters per request
Frequently Asked Questions
Q: Which file types can I upload?
A: PDF, DOCX, EPUB, TXT, MD, CSV, JSON, RTF and LOG files, up to 30 MB each.
Q: Why does my PDF show no text?
A: It is probably a scanned PDF made of images. Convert it with an OCR tool into a text PDF and upload that version.
Q: Can I listen in my own voice?
A: Yes. After the text opens in Text to Speech, choose your cloned or designed voice under “My voices”.
Q: Can it read a whole book at once?
A: Each generation has a character limit set by your plan. Split long books into chapters or sections and use multi-segment mode.