August 18, 2026
How to Turn Your Ebook into an Audiobook (Without a Studio)
Audio is the format most self-published authors skip, and it's usually not a decision — it's a stall. The manuscript is finished, the ebook is live, and then someone asks whether there's an audio version and the honest answer is "I looked into it and it seemed like a lot."
It is a lot, traditionally. But the reason to look again isn't that audiobooks are trendy: it's that an audiobook is the same content sold a second time to a group of people who were never going to read the text version. Commuters, walkers, people who consume non-fiction almost exclusively while doing something else. Skipping audio doesn't just cost you a format — it costs you an audience segment entirely.
The three routes, honestly compared
Hire a professional narrator. Best possible result, and for fiction — where performance, character voices, and pacing carry enormous weight — it remains meaningfully better than anything else. It's also the expensive route, typically priced per finished hour, plus a casting process, plus turnaround measured in weeks. For a first non-fiction book with unproven sales, the economics are hard to justify.
Narrate it yourself. Cheapest in money, most expensive in everything else. A decent microphone and a quiet room get you further than people expect, but the real cost is time: recording, re-recording every fluffed sentence, editing out breaths and room noise, matching levels across sessions recorded on different days. Expect several hours of work per finished hour of audio, and expect to hate chapter four by the third take.
Synthetic narration. Fast and cheap, and — this is the part that changed — no longer obviously robotic for straightforward prose. Modern text-to-speech handles clear declarative non-fiction well. It handles emotional fiction, distinct character voices, and heavy irony far less well.
The reasonable heuristic: practical non-fiction is where synthetic narration is genuinely good enough today. Literary fiction and memoir are where a human still wins by a wide margin.
What "good enough" actually depends on
Not all manuscripts narrate equally well, and this has less to do with the voice engine than with the text. Before generating anything, look at your manuscript for the things that read fine on a page and fall apart in audio:
- Tables and data grids. A table read aloud cell by cell is unlistenable. Either summarize the finding in prose or accept that the table is a text-only element.
- Long URLs and code. Same problem. Reference them by name ("the checklist on my site"), not by character-by-character reading.
- Dense bullet lists. Two or three items are fine spoken; a fourteen-item list is not. Convert long lists into prose or split them across paragraphs.
- Heavy cross-references. "As we saw in the table above" makes no sense to a listener. "Earlier" works.
- Footnotes and asides. They interrupt the spoken flow far more violently than the written one.
Fixing these is a manuscript edit, not an audio problem — and it's worth doing even if you never produce audio, because the same changes generally make the text clearer too.
Producing the audio
In Ebook Creator, audio generation runs from the export step on a finished book. What comes out is a package containing one MP3 per chapter plus a single chaptered .m4b file for the whole book — the format audiobook players use for chapter navigation, so listeners can jump between chapters and resume where they left off rather than scrubbing through one enormous file.
A few practical notes on how it works:
- The manuscript's Markdown is converted to speakable text first — headings, emphasis markers, tables, and callout boxes are stripped or dropped rather than read aloud as literal punctuation.
- Narration is available in the same six languages the app writes in, so a translated edition can have its own audio rather than being read with the wrong accent.
- It's charged per chapter (2 credits per chapter), unlike the text export formats, which are free and unlimited. Text-to-speech has a real per-character cost, so it's priced like chapter generation rather than bundled in.
- There's one narrator voice — a neutral, calm reader. Choosing between voices isn't something the app offers today, so if a specific vocal character is central to your book, that's an argument for the human route.
Generation runs in the background and you're emailed when it's ready — a full-length book is a lot of audio, and it isn't instant.
Before you upload it anywhere
Producing the files is the easy half. Distribution has its own requirements, and this is where authors get surprised.
Retailer specs are strict and specific. Audiobook distributors — Audible/ACX, Findaway Voices, Kobo, Apple — each publish technical requirements covering file format, bitrate, sample rate, peak and RMS loudness levels, and room tone. These are mastering specs, and meeting them may require a pass through audio software before submission. Read your chosen retailer's current spec sheet rather than assuming any generated file drops straight in.
Disclosure rules exist and they change. Several platforms now require you to declare whether narration is synthetic, and some have distinct policies or catalog placement for AI-narrated titles. Check the current policy of whichever platform you're submitting to — this area has moved fast and any advice more than a few months old should be treated as stale.
Listen to the whole thing at least once. Not skimmed — actually listened to. Text-to-speech mispronounces proper nouns, technical jargon, acronyms, and anything ambiguous between a noun and a verb. Names of people and places are the most common offenders. If a term is mangled, adjusting the spelling in the manuscript ahead of narration is usually the fastest fix.
Consider a sample-first approach. Generate audio for one representative chapter, listen critically, and decide whether the result meets your standard before committing to the full book. If it doesn't, you've learned that cheaply.
When it's worth it
Audio makes the most sense when your book is practical, your chapters are clean prose, and your readers are the kind of people who listen while doing other things — commuting, exercising, driving. It makes the least sense when the writing depends on voice, performance, or emotional delivery, or when your book leans heavily on tables, code, and visual reference material that simply doesn't survive the trip to audio.
The middle ground is real, too: some authors produce synthetic narration for a first book to test whether audio demand exists at all, then commission a human narration for the next one — or for the same book, once the sales data justifies it. Producing audio isn't a permanent commitment.
If your manuscript is finished and you're deciding what formats to publish in, the KDP self-publishing checklist covers the text side, and ebook marketing strategies covers what to do once it's live. And if the manuscript isn't finished yet, start writing — new accounts get 4 free credits, enough to outline and write a complete short book.