This guide turns a current search question into a repeatable production decision. It focuses on the source, controls, review, and destination checks that determine whether an output is actually useful.

Quick answer: Check the distribution platform and rights before generating anything. Build a narration bible, clean and segment the manuscript, audition voices on difficult passages, keep a pronunciation ledger, produce chapters in controlled batches, and listen to every exported file in sequence. Master to the destination specification. Do not assume AI narration is permitted: ACX currently requires human narration unless it has specifically authorized otherwise.

Start with the distribution gate

The first audiobook decision is not the voice. It is where the finished work may be distributed and what that platform permits. As of this guide's publication, ACX states that audiobooks must be narrated by a human unless ACX has otherwise authorized the use of text-to-speech or AI narration. Other storefronts, library services, podcast feeds, private learning systems, and direct-sales channels may use different rules.

Read the current agreement, audio specifications, disclosure requirements, and rights terms for every destination. Get permission for the text, translation, voice, music, and sound. Do not clone a narrator, actor, public figure, or private person without explicit authorization. If a platform prohibits the planned method, change the method or destination before spending time on production.

The audiobook production chain

PolicyPlatform and rights pass
BibleVoice, style, names, characters
ScriptClean chapters and markup
AuditionDifficult passage tests
GenerateControlled chapter batches
ListenMeaning and continuity QA
MasterDestination-compliant files

Create a narration bible

The bible preserves decisions across a project that may contain hundreds of pages. Record the intended listener, genre, point of view, pace range, emotional restraint, pronunciation standard, treatment of headings and footnotes, character voices, foreign-language policy, and forbidden mannerisms.

For fiction, define each recurring character with age presentation, energy, rhythm, accent boundaries, and relationship to the narrator. Keep voices distinct but sustainable. A theatrical voice that is impressive for 20 seconds may become exhausting over ten hours. For nonfiction, prioritize trust, clarity, hierarchy, and consistent treatment of lists, quotations, citations, numbers, and abbreviations.

Clean the manuscript for speech

Start from the approved manuscript, not a draft with tracked changes, headers, page numbers, or layout artifacts. Resolve footnotes, URLs, tables, equations, image captions, chapter numbers, and typographic symbols into a spoken policy. Expand ambiguous abbreviations and decide how dates, decimals, currency, units, and initials will be read.

Segment by chapter and scene, then into generation batches small enough to inspect and repair. Keep sentence and paragraph boundaries intact. Save a stable segment ID that links source text, generated audio, revisions, and final chapter position. Never make silent editorial changes during generation; corrections should return to the approved manuscript or a documented errata record.

Audition on the hardest passages

Do not choose a voice using only an easy promotional paragraph. Build an audition pack containing dialogue, narration, long sentences, short emotional lines, numbers, names, acronyms, foreign words, questions, parenthetical phrases, and a quiet passage. Listen on headphones and a phone speaker.

TestListen forReject when
Long sentenceBreath, phrasing, hierarchyThe meaning collapses mid-sentence
DialogueCharacter separation and restraintVoices become caricatures or swap identity
Names and termsRepeatable pronunciationThe same word changes across takes
Quiet proseNatural cadence without filler dramaEvery line receives the same emphasis
Numbers and symbolsCorrect spoken interpretationUnits, dates, or values change meaning

Maintain a pronunciation ledger

List every name, place, invented term, acronym, technical word, and foreign phrase. Store the approved spelling, phonetic cue, stress, language, reference recording when authorized, and the chapters where it appears. Ask a fluent speaker to review languages and names you do not know.

Test pronunciation in the full sentence because neighboring sounds and emphasis change delivery. Once approved, reuse the same instruction. If the model cannot produce a critical term reliably, generate that sentence separately, use an authorized alternate workflow, or change the production plan—never let inconsistency accumulate across chapters.

Generate chapters in controlled batches

Lock the chosen voice, model, pace, style, and pronunciation inputs. Generate one representative chapter before the full book. Review it end to end, revise the bible, then process the remaining chapters in documented batches. Avoid changing settings because one sentence feels different; isolate the sentence-level cause first.

Leave handles around section breaks, but avoid arbitrary silence between every paragraph. Watch for clipped initial consonants, swallowed endings, doubled words, missing lines, repeated phrases, odd breaths, abrupt room tone, and emotional resets at batch boundaries. Compare the audio to the source text while listening.

Run three distinct listening passes

  1. Accuracy pass: follow the manuscript and catch omissions, additions, changed numbers, pronunciation, and speaker errors.
  2. Continuity pass: listen across batch and chapter seams for pace, tone, character, loudness, and ambience changes.
  3. Audience pass: listen without reading, at normal speed, on the devices listeners will use. Mark confusion, fatigue, unnatural emphasis, and places where meaning is hard to follow.

Automated transcription, silence detection, and loudness tools can flag risk. They cannot decide whether a joke lands, a character remains respectful, a technical explanation is understandable, or six hours of delivery feels humanly listenable.

Master to the actual destination

Use the platform's current specification, not a generic “podcast preset.” ACX currently documents requirements including consistent sound, RMS between -23 dB and -18 dB, peaks no higher than -3 dB, a noise floor no higher than -60 dB RMS, 1–5 seconds of room tone at the beginning and end, and 192 kbps or higher constant-bit-rate 44.1 kHz MP3 files. It also requires separate chapter or section files and limits individual files to 120 minutes.

Those values are cited as an example of a destination specification, not permission to submit unauthorized AI narration. Check the current rules immediately before delivery. Export from a clean master, verify filenames and order, scan for clipping and encoding errors, and listen to every final file—not just the editing timeline.

Final authorization and quality checklist

  • The text, voice, language, music, and distribution method are authorized.
  • The destination currently permits the narration method or has explicitly authorized it.
  • The approved manuscript matches every chapter file.
  • Names, terms, numbers, and foreign-language passages follow the ledger.
  • Voice identity, character distinction, pace, and tone remain consistent.
  • All seams, starts, endings, and final exports pass listening QA.
  • Technical files meet the current destination specification.
  • The project retains a bible, source-to-audio lineage, approvals, and corrections.

Use AI narration where the policy allows it

QuestStudio Voice Lab can help audition authorized voices and produce controlled narration tests without a local model setup. Start with one difficult two-minute passage and the pronunciation ledger; do not generate an entire book before the method passes.

For the current ACX boundary and technical delivery details, read ACX's official audio submission requirements. Rules can change, and a platform's primary documentation takes precedence over this workflow.

Frequently asked questions

Does ACX allow AI-narrated audiobooks?

ACX currently says audiobooks must be human-narrated unless ACX has otherwise authorized text-to-speech or AI narration. Check the latest official rules before production.

How should I choose an AI audiobook voice?

Audition difficult passages with dialogue, long sentences, numbers, names, and quiet prose, then judge long-form listenability rather than a short demo.

What is a pronunciation ledger?

It is a project record of approved names, technical terms, stress, language, reference cues, and where each term occurs.

Can automated QA replace listening?

No. Tools can flag transcription, silence, loudness, and clipping risks, but people must review meaning, identity, continuity, and fatigue.

Should I generate the full book at once?

No. Complete and approve one representative chapter, refine the bible, then produce documented batches.