Reference
ACX audio requirements, explained one number at a time
ACX publishes these as a list. What the list does not say is how they are measured, why they conflict with each other, or which of them a piece of software can actually fix. That is what this page is for.
The five numbers
Every bound on its own scale
The tinted stretch is what ACX accepts. Open any card for how it is measured and why files fail it.
Average level (RMS)
−23 to −18 dBFS
−600 dBFS
Average energy across the whole file, not the loudest moment.
Peak
−3 dBFS or lower
−600 dBFS
The single loudest sample. This is the one that makes the level fix hard.
Noise floor
−60 dBFS or lower
−900 dBFS
The room between your words, and the one that is sometimes not fixable.
Room tone at both ends
1 to 5 seconds
0 s10 s
The sound of your room with nobody speaking — not digital silence.
Encoding
44.1 kHz · 192 kbps · CBR · MP3
“Or higher” is compliant. “Constant” is a claim about every frame.
Chapter length
120 minutes maximum
0 min180 min
Not fixable by software, and should not be — where a break belongs is editorial.
The question the list does not answer
Which of these can software actually fix?
| Requirement | Fixable here | Why |
|---|---|---|
| Average level (RMS) | Yes | Gain, but only in the order the other two numbers allow. |
| Peak | Yes | A limiter that sees the peak coming, applied to the signal that gets encoded. |
| Noise floor | Sometimes | Only the gaps between words can be lowered. Noise underneath the voice cannot. |
| Room tone | Yes | Built from your own quietest half second, never from digital silence. |
| Encoding | Yes | Re-encoded to 192 kbps CBR at 44.1 kHz and re-read frame by frame. |
| Chapter length | No | Splitting a chapter is an editorial decision, not an arithmetic one. |
| Human narration | No | ACX prohibits unauthorised text-to-speech. Nothing here changes that. |
The official analyzer
What ACX’s own Audio Lab checks — and what it leaves to humans
The eight metrics it measures
Per the ACX Audio Lab FAQ, the analysis tool checks RMS, peak levels, bitrate, bitrate method, sample rate, mixed channels, duplicate files, and spacing. “Spacing” is ACX’s word for the head and tail silence this page calls room tone, and “bitrate method” is the constant-bitrate requirement. Results arrive as a downloadable spreadsheet, and there is no dispute process for its verdicts.
What it does not measure
The same FAQ states the Audio Lab does not check noise floor and does not catch editing errors. It also says passing analysis does not guarantee approval. The checker on this site’s home page measures the noise floor from your file’s own bytes, which is exactly the gap the official tool leaves open.
ACX also answers why your DAW’s RMS differs from theirs: RMS can be measured in multiple ways, so meters legitimately disagree. This site measures whole-file RMS in dBFS and applies the same instrument to its own output.
Primary source
What ACX says, in its own words
Volume is between -23dB and -18dB RMS: Each file needs to fall between the specific volume range of -23dB and -18dB RMS for consistent volume.
Peak levels are less than -3dB: Each file must have peak values no higher than -3dB to avoid distortion.
Noise floor is less than -60dB RMS: Each file must have a noise floor no higher than -60dB RMS to avoid background noise distractions.
Keep room tone less than 5 seconds: We recommend between 1 and 5 seconds of room tone at the beginning and end of each file for an ideal listening experience. Room tone spacing must not exceed 5 seconds.
File format is 192 kbps or higher CBR: Each file must be a 192 kbps or higher CBR, 44.1kHz MP3.
Files are in either mono or stereo: All files must be in the same channel format.
Each file should be no longer than 120 minutes.
Source: ACX audio submission requirements, read 9 August 2026.
Average level: between −23 and −18 dBFS
RMS is the root mean square of every sample in the file — the average energy, not the loudest moment. It is a single number for the whole chapter, so a passage of shouting cannot be balanced out by a whisper somewhere else; both move the average.
It is not loudness in the broadcast sense. If you have used a LUFS meter, that measurement is weighted by frequency and gated to ignore silence, and it will not give you this number. ACX asks for plain RMS in decibels relative to full scale.
Why files fail: recorded conservatively to avoid clipping, which is good practice and leaves the file below −23 dB. Fixable: yes, but see the peak below.
Peak: −3 dBFS or lower
The single loudest sample in the file. Almost always a plosive — a “p” or “b” hitting the microphone — or a chair, or a page.
Why this one matters more than it looks: it is what makes the level fix hard. A chapter at −27 dB RMS needs about six decibels. If the loudest plosive already sits at −4 dBFS, six decibels puts it at +2 dBFS: clipped, and well over ACX’s ceiling. Gain alone cannot satisfy both bounds, which is why “normalise” fails this specification so often. What is needed is a limiter that sees the peak coming and reduces gain before it arrives.
Noise floor: −60 dBFS or lower
The level of the room between your words. Not the quietest single sample — speech crosses zero constantly — but the level of a sustained quiet passage.
The interaction nobody mentions: every decibel of gain you add to reach the RMS window raises the noise floor by exactly the same decibel. A file with a −62 dB floor passes today and fails after the six decibels it needs. So the floor has to be corrected against the level it must reach before the gain, not after it.
What can be fixed, and where it stops: a gate lowers a measured floor by pulling down the gaps between words, and that is all a gate can do — it never touches the noise sitting underneath your voice. Subtracting a measured noise spectrum does reach under the voice, and this tool does that when gating alone cannot: it profiles the room from your file’s own quietest half second and subtracts it, capped at 12 decibels. The cap is not timidity. Past roughly that much, subtraction leaves the warbling artefacts people describe as underwater, and a chapter that passes a meter while sounding processed fails the human review it was supposed to survive.
When it still cannot be fixed: below about 18 dB of separation between your speech and your room, twelve decibels does not bridge the gap and nothing else will either. That is a re-record, and this tool says so rather than selling you an export. Either way the corrected file is re-measured before it is offered, so the answer you get is measured rather than promised.
44.1 kHz, 192 kbps or higher, constant bit rate, MP3
Note the words or higher. A 256 kbps constant-bitrate file at 44.1 kHz is compliant, and any tool that reports it as a failure is wrong.
“Constant” is a claim about every frame in the file, not about the average. A variable-bitrate export can hold one bitrate through a quiet passage for tens of seconds, so checking the first frames proves nothing. This tool reads every frame header in the file.
Room tone at both ends, at most five seconds
Room tone is the sound of your recording space with nobody speaking. ACX requires no more than five seconds at each end, and recommends between one and five. This tool builds one second when a file has none, which sits inside that recommendation.
The mistake that looks like a fix: padding the file with silence. Digital silence is not room tone. The noise floor drops to nothing and then the room arrives with your first word, which a listener hears as a click and a reviewer hears as an edit. Tone has to be built from the recording’s own quiet passage.
120 minutes maximum
This one is not fixable by software, and should not be: splitting a chapter is an editorial decision about where a break belongs, not an arithmetic one.
One requirement no software can help with
ACX prohibits unauthorised text-to-speech, AI and automated recordings. This tool masters a recording of a human voice; it cannot make anything else acceptable, and nothing here should be read as suggesting otherwise.
What this tool does not check
Consistency across the whole book, extraneous sounds, opening and closing credits, the retail sample, and whether the section header is read aloud. Those are judged by a person, and no measurement here can stand in for that. Passing all five numbers means the machine will not reject your file before a human hears it — not that the book is finished.
Measure it