TTS Talking Book
Working sample
Turns an EPUB into a DAISY digital talking book read by a synthetic voice, built to the Library of Congress NLS construction spec. Every file is checked against the standard, clip timing is measured from the audio, and the narration is transcribed back to measure how clear it is.
Listen to the sample
Current buildDevelopment details
- Source code
- Public
Built with
Screenshots

Every navigation point can be played: its spoken label from the headings file, or the book from that point.

30 checks run on the finished files, including offline DTD validation and clip timing measured from the audio.

Each track is transcribed back and compared with the text the voice was given: 1.2% word error rate overall.
What it does
Reads an EPUB, finds its chapters and sections, narrates it with an open-weight voice, and packages a DAISY talking book: audio, a headings file, and the OPF, NCX, SMIL, and checksum files that players use to navigate.
Built to the spec
File names, navigation classes, checksums, and clip timing follow NLS Specification 1203. Each clip starts 100 ms before the voice and ends 200 ms after it, inside the spec's windows, and the validator measures this from the audio.
Measured
The sample is the U.S. Constitution: 29.6 minutes of narration, built in about five minutes on a laptop CPU. All 30 checks pass, and transcribing the audio back gives a 1.2% word error rate.
No license fees
The voice is Kokoro-82M under the Apache 2.0 license, running locally, with no per-use fee or outside service. NLS delivery would also need AMR-WB+ audio and PDTB2 encryption, which the sample does not do yet.