Chingie, the Studio Chingie mascot

Studio Chingie

TTS Talking Book

Working sample

Turns an EPUB into a DAISY digital talking book read by a synthetic voice, built to the Library of Congress NLS construction spec. Every file is checked against the standard, clip timing is measured from the audio, and the narration is transcribed back to measure how clear it is.

Listen to the sample
AccessibilitySpeech synthesisDAISY / Z39.86Measured qualityOpen source
Sample talking book of the U.S. Constitution: 30 minutes of narration, 31 navigation points, 170 clips, 30 of 30 checks passedCurrent build
The U.S. Constitution as a talking book, built from an EPUB with no manual steps. All 30 checks pass.

Development details

Source code
Public

Built with

PythonKokoro-82M (local)whisper.cppffmpegxmllint

Screenshots

Navigation tree with Label and Play from here buttons for each article and section
Current build

Every navigation point can be played: its spoken label from the headings file, or the book from that point.

Table of 30 passing validation checks
Current build

30 checks run on the finished files, including offline DTD validation and clip timing measured from the audio.

Intelligibility table with word error rate per track, and the pronunciation lexicon
Current build

Each track is transcribed back and compared with the text the voice was given: 1.2% word error rate overall.

What it does

Reads an EPUB, finds its chapters and sections, narrates it with an open-weight voice, and packages a DAISY talking book: audio, a headings file, and the OPF, NCX, SMIL, and checksum files that players use to navigate.

Built to the spec

File names, navigation classes, checksums, and clip timing follow NLS Specification 1203. Each clip starts 100 ms before the voice and ends 200 ms after it, inside the spec's windows, and the validator measures this from the audio.

Measured

The sample is the U.S. Constitution: 29.6 minutes of narration, built in about five minutes on a laptop CPU. All 30 checks pass, and transcribing the audio back gives a 1.2% word error rate.

No license fees

The voice is Kokoro-82M under the Apache 2.0 license, running locally, with no per-use fee or outside service. NLS delivery would also need AMR-WB+ audio and PDTB2 encryption, which the sample does not do yet.