reddex/saved/pokedex_voice_spike_scripts/README.md
forgejoadmin 71db2d1ab9 Initial commit
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-18 03:16:42 -04:00

3 KiB

Pokedex voice spike — scripts

De-risking spike for the Pokédex app's robotic voice output. Full pipeline: text -> TTS -> optional DSP "robot" filter -> mp3.

Files

  • generate_samples.sh — end-to-end driver: installs are documented at the top, then runs espeak-ng and Piper TTS, then the DSP filters, then converts everything to mp3. Run it from this directory (./generate_samples.sh).
  • robot_filter.py — heavy/original robot filter: square-wave ring modulation, bitcrush (sample-and-hold + bit-depth reduction), narrow band-pass, hard drive. This is what "too robotic" was generated with.
  • robot_filter_light.py — barely-there version: sine-wave ring mod at low mix, wide band-pass, mild soft-clip. Meant to add a hint of synthesized edge to an already-natural neural voice without disturbing cadence.
  • robot_filter_v2.py — the current one: parameterized, takes an intensity argument from 0.0 (untouched) to 1.0 (full heavy filter) and interpolates ring-mod mix/carrier, bandpass width, bitcrush, and drive along that scale. Usage: python3 robot_filter_v2.py in.wav out.wav 0.45. This is the one to keep tuning going forward — texted intensity numbers ("try 0.6") map directly to its third argument.

Voice model

TTS engine is Piper — a small, fully offline neural TTS with real Android ports available, which matters for the app's "works with no signal" requirement.

The voice used is en_US-joe-medium.onnx, fine-tuned from Piper's stock "lessac" voice on a CC0-licensed dataset (see MODEL_CARD if you unpack the wheel) — no attribution/licensing issue to ship it.

How I got the model file in this sandbox: the sandbox's network allowlist didn't reach huggingface.co, where Piper's official voices are hosted, only package registries. I found a community-published PyPI wheel (joe-us-piper-voice) that bundles the .onnx file directly, so a plain pip install pulled it down. That's a workaround specific to this sandbox — for the real project, get voices the normal way from Piper's official releases/voice list, which gives you far more voice choices (different speakers, accents, quality tiers) than this one bundled option.

What's still open

  • Pronunciation of individual Pokémon species names hasn't been systematically checked — only "Pokémon," "Bulbasaur," "Charizard," "Pikachu" were tested. Expect to need a pass listening to all ~1000+ species names and hand-fixing the ones the phonemizer (espeak-ng, under Piper's hood) gets wrong, either by respelling in the source text or via espeak's phoneme-override escape syntax.
  • Robot filter intensity (robot_filter_v2.py's intensity arg) needs to land wherever the actual desired "amount of robotic" ends up — currently parked at 0.45 pending feedback.
  • This whole pipeline is meant to run once, offline, over every dex entry ahead of time (batch pre-generation), not live on-device — see the earlier discussion in this conversation for why.