Initial commit
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
commit
71db2d1ab9
126 changed files with 3198 additions and 0 deletions
57
saved/pokedex_voice_spike_scripts/README.md
Normal file
57
saved/pokedex_voice_spike_scripts/README.md
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
# Pokedex voice spike — scripts
|
||||
|
||||
De-risking spike for the Pokédex app's robotic voice output. Full pipeline:
|
||||
text -> TTS -> optional DSP "robot" filter -> mp3.
|
||||
|
||||
## Files
|
||||
|
||||
- `generate_samples.sh` — end-to-end driver: installs are documented at the
|
||||
top, then runs espeak-ng and Piper TTS, then the DSP filters, then
|
||||
converts everything to mp3. Run it from this directory (`./generate_samples.sh`).
|
||||
- `robot_filter.py` — heavy/original robot filter: square-wave ring
|
||||
modulation, bitcrush (sample-and-hold + bit-depth reduction), narrow
|
||||
band-pass, hard drive. This is what "too robotic" was generated with.
|
||||
- `robot_filter_light.py` — barely-there version: sine-wave ring mod at
|
||||
low mix, wide band-pass, mild soft-clip. Meant to add a hint of
|
||||
synthesized edge to an already-natural neural voice without disturbing
|
||||
cadence.
|
||||
- `robot_filter_v2.py` — the current one: parameterized, takes an
|
||||
`intensity` argument from 0.0 (untouched) to 1.0 (full heavy filter) and
|
||||
interpolates ring-mod mix/carrier, bandpass width, bitcrush, and drive
|
||||
along that scale. Usage: `python3 robot_filter_v2.py in.wav out.wav 0.45`.
|
||||
This is the one to keep tuning going forward — texted intensity numbers
|
||||
("try 0.6") map directly to its third argument.
|
||||
|
||||
## Voice model
|
||||
|
||||
TTS engine is [Piper](https://github.com/OHF-Voice/piper1-gpl) — a small,
|
||||
fully offline neural TTS with real Android ports available, which matters
|
||||
for the app's "works with no signal" requirement.
|
||||
|
||||
The voice used is `en_US-joe-medium.onnx`, fine-tuned from Piper's stock
|
||||
"lessac" voice on a CC0-licensed dataset (see MODEL_CARD if you unpack the
|
||||
wheel) — no attribution/licensing issue to ship it.
|
||||
|
||||
**How I got the model file in this sandbox:** the sandbox's network
|
||||
allowlist didn't reach huggingface.co, where Piper's official voices are
|
||||
hosted, only package registries. I found a community-published PyPI wheel
|
||||
(`joe-us-piper-voice`) that bundles the .onnx file directly, so a plain
|
||||
`pip install` pulled it down. That's a workaround specific to *this*
|
||||
sandbox — for the real project, get voices the normal way from Piper's
|
||||
official releases/voice list, which gives you far more voice choices
|
||||
(different speakers, accents, quality tiers) than this one bundled option.
|
||||
|
||||
## What's still open
|
||||
|
||||
- Pronunciation of individual Pokémon species names hasn't been
|
||||
systematically checked — only "Pokémon," "Bulbasaur," "Charizard,"
|
||||
"Pikachu" were tested. Expect to need a pass listening to all ~1000+
|
||||
species names and hand-fixing the ones the phonemizer (espeak-ng, under
|
||||
Piper's hood) gets wrong, either by respelling in the source text or via
|
||||
espeak's phoneme-override escape syntax.
|
||||
- Robot filter intensity (`robot_filter_v2.py`'s `intensity` arg) needs to
|
||||
land wherever the actual desired "amount of robotic" ends up — currently
|
||||
parked at 0.45 pending feedback.
|
||||
- This whole pipeline is meant to run once, offline, over every dex entry
|
||||
ahead of time (batch pre-generation), not live on-device — see the
|
||||
earlier discussion in this conversation for why.
|
||||
Loading…
Add table
Add a link
Reference in a new issue