Initial commit

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
forgejoadmin 2026-08-18 03:16:42 -04:00
commit 71db2d1ab9
126 changed files with 3198 additions and 0 deletions

View file

@ -0,0 +1,57 @@
# Pokedex voice spike — scripts
De-risking spike for the Pokédex app's robotic voice output. Full pipeline:
text -> TTS -> optional DSP "robot" filter -> mp3.
## Files
- `generate_samples.sh` — end-to-end driver: installs are documented at the
top, then runs espeak-ng and Piper TTS, then the DSP filters, then
converts everything to mp3. Run it from this directory (`./generate_samples.sh`).
- `robot_filter.py` — heavy/original robot filter: square-wave ring
modulation, bitcrush (sample-and-hold + bit-depth reduction), narrow
band-pass, hard drive. This is what "too robotic" was generated with.
- `robot_filter_light.py` — barely-there version: sine-wave ring mod at
low mix, wide band-pass, mild soft-clip. Meant to add a hint of
synthesized edge to an already-natural neural voice without disturbing
cadence.
- `robot_filter_v2.py` — the current one: parameterized, takes an
`intensity` argument from 0.0 (untouched) to 1.0 (full heavy filter) and
interpolates ring-mod mix/carrier, bandpass width, bitcrush, and drive
along that scale. Usage: `python3 robot_filter_v2.py in.wav out.wav 0.45`.
This is the one to keep tuning going forward — texted intensity numbers
("try 0.6") map directly to its third argument.
## Voice model
TTS engine is [Piper](https://github.com/OHF-Voice/piper1-gpl) — a small,
fully offline neural TTS with real Android ports available, which matters
for the app's "works with no signal" requirement.
The voice used is `en_US-joe-medium.onnx`, fine-tuned from Piper's stock
"lessac" voice on a CC0-licensed dataset (see MODEL_CARD if you unpack the
wheel) — no attribution/licensing issue to ship it.
**How I got the model file in this sandbox:** the sandbox's network
allowlist didn't reach huggingface.co, where Piper's official voices are
hosted, only package registries. I found a community-published PyPI wheel
(`joe-us-piper-voice`) that bundles the .onnx file directly, so a plain
`pip install` pulled it down. That's a workaround specific to *this*
sandbox — for the real project, get voices the normal way from Piper's
official releases/voice list, which gives you far more voice choices
(different speakers, accents, quality tiers) than this one bundled option.
## What's still open
- Pronunciation of individual Pokémon species names hasn't been
systematically checked — only "Pokémon," "Bulbasaur," "Charizard,"
"Pikachu" were tested. Expect to need a pass listening to all ~1000+
species names and hand-fixing the ones the phonemizer (espeak-ng, under
Piper's hood) gets wrong, either by respelling in the source text or via
espeak's phoneme-override escape syntax.
- Robot filter intensity (`robot_filter_v2.py`'s `intensity` arg) needs to
land wherever the actual desired "amount of robotic" ends up — currently
parked at 0.45 pending feedback.
- This whole pipeline is meant to run once, offline, over every dex entry
ahead of time (batch pre-generation), not live on-device — see the
earlier discussion in this conversation for why.

View file

@ -0,0 +1,77 @@
"""
Windows-friendly regeneration of the voice spike, extended to all 5
Pokemon the app currently knows about. Ports generate_samples.sh's
"fixed" (cadence-corrected) + robot_filter_v2.py (intensity 0.45) stages
to Python, using piper's Python API directly instead of the `piper` CLI
and skipping the ffmpeg mp3 step (WAV plays fine on Android, and ffmpeg
isn't installed on this machine).
Usage: .venv\\Scripts\\python.exe generate_all.py
"""
import wave
from pathlib import Path
from piper import PiperVoice
from piper.config import SynthesisConfig
import robot_filter_v2
HERE = Path(__file__).parent
VOICE_MODEL = HERE / "voices" / "en_US-joe-medium.onnx"
OUT_DIR = HERE / "output"
ENTRIES = {
"bulbasaur": (
"Bulbasaur, the seed Pokémon. It can be seen napping in bright "
"sunlight. There is a seed on its back. By soaking up the sun's "
"rays, the seed grows progressively larger."
),
"charizard": (
"Charizard, the flame Pokémon. Charizard flies around the sky in "
"search of powerful opponents. It breathes fire of such great "
"heat that it melts anything."
),
"pikachu": (
"Pikachu, the mouse Pokémon. When several of these Pokémon "
"gather, their electricity could build and cause lightning "
"storms."
),
"eevee": (
"Eevee, the evolution Pokémon. Its genetic code is irregular. It "
"may mutate if it is exposed to radiation from element stones."
),
"squirtle": (
"Squirtle, the tiny turtle Pokémon. After birth, its back swells "
"and hardens into a shell. It powerfully sprays foam from its "
"mouth."
),
}
ROBOT_INTENSITY = 0.45
SYN_CONFIG = SynthesisConfig(noise_scale=0.5, noise_w_scale=0.3, length_scale=0.98)
def main():
OUT_DIR.mkdir(exist_ok=True)
print(f"Loading voice model {VOICE_MODEL.name}...")
voice = PiperVoice.load(str(VOICE_MODEL))
for name, text in ENTRIES.items():
fixed_path = OUT_DIR / f"{name}_fixed.wav"
robot_path = OUT_DIR / f"{name}_fixed_robot.wav"
print(f"Synthesizing {name}...")
with wave.open(str(fixed_path), "wb") as wav_file:
voice.synthesize_wav(text, wav_file, syn_config=SYN_CONFIG)
sr, x = robot_filter_v2.load(str(fixed_path))
y = robot_filter_v2.robotize(x, sr, intensity=ROBOT_INTENSITY)
robot_filter_v2.save(str(robot_path), sr, y)
print(f" -> {robot_path.name}")
print("\nDone. See output/*_fixed_robot.wav")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,74 @@
#!/usr/bin/env bash
# Reproduces the Pokedex voice spike end to end:
# TTS (espeak-ng and Piper) -> robot-voice DSP filter -> mp3
#
# Setup (Debian/Ubuntu):
# sudo apt-get install -y espeak-ng ffmpeg
# pip install piper-tts --break-system-packages
# pip install numpy scipy --break-system-packages
#
# Voice model:
# Official route: download a Piper voice (.onnx + .onnx.json) from the
# Piper voices repo and point --model at it, e.g. en_US-lessac-medium.
# In this sandbox, huggingface.co wasn't reachable, so I instead pulled a
# community PyPI wheel that bundles the model file directly:
# pip download --no-deps joe-us-piper-voice
# (unzip the wheel; the .onnx/.onnx.json live under
# joe_us_piper_voice/data/). Not an official source -- for production,
# get voices from Piper's own releases instead.
set -euo pipefail
cd "$(dirname "$0")"
VOICE_MODEL="voices/en_US-joe-medium.onnx"
export ALSA_CONFIG_PATH=/dev/null # silence ALSA warnings in headless envs
# ---- Stage 1: raw espeak-ng (formant synth, offline, inherently "robotic") ----
espeak-ng -v en-us -s 150 -w bulbasaur_raw.wav \
"Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
espeak-ng -v en-us -s 150 -w charizard_raw.wav \
"Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
espeak-ng -v en-us -s 150 -w pikachu_raw.wav \
"Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
# ---- Stage 2: Piper neural TTS, first pass (natural but bad cadence on rare words) ----
gen_neural() {
local name="$1" text="$2"
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_neural.wav"
}
gen_neural bulbasaur "Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
gen_neural charizard "Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
gen_neural pikachu "Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
# ---- Stage 3: Piper neural TTS, cadence-fixed pass ----
# Fix = (a) standard "X, the Y Pokemon." phrasing instead of two short
# sentences, which stopped "Pokemon" from landing phrase-final where
# duration models over-lengthen it, and (b) reduced noise-w-scale (duration
# randomness) and noise-scale (audio variance) so rare proper nouns don't
# get a random stretched/warped rendering.
gen_fixed() {
local name="$1" text="$2"
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_fixed.wav" \
--noise-scale 0.5 --noise-w-scale 0.3 --length-scale 0.98
}
gen_fixed bulbasaur "Bulbasaur, the seed Pokémon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
gen_fixed charizard "Charizard, the flame Pokémon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
gen_fixed pikachu "Pikachu, the mouse Pokémon. When several of these Pokémon gather, their electricity could build and cause lightning storms."
# ---- Stage 4: robot-voice DSP filter passes ----
# robot_filter.py -> heavy/original (square-wave ring mod + bitcrush + narrow bandpass)
# robot_filter_light.py -> barely-there (sine ring mod, wide bandpass, mild drive)
# robot_filter_v2.py -> parameterized 0..1 intensity knob (used at 0.45 = "medium")
for name in bulbasaur charizard pikachu; do
python3 robot_filter.py "${name}_raw.wav" "${name}_robot.wav"
python3 robot_filter_light.py "${name}_neural.wav" "${name}_neural_light.wav"
python3 robot_filter_v2.py "${name}_fixed.wav" "${name}_fixed_robot.wav" 0.45
done
# ---- Stage 5: mp3 for easy playback/delivery ----
for f in *_raw.wav *_robot.wav *_neural.wav *_neural_light.wav *_fixed.wav *_fixed_robot.wav; do
[ -f "$f" ] && ffmpeg -y -loglevel error -i "$f" -codec:a libmp3lame -qscale:a 4 "${f%.wav}.mp3"
done
echo "Done. See *.mp3 for output."

View file

@ -0,0 +1,79 @@
"""
Quick-and-dirty 'robot voice' post-processor for the Pokedex voice spike.
Pipeline: TTS wav in -> ring modulation + bitcrush + band-pass (speaker-grille
coloring) + light hard-clip drive -> wav out.
Usage: python3 robot_filter.py in.wav out.wav
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
elif data.dtype == np.uint8:
x = (data.astype(np.float64) - 128) / 128.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz=45.0, mix=0.55):
t = np.arange(len(x)) / sr
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) # square carrier = harsher/metallic
modulated = x * carrier
return (1 - mix) * x + mix * modulated
def bitcrush(x, bit_depth=6, sr=None, downsample_factor=3):
# sample-and-hold downsample (classic lo-fi/robotic stepping)
if downsample_factor > 1:
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
if len(held) < len(x):
held = np.pad(held, (0, len(x) - len(held)))
x = held
levels = 2 ** bit_depth
x = np.round(x * levels) / levels
return x
def bandpass(x, sr, low=350.0, high=3400.0, order=4):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount=1.8):
return np.tanh(x * amount) / np.tanh(amount)
def robotize(x, sr):
y = ring_modulate(x, sr, carrier_hz=45.0, mix=0.5)
y = bitcrush(y, bit_depth=7, downsample_factor=2)
y = bandpass(y, sr, low=300, high=3800)
y = drive(y, amount=1.6)
# normalize
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
sr, x = load(src)
y = robotize(x, sr)
save(dst, sr, y)
print(f"wrote {dst}")

View file

@ -0,0 +1,58 @@
"""
Light-touch robot coloring for a neural (already-natural-sounding) TTS voice.
Goal: keep natural cadence/prosody intact, just add a synthesized/device edge.
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz=90.0, mix=0.12):
t = np.arange(len(x)) / sr
carrier = np.sin(2 * np.pi * carrier_hz * t) # sine carrier = subtler than square
return (1 - mix) * x + mix * (x * carrier)
def bandpass(x, sr, low=200.0, high=5500.0, order=2):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount=1.15):
return np.tanh(x * amount) / np.tanh(amount)
def robotize_light(x, sr):
y = ring_modulate(x, sr, carrier_hz=90.0, mix=0.12)
y = bandpass(y, sr, low=200, high=5500)
y = drive(y, amount=1.15)
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
sr, x = load(src)
y = robotize_light(x, sr)
save(dst, sr, y)
print(f"wrote {dst}")

View file

@ -0,0 +1,72 @@
"""
Parameterized robot-voice filter, v2.
intensity in [0, 1]: 0 = untouched, 1 = full "heavy" robot from round 1.
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz, mix, square=False):
t = np.arange(len(x)) / sr
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) if square else np.sin(2 * np.pi * carrier_hz * t)
return (1 - mix) * x + mix * (x * carrier)
def bitcrush(x, bit_depth, downsample_factor):
if downsample_factor > 1:
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
if len(held) < len(x):
held = np.pad(held, (0, len(x) - len(held)))
x = held
levels = 2 ** bit_depth
return np.round(x * levels) / levels
def bandpass(x, sr, low, high, order=3):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount):
return np.tanh(x * amount) / np.tanh(amount)
def robotize(x, sr, intensity=0.45):
i = intensity
y = ring_modulate(x, sr, carrier_hz=60 - 15 * i, mix=0.1 + 0.45 * i, square=(i > 0.6))
if i > 0.35:
y = bitcrush(y, bit_depth=int(12 - 6 * i), downsample_factor=1 if i < 0.7 else 2)
y = bandpass(y, sr, low=250 - 100 * i, high=6000 - 2200 * i)
y = drive(y, amount=1.0 + 0.9 * i)
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
intensity = float(sys.argv[3]) if len(sys.argv) > 3 else 0.45
sr, x = load(src)
y = robotize(x, sr, intensity=intensity)
save(dst, sr, y)
print(f"wrote {dst} (intensity={intensity})")

View file

@ -0,0 +1,484 @@
{
"audio": {
"sample_rate": 22050,
"quality": "medium"
},
"espeak": {
"voice": "en-us"
},
"inference": {
"noise_scale": 0.667,
"length_scale": 1,
"noise_w": 0.8
},
"phoneme_type": "espeak",
"phoneme_map": {},
"phoneme_id_map": {
"_": [
0
],
"^": [
1
],
"$": [
2
],
" ": [
3
],
"!": [
4
],
"'": [
5
],
"(": [
6
],
")": [
7
],
",": [
8
],
"-": [
9
],
".": [
10
],
":": [
11
],
";": [
12
],
"?": [
13
],
"a": [
14
],
"b": [
15
],
"c": [
16
],
"d": [
17
],
"e": [
18
],
"f": [
19
],
"h": [
20
],
"i": [
21
],
"j": [
22
],
"k": [
23
],
"l": [
24
],
"m": [
25
],
"n": [
26
],
"o": [
27
],
"p": [
28
],
"q": [
29
],
"r": [
30
],
"s": [
31
],
"t": [
32
],
"u": [
33
],
"v": [
34
],
"w": [
35
],
"x": [
36
],
"y": [
37
],
"z": [
38
],
"æ": [
39
],
"ç": [
40
],
"ð": [
41
],
"ø": [
42
],
"ħ": [
43
],
"ŋ": [
44
],
"œ": [
45
],
"ǀ": [
46
],
"ǁ": [
47
],
"ǂ": [
48
],
"ǃ": [
49
],
"ɐ": [
50
],
"ɑ": [
51
],
"ɒ": [
52
],
"ɓ": [
53
],
"ɔ": [
54
],
"ɕ": [
55
],
"ɖ": [
56
],
"ɗ": [
57
],
"ɘ": [
58
],
"ə": [
59
],
"ɚ": [
60
],
"ɛ": [
61
],
"ɜ": [
62
],
"ɞ": [
63
],
"ɟ": [
64
],
"ɠ": [
65
],
"ɡ": [
66
],
"ɢ": [
67
],
"ɣ": [
68
],
"ɤ": [
69
],
"ɥ": [
70
],
"ɦ": [
71
],
"ɧ": [
72
],
"ɨ": [
73
],
"ɪ": [
74
],
"ɫ": [
75
],
"ɬ": [
76
],
"ɭ": [
77
],
"ɮ": [
78
],
"ɯ": [
79
],
"ɰ": [
80
],
"ɱ": [
81
],
"ɲ": [
82
],
"ɳ": [
83
],
"ɴ": [
84
],
"ɵ": [
85
],
"ɶ": [
86
],
"ɸ": [
87
],
"ɹ": [
88
],
"ɺ": [
89
],
"ɻ": [
90
],
"ɽ": [
91
],
"ɾ": [
92
],
"ʀ": [
93
],
"ʁ": [
94
],
"ʂ": [
95
],
"ʃ": [
96
],
"ʄ": [
97
],
"ʈ": [
98
],
"ʉ": [
99
],
"ʊ": [
100
],
"ʋ": [
101
],
"ʌ": [
102
],
"ʍ": [
103
],
"ʎ": [
104
],
"ʏ": [
105
],
"ʐ": [
106
],
"ʑ": [
107
],
"ʒ": [
108
],
"ʔ": [
109
],
"ʕ": [
110
],
"ʘ": [
111
],
"ʙ": [
112
],
"ʛ": [
113
],
"ʜ": [
114
],
"ʝ": [
115
],
"ʟ": [
116
],
"ʡ": [
117
],
"ʢ": [
118
],
"ʲ": [
119
],
"ˈ": [
120
],
"ˌ": [
121
],
"ː": [
122
],
"ˑ": [
123
],
"˞": [
124
],
"β": [
125
],
"θ": [
126
],
"χ": [
127
],
"ᵻ": [
128
],
"ⱱ": [
129
],
"0": [
130
],
"1": [
131
],
"2": [
132
],
"3": [
133
],
"4": [
134
],
"5": [
135
],
"6": [
136
],
"7": [
137
],
"8": [
138
],
"9": [
139
],
"̧": [
140
],
"̃": [
141
],
"̪": [
142
],
"̯": [
143
],
"̩": [
144
],
"ʰ": [
145
],
"ˤ": [
146
],
"ε": [
147
],
"↓": [
148
],
"#": [
149
],
"\"": [
150
]
},
"num_symbols": 256,
"num_speakers": 1,
"speaker_id_map": {},
"piper_version": "1.0.0",
"language": {
"code": "en_US",
"family": "en",
"region": "US",
"name_native": "English",
"name_english": "English",
"country_english": "United States"
},
"dataset": "joe"
}