Initial commit

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
forgejoadmin 2026-08-18 03:16:42 -04:00
commit 71db2d1ab9
126 changed files with 3198 additions and 0 deletions

View file

@ -0,0 +1,87 @@
# Image recognition de-risking — report
Goal: figure out how the app recognizes a Pokemon when the phone is
pointed at a figure/toy, before committing to an architecture.
## Test set
5 classes (Bulbasaur, Squirtle, Pikachu, Eevee, Charizard), grown over the
course of the spike to 12 query images spanning very different visual
styles on purpose: stock photos, Pokemon GO AR screenshots, plush toys
(including two real photos of figures the user owns), a TCG card, a tiny
battle sprite, and anime screencaps. Two reference sets were tried:
scraped stock photos, and Bulbapedia's official artwork.
## Approaches tried, in order
### 1. CLIP (ViT-B-32, OpenAI weights) + nearest-neighbor
Embed a reference image per species, embed the query, cosine-similarity
match. Initial run used the wrong open_clip model variant (`ViT-B-32`
instead of `ViT-B-32-quickgelu`), which silently degrades OpenAI-weight
embeddings — worth remembering if this comes up again.
- Photo references: **4/12** correct top-1.
- Bulbapedia art references: **8/12** correct top-1.
### 2. DINOv2 (facebook/dinov2-base) + nearest-neighbor
Self-supervised, built for instance/visual similarity rather than
text-image alignment — expected to beat CLIP at this specific task.
- Photo references: **6/12**.
- Bulbapedia art references: **8/12**.
**Finding across both:** official art references consistently beat random
stock-photo references, regardless of model — the illustration-vs-photo
domain gap we worried about mattered less than material/form-factor
mismatch (plush vs. rigid figure vs. flat art). `charizard_plush` failed
in literally every embedding config tried (0/4) — plush toys are the
genuinely hard case for this whole approach, not photos-vs-domain style.
### 3. Direct vision-LLM recognition (the "cheat")
Skip reference images and embeddings entirely — ask a vision-language
model "what Pokemon is this" and let its pretrained world knowledge do
the work.
- Claude (me, just looking at the images): **12/12**.
- Self-hosted Qwen2.5-VL-3B-Instruct, run locally on the RTX 5070 Ti:
**10/12** cold, no fine-tuning, no reference images at all. The 2
misses were the two genuinely hardest images in the set (a tiny
213x240 keychain thumbnail → correctly returned "unknown" rather than
a wrong guess; and a plush the user themselves said "looks like shit,
not even sure that's Charizard").
## Decision
Went with **self-hosted VLM recognition** (Qwen2.5-VL-3B-Instruct) over
the embedding/nearest-neighbor approach. Reasons:
- Meaningfully higher accuracy (10/12 vs. best embedding score of 8/12).
- No reference-image sourcing/maintenance needed for 1000+ species —
eliminates the "content volume" risk from the original risk assessment
entirely.
- Degrades safely: genuinely ambiguous images tend to get "unknown"
rather than a confident wrong answer.
Trade-off accepted: requires a GPU server reachable over the network at
recognition time (already an accepted dependency — the original plan
always involved uploading the photo to a home server).
## What shipped from this
`server/` — FastAPI wrapper around the same Qwen2.5-VL-3B pipeline,
running on this Windows machine (chosen over buying a GPU for Unraid).
`POST /identify` takes a photo, returns `{recognized, species,
raw_response}`. Verified working end-to-end over real HTTP.
## Open questions / not yet tested
- Accuracy at real scale (1000+ candidate species) is untested — only
ever tried 5 classes. Confusion likely increases with more classes.
- Never tested against the user's own figures except for two Charizard
photos (both plush) — the real target (rigid painted figures) hasn't
been tried.
- Larger models (Qwen2.5-VL-7B+) not tried — likely closes some of the
remaining gap to the 12/12 upper bound, at the cost of latency/VRAM.
- No latency/throughput measurement done — only correctness.

View file

@ -0,0 +1,83 @@
"""
De-risking spike: can a generic vision embedding model (CLIP) tell Pokemon
figures/toys apart via nearest-neighbor lookup, with no training?
Pipeline: embed every image under images/reference/ (one file per Pokemon,
filename = class name) to build a small reference index, then embed every
image under images/query/ and report the nearest reference neighbor +
cosine similarity. A "pass" is query/<name>.jpg matching reference/<name>.jpg
as the top-1 hit.
Usage: .venv\\Scripts\\python.exe match_spike.py
"""
import sys
from pathlib import Path
import open_clip
import torch
from PIL import Image
HERE = Path(__file__).parent
REF_DIR = HERE / "images" / (sys.argv[1] if len(sys.argv) > 1 else "reference")
QUERY_DIR = HERE / "images" / "query"
MODEL_NAME = "ViT-B-32-quickgelu"
PRETRAINED = "openai"
def load_model():
model, _, preprocess = open_clip.create_model_and_transforms(
MODEL_NAME, pretrained=PRETRAINED
)
model.eval()
return model, preprocess
def embed_image(model, preprocess, path: Path) -> torch.Tensor:
image = preprocess(Image.open(path).convert("RGB")).unsqueeze(0)
with torch.no_grad():
features = model.encode_image(image)
features = features / features.norm(dim=-1, keepdim=True)
return features.squeeze(0)
def main():
print(f"Loading {MODEL_NAME} ({PRETRAINED})...")
model, preprocess = load_model()
ref_paths = (
sorted(REF_DIR.glob("*.jpg"))
+ sorted(REF_DIR.glob("*.webp"))
+ sorted(REF_DIR.glob("*.png"))
)
query_paths = (
sorted(QUERY_DIR.glob("*.jpg"))
+ sorted(QUERY_DIR.glob("*.webp"))
+ sorted(QUERY_DIR.glob("*.png"))
)
print(f"Embedding {len(ref_paths)} reference images...")
ref_names = [p.stem for p in ref_paths]
ref_embeds = torch.stack([embed_image(model, preprocess, p) for p in ref_paths])
print(f"Embedding {len(query_paths)} query images...\n")
correct = 0
for qpath in query_paths:
qembed = embed_image(model, preprocess, qpath)
sims = ref_embeds @ qembed # cosine similarity, both sides unit-norm
ranked = sorted(zip(ref_names, sims.tolist()), key=lambda x: -x[1])
top_name, top_sim = ranked[0]
expected = qpath.stem.split("_")[0]
is_correct = top_name == expected
correct += is_correct
marker = "OK " if is_correct else "MISS"
print(f"[{marker}] query={qpath.stem:<18} -> best={top_name:<10} sim={top_sim:.4f}")
runner_up = ", ".join(f"{n}={s:.3f}" for n, s in ranked[1:4])
print(f" runner-up: {runner_up}")
print(f"\n{correct}/{len(query_paths)} correct top-1 matches")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,82 @@
"""
Same spike as match_spike.py, but using DINOv2 instead of CLIP.
DINOv2 is trained with a self-supervised image-only objective specifically
aimed at instance/fine-grained visual similarity, which is a better fit for
"is this query photo the same object as this reference photo" than CLIP
(CLIP is trained for text-image alignment and tends to cluster images by
generic scene/style rather than object identity).
Usage: .venv\\Scripts\\python.exe match_spike_dinov2.py
"""
import sys
from pathlib import Path
import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModel
HERE = Path(__file__).parent
REF_DIR = HERE / "images" / (sys.argv[1] if len(sys.argv) > 1 else "reference")
QUERY_DIR = HERE / "images" / "query"
MODEL_NAME = "facebook/dinov2-base"
def load_model():
processor = AutoImageProcessor.from_pretrained(MODEL_NAME)
model = AutoModel.from_pretrained(MODEL_NAME)
model.eval()
return model, processor
def embed_image(model, processor, path: Path) -> torch.Tensor:
image = Image.open(path).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
features = outputs.last_hidden_state[:, 0, :] # CLS token
features = features / features.norm(dim=-1, keepdim=True)
return features.squeeze(0)
def main():
print(f"Loading {MODEL_NAME}...")
model, processor = load_model()
ref_paths = (
sorted(REF_DIR.glob("*.jpg"))
+ sorted(REF_DIR.glob("*.webp"))
+ sorted(REF_DIR.glob("*.png"))
)
query_paths = (
sorted(QUERY_DIR.glob("*.jpg"))
+ sorted(QUERY_DIR.glob("*.webp"))
+ sorted(QUERY_DIR.glob("*.png"))
)
print(f"Embedding {len(ref_paths)} reference images...")
ref_names = [p.stem for p in ref_paths]
ref_embeds = torch.stack([embed_image(model, processor, p) for p in ref_paths])
print(f"Embedding {len(query_paths)} query images...\n")
correct = 0
for qpath in query_paths:
qembed = embed_image(model, processor, qpath)
sims = ref_embeds @ qembed
ranked = sorted(zip(ref_names, sims.tolist()), key=lambda x: -x[1])
top_name, top_sim = ranked[0]
expected = qpath.stem.split("_")[0]
is_correct = top_name == expected
correct += is_correct
marker = "OK " if is_correct else "MISS"
print(f"[{marker}] query={qpath.stem:<18} -> best={top_name:<10} sim={top_sim:.4f}")
runner_up = ", ".join(f"{n}={s:.3f}" for n, s in ranked[1:4])
print(f" runner-up: {runner_up}")
print(f"\n{correct}/{len(query_paths)} correct top-1 matches")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,87 @@
"""
De-risking spike, take 3: instead of embedding + nearest-neighbor lookup,
just ask a self-hosted vision-language model directly what Pokemon is in
the photo. No reference images, no vector index -- the model's own
pretrained world knowledge does the recognition.
Usage: .venv_vlm\\Scripts\\python.exe match_spike_vlm.py
"""
from pathlib import Path
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
HERE = Path(__file__).parent
QUERY_DIR = HERE / "images" / "query"
MODEL_NAME = "Qwen/Qwen2.5-VL-3B-Instruct"
PROMPT = (
"You are the recognition system inside a Pokedex app. Identify the "
"Pokemon species shown in this image, even if it's a toy, plush, "
"trading card, sprite, fan art, or in-game screenshot of it. "
"Reply with ONLY the species name, nothing else. If no Pokemon is "
"clearly depicted, reply with exactly: unknown"
)
def load_model():
processor = AutoProcessor.from_pretrained(MODEL_NAME)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
MODEL_NAME, torch_dtype=torch.bfloat16, device_map="cuda"
)
model.eval()
return model, processor
def identify(model, processor, path: Path) -> str:
image = Image.open(path).convert("RGB")
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": PROMPT},
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(text=[text], images=[image], return_tensors="pt").to("cuda")
with torch.no_grad():
generated = model.generate(**inputs, max_new_tokens=16)
trimmed = generated[:, inputs["input_ids"].shape[1] :]
output = processor.batch_decode(
trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=True
)[0]
return output.strip()
def main():
print(f"Loading {MODEL_NAME} on GPU...")
model, processor = load_model()
query_paths = (
sorted(QUERY_DIR.glob("*.jpg"))
+ sorted(QUERY_DIR.glob("*.webp"))
+ sorted(QUERY_DIR.glob("*.png"))
)
print(f"Identifying {len(query_paths)} query images...\n")
correct = 0
for qpath in query_paths:
expected = qpath.stem.split("_")[0]
answer = identify(model, processor, qpath)
is_correct = answer.strip().lower() == expected.lower()
correct += is_correct
marker = "OK " if is_correct else "MISS"
print(f"[{marker}] query={qpath.stem:<18} expected={expected:<10} model_said={answer!r}")
print(f"\n{correct}/{len(query_paths)} correct")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,4 @@
torch --index-url https://download.pytorch.org/whl/cpu
open_clip_torch
pillow
numpy

View file

@ -0,0 +1,57 @@
# Pokedex voice spike — scripts
De-risking spike for the Pokédex app's robotic voice output. Full pipeline:
text -> TTS -> optional DSP "robot" filter -> mp3.
## Files
- `generate_samples.sh` — end-to-end driver: installs are documented at the
top, then runs espeak-ng and Piper TTS, then the DSP filters, then
converts everything to mp3. Run it from this directory (`./generate_samples.sh`).
- `robot_filter.py` — heavy/original robot filter: square-wave ring
modulation, bitcrush (sample-and-hold + bit-depth reduction), narrow
band-pass, hard drive. This is what "too robotic" was generated with.
- `robot_filter_light.py` — barely-there version: sine-wave ring mod at
low mix, wide band-pass, mild soft-clip. Meant to add a hint of
synthesized edge to an already-natural neural voice without disturbing
cadence.
- `robot_filter_v2.py` — the current one: parameterized, takes an
`intensity` argument from 0.0 (untouched) to 1.0 (full heavy filter) and
interpolates ring-mod mix/carrier, bandpass width, bitcrush, and drive
along that scale. Usage: `python3 robot_filter_v2.py in.wav out.wav 0.45`.
This is the one to keep tuning going forward — texted intensity numbers
("try 0.6") map directly to its third argument.
## Voice model
TTS engine is [Piper](https://github.com/OHF-Voice/piper1-gpl) — a small,
fully offline neural TTS with real Android ports available, which matters
for the app's "works with no signal" requirement.
The voice used is `en_US-joe-medium.onnx`, fine-tuned from Piper's stock
"lessac" voice on a CC0-licensed dataset (see MODEL_CARD if you unpack the
wheel) — no attribution/licensing issue to ship it.
**How I got the model file in this sandbox:** the sandbox's network
allowlist didn't reach huggingface.co, where Piper's official voices are
hosted, only package registries. I found a community-published PyPI wheel
(`joe-us-piper-voice`) that bundles the .onnx file directly, so a plain
`pip install` pulled it down. That's a workaround specific to *this*
sandbox — for the real project, get voices the normal way from Piper's
official releases/voice list, which gives you far more voice choices
(different speakers, accents, quality tiers) than this one bundled option.
## What's still open
- Pronunciation of individual Pokémon species names hasn't been
systematically checked — only "Pokémon," "Bulbasaur," "Charizard,"
"Pikachu" were tested. Expect to need a pass listening to all ~1000+
species names and hand-fixing the ones the phonemizer (espeak-ng, under
Piper's hood) gets wrong, either by respelling in the source text or via
espeak's phoneme-override escape syntax.
- Robot filter intensity (`robot_filter_v2.py`'s `intensity` arg) needs to
land wherever the actual desired "amount of robotic" ends up — currently
parked at 0.45 pending feedback.
- This whole pipeline is meant to run once, offline, over every dex entry
ahead of time (batch pre-generation), not live on-device — see the
earlier discussion in this conversation for why.

View file

@ -0,0 +1,77 @@
"""
Windows-friendly regeneration of the voice spike, extended to all 5
Pokemon the app currently knows about. Ports generate_samples.sh's
"fixed" (cadence-corrected) + robot_filter_v2.py (intensity 0.45) stages
to Python, using piper's Python API directly instead of the `piper` CLI
and skipping the ffmpeg mp3 step (WAV plays fine on Android, and ffmpeg
isn't installed on this machine).
Usage: .venv\\Scripts\\python.exe generate_all.py
"""
import wave
from pathlib import Path
from piper import PiperVoice
from piper.config import SynthesisConfig
import robot_filter_v2
HERE = Path(__file__).parent
VOICE_MODEL = HERE / "voices" / "en_US-joe-medium.onnx"
OUT_DIR = HERE / "output"
ENTRIES = {
"bulbasaur": (
"Bulbasaur, the seed Pokémon. It can be seen napping in bright "
"sunlight. There is a seed on its back. By soaking up the sun's "
"rays, the seed grows progressively larger."
),
"charizard": (
"Charizard, the flame Pokémon. Charizard flies around the sky in "
"search of powerful opponents. It breathes fire of such great "
"heat that it melts anything."
),
"pikachu": (
"Pikachu, the mouse Pokémon. When several of these Pokémon "
"gather, their electricity could build and cause lightning "
"storms."
),
"eevee": (
"Eevee, the evolution Pokémon. Its genetic code is irregular. It "
"may mutate if it is exposed to radiation from element stones."
),
"squirtle": (
"Squirtle, the tiny turtle Pokémon. After birth, its back swells "
"and hardens into a shell. It powerfully sprays foam from its "
"mouth."
),
}
ROBOT_INTENSITY = 0.45
SYN_CONFIG = SynthesisConfig(noise_scale=0.5, noise_w_scale=0.3, length_scale=0.98)
def main():
OUT_DIR.mkdir(exist_ok=True)
print(f"Loading voice model {VOICE_MODEL.name}...")
voice = PiperVoice.load(str(VOICE_MODEL))
for name, text in ENTRIES.items():
fixed_path = OUT_DIR / f"{name}_fixed.wav"
robot_path = OUT_DIR / f"{name}_fixed_robot.wav"
print(f"Synthesizing {name}...")
with wave.open(str(fixed_path), "wb") as wav_file:
voice.synthesize_wav(text, wav_file, syn_config=SYN_CONFIG)
sr, x = robot_filter_v2.load(str(fixed_path))
y = robot_filter_v2.robotize(x, sr, intensity=ROBOT_INTENSITY)
robot_filter_v2.save(str(robot_path), sr, y)
print(f" -> {robot_path.name}")
print("\nDone. See output/*_fixed_robot.wav")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,74 @@
#!/usr/bin/env bash
# Reproduces the Pokedex voice spike end to end:
# TTS (espeak-ng and Piper) -> robot-voice DSP filter -> mp3
#
# Setup (Debian/Ubuntu):
# sudo apt-get install -y espeak-ng ffmpeg
# pip install piper-tts --break-system-packages
# pip install numpy scipy --break-system-packages
#
# Voice model:
# Official route: download a Piper voice (.onnx + .onnx.json) from the
# Piper voices repo and point --model at it, e.g. en_US-lessac-medium.
# In this sandbox, huggingface.co wasn't reachable, so I instead pulled a
# community PyPI wheel that bundles the model file directly:
# pip download --no-deps joe-us-piper-voice
# (unzip the wheel; the .onnx/.onnx.json live under
# joe_us_piper_voice/data/). Not an official source -- for production,
# get voices from Piper's own releases instead.
set -euo pipefail
cd "$(dirname "$0")"
VOICE_MODEL="voices/en_US-joe-medium.onnx"
export ALSA_CONFIG_PATH=/dev/null # silence ALSA warnings in headless envs
# ---- Stage 1: raw espeak-ng (formant synth, offline, inherently "robotic") ----
espeak-ng -v en-us -s 150 -w bulbasaur_raw.wav \
"Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
espeak-ng -v en-us -s 150 -w charizard_raw.wav \
"Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
espeak-ng -v en-us -s 150 -w pikachu_raw.wav \
"Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
# ---- Stage 2: Piper neural TTS, first pass (natural but bad cadence on rare words) ----
gen_neural() {
local name="$1" text="$2"
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_neural.wav"
}
gen_neural bulbasaur "Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
gen_neural charizard "Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
gen_neural pikachu "Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
# ---- Stage 3: Piper neural TTS, cadence-fixed pass ----
# Fix = (a) standard "X, the Y Pokemon." phrasing instead of two short
# sentences, which stopped "Pokemon" from landing phrase-final where
# duration models over-lengthen it, and (b) reduced noise-w-scale (duration
# randomness) and noise-scale (audio variance) so rare proper nouns don't
# get a random stretched/warped rendering.
gen_fixed() {
local name="$1" text="$2"
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_fixed.wav" \
--noise-scale 0.5 --noise-w-scale 0.3 --length-scale 0.98
}
gen_fixed bulbasaur "Bulbasaur, the seed Pokémon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
gen_fixed charizard "Charizard, the flame Pokémon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
gen_fixed pikachu "Pikachu, the mouse Pokémon. When several of these Pokémon gather, their electricity could build and cause lightning storms."
# ---- Stage 4: robot-voice DSP filter passes ----
# robot_filter.py -> heavy/original (square-wave ring mod + bitcrush + narrow bandpass)
# robot_filter_light.py -> barely-there (sine ring mod, wide bandpass, mild drive)
# robot_filter_v2.py -> parameterized 0..1 intensity knob (used at 0.45 = "medium")
for name in bulbasaur charizard pikachu; do
python3 robot_filter.py "${name}_raw.wav" "${name}_robot.wav"
python3 robot_filter_light.py "${name}_neural.wav" "${name}_neural_light.wav"
python3 robot_filter_v2.py "${name}_fixed.wav" "${name}_fixed_robot.wav" 0.45
done
# ---- Stage 5: mp3 for easy playback/delivery ----
for f in *_raw.wav *_robot.wav *_neural.wav *_neural_light.wav *_fixed.wav *_fixed_robot.wav; do
[ -f "$f" ] && ffmpeg -y -loglevel error -i "$f" -codec:a libmp3lame -qscale:a 4 "${f%.wav}.mp3"
done
echo "Done. See *.mp3 for output."

View file

@ -0,0 +1,79 @@
"""
Quick-and-dirty 'robot voice' post-processor for the Pokedex voice spike.
Pipeline: TTS wav in -> ring modulation + bitcrush + band-pass (speaker-grille
coloring) + light hard-clip drive -> wav out.
Usage: python3 robot_filter.py in.wav out.wav
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
elif data.dtype == np.uint8:
x = (data.astype(np.float64) - 128) / 128.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz=45.0, mix=0.55):
t = np.arange(len(x)) / sr
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) # square carrier = harsher/metallic
modulated = x * carrier
return (1 - mix) * x + mix * modulated
def bitcrush(x, bit_depth=6, sr=None, downsample_factor=3):
# sample-and-hold downsample (classic lo-fi/robotic stepping)
if downsample_factor > 1:
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
if len(held) < len(x):
held = np.pad(held, (0, len(x) - len(held)))
x = held
levels = 2 ** bit_depth
x = np.round(x * levels) / levels
return x
def bandpass(x, sr, low=350.0, high=3400.0, order=4):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount=1.8):
return np.tanh(x * amount) / np.tanh(amount)
def robotize(x, sr):
y = ring_modulate(x, sr, carrier_hz=45.0, mix=0.5)
y = bitcrush(y, bit_depth=7, downsample_factor=2)
y = bandpass(y, sr, low=300, high=3800)
y = drive(y, amount=1.6)
# normalize
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
sr, x = load(src)
y = robotize(x, sr)
save(dst, sr, y)
print(f"wrote {dst}")

View file

@ -0,0 +1,58 @@
"""
Light-touch robot coloring for a neural (already-natural-sounding) TTS voice.
Goal: keep natural cadence/prosody intact, just add a synthesized/device edge.
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz=90.0, mix=0.12):
t = np.arange(len(x)) / sr
carrier = np.sin(2 * np.pi * carrier_hz * t) # sine carrier = subtler than square
return (1 - mix) * x + mix * (x * carrier)
def bandpass(x, sr, low=200.0, high=5500.0, order=2):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount=1.15):
return np.tanh(x * amount) / np.tanh(amount)
def robotize_light(x, sr):
y = ring_modulate(x, sr, carrier_hz=90.0, mix=0.12)
y = bandpass(y, sr, low=200, high=5500)
y = drive(y, amount=1.15)
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
sr, x = load(src)
y = robotize_light(x, sr)
save(dst, sr, y)
print(f"wrote {dst}")

View file

@ -0,0 +1,72 @@
"""
Parameterized robot-voice filter, v2.
intensity in [0, 1]: 0 = untouched, 1 = full "heavy" robot from round 1.
"""
import sys
import numpy as np
from scipy.io import wavfile
from scipy.signal import butter, sosfilt
def load(path):
sr, data = wavfile.read(path)
if data.dtype == np.int16:
x = data.astype(np.float64) / 32768.0
elif data.dtype == np.int32:
x = data.astype(np.float64) / 2147483648.0
else:
x = data.astype(np.float64)
if x.ndim > 1:
x = x.mean(axis=1)
return sr, x
def save(path, sr, x):
x = np.clip(x, -1.0, 1.0)
wavfile.write(path, sr, (x * 32767).astype(np.int16))
def ring_modulate(x, sr, carrier_hz, mix, square=False):
t = np.arange(len(x)) / sr
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) if square else np.sin(2 * np.pi * carrier_hz * t)
return (1 - mix) * x + mix * (x * carrier)
def bitcrush(x, bit_depth, downsample_factor):
if downsample_factor > 1:
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
if len(held) < len(x):
held = np.pad(held, (0, len(x) - len(held)))
x = held
levels = 2 ** bit_depth
return np.round(x * levels) / levels
def bandpass(x, sr, low, high, order=3):
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
return sosfilt(sos, x)
def drive(x, amount):
return np.tanh(x * amount) / np.tanh(amount)
def robotize(x, sr, intensity=0.45):
i = intensity
y = ring_modulate(x, sr, carrier_hz=60 - 15 * i, mix=0.1 + 0.45 * i, square=(i > 0.6))
if i > 0.35:
y = bitcrush(y, bit_depth=int(12 - 6 * i), downsample_factor=1 if i < 0.7 else 2)
y = bandpass(y, sr, low=250 - 100 * i, high=6000 - 2200 * i)
y = drive(y, amount=1.0 + 0.9 * i)
peak = np.max(np.abs(y)) or 1.0
y = y / peak * 0.9
return y
if __name__ == "__main__":
src, dst = sys.argv[1], sys.argv[2]
intensity = float(sys.argv[3]) if len(sys.argv) > 3 else 0.45
sr, x = load(src)
y = robotize(x, sr, intensity=intensity)
save(dst, sr, y)
print(f"wrote {dst} (intensity={intensity})")

View file

@ -0,0 +1,484 @@
{
"audio": {
"sample_rate": 22050,
"quality": "medium"
},
"espeak": {
"voice": "en-us"
},
"inference": {
"noise_scale": 0.667,
"length_scale": 1,
"noise_w": 0.8
},
"phoneme_type": "espeak",
"phoneme_map": {},
"phoneme_id_map": {
"_": [
0
],
"^": [
1
],
"$": [
2
],
" ": [
3
],
"!": [
4
],
"'": [
5
],
"(": [
6
],
")": [
7
],
",": [
8
],
"-": [
9
],
".": [
10
],
":": [
11
],
";": [
12
],
"?": [
13
],
"a": [
14
],
"b": [
15
],
"c": [
16
],
"d": [
17
],
"e": [
18
],
"f": [
19
],
"h": [
20
],
"i": [
21
],
"j": [
22
],
"k": [
23
],
"l": [
24
],
"m": [
25
],
"n": [
26
],
"o": [
27
],
"p": [
28
],
"q": [
29
],
"r": [
30
],
"s": [
31
],
"t": [
32
],
"u": [
33
],
"v": [
34
],
"w": [
35
],
"x": [
36
],
"y": [
37
],
"z": [
38
],
"æ": [
39
],
"ç": [
40
],
"ð": [
41
],
"ø": [
42
],
"ħ": [
43
],
"ŋ": [
44
],
"œ": [
45
],
"ǀ": [
46
],
"ǁ": [
47
],
"ǂ": [
48
],
"ǃ": [
49
],
"ɐ": [
50
],
"ɑ": [
51
],
"ɒ": [
52
],
"ɓ": [
53
],
"ɔ": [
54
],
"ɕ": [
55
],
"ɖ": [
56
],
"ɗ": [
57
],
"ɘ": [
58
],
"ə": [
59
],
"ɚ": [
60
],
"ɛ": [
61
],
"ɜ": [
62
],
"ɞ": [
63
],
"ɟ": [
64
],
"ɠ": [
65
],
"ɡ": [
66
],
"ɢ": [
67
],
"ɣ": [
68
],
"ɤ": [
69
],
"ɥ": [
70
],
"ɦ": [
71
],
"ɧ": [
72
],
"ɨ": [
73
],
"ɪ": [
74
],
"ɫ": [
75
],
"ɬ": [
76
],
"ɭ": [
77
],
"ɮ": [
78
],
"ɯ": [
79
],
"ɰ": [
80
],
"ɱ": [
81
],
"ɲ": [
82
],
"ɳ": [
83
],
"ɴ": [
84
],
"ɵ": [
85
],
"ɶ": [
86
],
"ɸ": [
87
],
"ɹ": [
88
],
"ɺ": [
89
],
"ɻ": [
90
],
"ɽ": [
91
],
"ɾ": [
92
],
"ʀ": [
93
],
"ʁ": [
94
],
"ʂ": [
95
],
"ʃ": [
96
],
"ʄ": [
97
],
"ʈ": [
98
],
"ʉ": [
99
],
"ʊ": [
100
],
"ʋ": [
101
],
"ʌ": [
102
],
"ʍ": [
103
],
"ʎ": [
104
],
"ʏ": [
105
],
"ʐ": [
106
],
"ʑ": [
107
],
"ʒ": [
108
],
"ʔ": [
109
],
"ʕ": [
110
],
"ʘ": [
111
],
"ʙ": [
112
],
"ʛ": [
113
],
"ʜ": [
114
],
"ʝ": [
115
],
"ʟ": [
116
],
"ʡ": [
117
],
"ʢ": [
118
],
"ʲ": [
119
],
"ˈ": [
120
],
"ˌ": [
121
],
"ː": [
122
],
"ˑ": [
123
],
"˞": [
124
],
"β": [
125
],
"θ": [
126
],
"χ": [
127
],
"ᵻ": [
128
],
"ⱱ": [
129
],
"0": [
130
],
"1": [
131
],
"2": [
132
],
"3": [
133
],
"4": [
134
],
"5": [
135
],
"6": [
136
],
"7": [
137
],
"8": [
138
],
"9": [
139
],
"̧": [
140
],
"̃": [
141
],
"̪": [
142
],
"̯": [
143
],
"̩": [
144
],
"ʰ": [
145
],
"ˤ": [
146
],
"ε": [
147
],
"↓": [
148
],
"#": [
149
],
"\"": [
150
]
},
"num_symbols": 256,
"num_speakers": 1,
"speaker_id_map": {},
"piper_version": "1.0.0",
"language": {
"code": "en_US",
"family": "en",
"region": "US",
"name_native": "English",
"name_english": "English",
"country_english": "United States"
},
"dataset": "joe"
}

Binary file not shown.

After

Width:  |  Height:  |  Size: 104 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 135 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 42 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 49 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 15 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 45 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 96 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 114 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 161 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 145 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 57 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 318 KiB