Initial commit
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
87
saved/image_recognition_spike/REPORT.md
Normal file
|
|
@ -0,0 +1,87 @@
|
|||
# Image recognition de-risking — report
|
||||
|
||||
Goal: figure out how the app recognizes a Pokemon when the phone is
|
||||
pointed at a figure/toy, before committing to an architecture.
|
||||
|
||||
## Test set
|
||||
|
||||
5 classes (Bulbasaur, Squirtle, Pikachu, Eevee, Charizard), grown over the
|
||||
course of the spike to 12 query images spanning very different visual
|
||||
styles on purpose: stock photos, Pokemon GO AR screenshots, plush toys
|
||||
(including two real photos of figures the user owns), a TCG card, a tiny
|
||||
battle sprite, and anime screencaps. Two reference sets were tried:
|
||||
scraped stock photos, and Bulbapedia's official artwork.
|
||||
|
||||
## Approaches tried, in order
|
||||
|
||||
### 1. CLIP (ViT-B-32, OpenAI weights) + nearest-neighbor
|
||||
|
||||
Embed a reference image per species, embed the query, cosine-similarity
|
||||
match. Initial run used the wrong open_clip model variant (`ViT-B-32`
|
||||
instead of `ViT-B-32-quickgelu`), which silently degrades OpenAI-weight
|
||||
embeddings — worth remembering if this comes up again.
|
||||
|
||||
- Photo references: **4/12** correct top-1.
|
||||
- Bulbapedia art references: **8/12** correct top-1.
|
||||
|
||||
### 2. DINOv2 (facebook/dinov2-base) + nearest-neighbor
|
||||
|
||||
Self-supervised, built for instance/visual similarity rather than
|
||||
text-image alignment — expected to beat CLIP at this specific task.
|
||||
|
||||
- Photo references: **6/12**.
|
||||
- Bulbapedia art references: **8/12**.
|
||||
|
||||
**Finding across both:** official art references consistently beat random
|
||||
stock-photo references, regardless of model — the illustration-vs-photo
|
||||
domain gap we worried about mattered less than material/form-factor
|
||||
mismatch (plush vs. rigid figure vs. flat art). `charizard_plush` failed
|
||||
in literally every embedding config tried (0/4) — plush toys are the
|
||||
genuinely hard case for this whole approach, not photos-vs-domain style.
|
||||
|
||||
### 3. Direct vision-LLM recognition (the "cheat")
|
||||
|
||||
Skip reference images and embeddings entirely — ask a vision-language
|
||||
model "what Pokemon is this" and let its pretrained world knowledge do
|
||||
the work.
|
||||
|
||||
- Claude (me, just looking at the images): **12/12**.
|
||||
- Self-hosted Qwen2.5-VL-3B-Instruct, run locally on the RTX 5070 Ti:
|
||||
**10/12** cold, no fine-tuning, no reference images at all. The 2
|
||||
misses were the two genuinely hardest images in the set (a tiny
|
||||
213x240 keychain thumbnail → correctly returned "unknown" rather than
|
||||
a wrong guess; and a plush the user themselves said "looks like shit,
|
||||
not even sure that's Charizard").
|
||||
|
||||
## Decision
|
||||
|
||||
Went with **self-hosted VLM recognition** (Qwen2.5-VL-3B-Instruct) over
|
||||
the embedding/nearest-neighbor approach. Reasons:
|
||||
- Meaningfully higher accuracy (10/12 vs. best embedding score of 8/12).
|
||||
- No reference-image sourcing/maintenance needed for 1000+ species —
|
||||
eliminates the "content volume" risk from the original risk assessment
|
||||
entirely.
|
||||
- Degrades safely: genuinely ambiguous images tend to get "unknown"
|
||||
rather than a confident wrong answer.
|
||||
|
||||
Trade-off accepted: requires a GPU server reachable over the network at
|
||||
recognition time (already an accepted dependency — the original plan
|
||||
always involved uploading the photo to a home server).
|
||||
|
||||
## What shipped from this
|
||||
|
||||
`server/` — FastAPI wrapper around the same Qwen2.5-VL-3B pipeline,
|
||||
running on this Windows machine (chosen over buying a GPU for Unraid).
|
||||
`POST /identify` takes a photo, returns `{recognized, species,
|
||||
raw_response}`. Verified working end-to-end over real HTTP.
|
||||
|
||||
## Open questions / not yet tested
|
||||
|
||||
- Accuracy at real scale (1000+ candidate species) is untested — only
|
||||
ever tried 5 classes. Confusion likely increases with more classes.
|
||||
- Never tested against the user's own figures except for two Charizard
|
||||
photos (both plush) — the real target (rigid painted figures) hasn't
|
||||
been tried.
|
||||
- Larger models (Qwen2.5-VL-7B+) not tried — likely closes some of the
|
||||
remaining gap to the 12/12 upper bound, at the cost of latency/VRAM.
|
||||
- No latency/throughput measurement done — only correctness.
|
||||
83
saved/image_recognition_spike/match_spike.py
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
"""
|
||||
De-risking spike: can a generic vision embedding model (CLIP) tell Pokemon
|
||||
figures/toys apart via nearest-neighbor lookup, with no training?
|
||||
|
||||
Pipeline: embed every image under images/reference/ (one file per Pokemon,
|
||||
filename = class name) to build a small reference index, then embed every
|
||||
image under images/query/ and report the nearest reference neighbor +
|
||||
cosine similarity. A "pass" is query/<name>.jpg matching reference/<name>.jpg
|
||||
as the top-1 hit.
|
||||
|
||||
Usage: .venv\\Scripts\\python.exe match_spike.py
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import open_clip
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
REF_DIR = HERE / "images" / (sys.argv[1] if len(sys.argv) > 1 else "reference")
|
||||
QUERY_DIR = HERE / "images" / "query"
|
||||
|
||||
MODEL_NAME = "ViT-B-32-quickgelu"
|
||||
PRETRAINED = "openai"
|
||||
|
||||
|
||||
def load_model():
|
||||
model, _, preprocess = open_clip.create_model_and_transforms(
|
||||
MODEL_NAME, pretrained=PRETRAINED
|
||||
)
|
||||
model.eval()
|
||||
return model, preprocess
|
||||
|
||||
|
||||
def embed_image(model, preprocess, path: Path) -> torch.Tensor:
|
||||
image = preprocess(Image.open(path).convert("RGB")).unsqueeze(0)
|
||||
with torch.no_grad():
|
||||
features = model.encode_image(image)
|
||||
features = features / features.norm(dim=-1, keepdim=True)
|
||||
return features.squeeze(0)
|
||||
|
||||
|
||||
def main():
|
||||
print(f"Loading {MODEL_NAME} ({PRETRAINED})...")
|
||||
model, preprocess = load_model()
|
||||
|
||||
ref_paths = (
|
||||
sorted(REF_DIR.glob("*.jpg"))
|
||||
+ sorted(REF_DIR.glob("*.webp"))
|
||||
+ sorted(REF_DIR.glob("*.png"))
|
||||
)
|
||||
query_paths = (
|
||||
sorted(QUERY_DIR.glob("*.jpg"))
|
||||
+ sorted(QUERY_DIR.glob("*.webp"))
|
||||
+ sorted(QUERY_DIR.glob("*.png"))
|
||||
)
|
||||
|
||||
print(f"Embedding {len(ref_paths)} reference images...")
|
||||
ref_names = [p.stem for p in ref_paths]
|
||||
ref_embeds = torch.stack([embed_image(model, preprocess, p) for p in ref_paths])
|
||||
|
||||
print(f"Embedding {len(query_paths)} query images...\n")
|
||||
correct = 0
|
||||
for qpath in query_paths:
|
||||
qembed = embed_image(model, preprocess, qpath)
|
||||
sims = ref_embeds @ qembed # cosine similarity, both sides unit-norm
|
||||
ranked = sorted(zip(ref_names, sims.tolist()), key=lambda x: -x[1])
|
||||
top_name, top_sim = ranked[0]
|
||||
expected = qpath.stem.split("_")[0]
|
||||
is_correct = top_name == expected
|
||||
correct += is_correct
|
||||
marker = "OK " if is_correct else "MISS"
|
||||
print(f"[{marker}] query={qpath.stem:<18} -> best={top_name:<10} sim={top_sim:.4f}")
|
||||
runner_up = ", ".join(f"{n}={s:.3f}" for n, s in ranked[1:4])
|
||||
print(f" runner-up: {runner_up}")
|
||||
|
||||
print(f"\n{correct}/{len(query_paths)} correct top-1 matches")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
82
saved/image_recognition_spike/match_spike_dinov2.py
Normal file
|
|
@ -0,0 +1,82 @@
|
|||
"""
|
||||
Same spike as match_spike.py, but using DINOv2 instead of CLIP.
|
||||
|
||||
DINOv2 is trained with a self-supervised image-only objective specifically
|
||||
aimed at instance/fine-grained visual similarity, which is a better fit for
|
||||
"is this query photo the same object as this reference photo" than CLIP
|
||||
(CLIP is trained for text-image alignment and tends to cluster images by
|
||||
generic scene/style rather than object identity).
|
||||
|
||||
Usage: .venv\\Scripts\\python.exe match_spike_dinov2.py
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
from PIL import Image
|
||||
from transformers import AutoImageProcessor, AutoModel
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
REF_DIR = HERE / "images" / (sys.argv[1] if len(sys.argv) > 1 else "reference")
|
||||
QUERY_DIR = HERE / "images" / "query"
|
||||
|
||||
MODEL_NAME = "facebook/dinov2-base"
|
||||
|
||||
|
||||
def load_model():
|
||||
processor = AutoImageProcessor.from_pretrained(MODEL_NAME)
|
||||
model = AutoModel.from_pretrained(MODEL_NAME)
|
||||
model.eval()
|
||||
return model, processor
|
||||
|
||||
|
||||
def embed_image(model, processor, path: Path) -> torch.Tensor:
|
||||
image = Image.open(path).convert("RGB")
|
||||
inputs = processor(images=image, return_tensors="pt")
|
||||
with torch.no_grad():
|
||||
outputs = model(**inputs)
|
||||
features = outputs.last_hidden_state[:, 0, :] # CLS token
|
||||
features = features / features.norm(dim=-1, keepdim=True)
|
||||
return features.squeeze(0)
|
||||
|
||||
|
||||
def main():
|
||||
print(f"Loading {MODEL_NAME}...")
|
||||
model, processor = load_model()
|
||||
|
||||
ref_paths = (
|
||||
sorted(REF_DIR.glob("*.jpg"))
|
||||
+ sorted(REF_DIR.glob("*.webp"))
|
||||
+ sorted(REF_DIR.glob("*.png"))
|
||||
)
|
||||
query_paths = (
|
||||
sorted(QUERY_DIR.glob("*.jpg"))
|
||||
+ sorted(QUERY_DIR.glob("*.webp"))
|
||||
+ sorted(QUERY_DIR.glob("*.png"))
|
||||
)
|
||||
|
||||
print(f"Embedding {len(ref_paths)} reference images...")
|
||||
ref_names = [p.stem for p in ref_paths]
|
||||
ref_embeds = torch.stack([embed_image(model, processor, p) for p in ref_paths])
|
||||
|
||||
print(f"Embedding {len(query_paths)} query images...\n")
|
||||
correct = 0
|
||||
for qpath in query_paths:
|
||||
qembed = embed_image(model, processor, qpath)
|
||||
sims = ref_embeds @ qembed
|
||||
ranked = sorted(zip(ref_names, sims.tolist()), key=lambda x: -x[1])
|
||||
top_name, top_sim = ranked[0]
|
||||
expected = qpath.stem.split("_")[0]
|
||||
is_correct = top_name == expected
|
||||
correct += is_correct
|
||||
marker = "OK " if is_correct else "MISS"
|
||||
print(f"[{marker}] query={qpath.stem:<18} -> best={top_name:<10} sim={top_sim:.4f}")
|
||||
runner_up = ", ".join(f"{n}={s:.3f}" for n, s in ranked[1:4])
|
||||
print(f" runner-up: {runner_up}")
|
||||
|
||||
print(f"\n{correct}/{len(query_paths)} correct top-1 matches")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
87
saved/image_recognition_spike/match_spike_vlm.py
Normal file
|
|
@ -0,0 +1,87 @@
|
|||
"""
|
||||
De-risking spike, take 3: instead of embedding + nearest-neighbor lookup,
|
||||
just ask a self-hosted vision-language model directly what Pokemon is in
|
||||
the photo. No reference images, no vector index -- the model's own
|
||||
pretrained world knowledge does the recognition.
|
||||
|
||||
Usage: .venv_vlm\\Scripts\\python.exe match_spike_vlm.py
|
||||
"""
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
from PIL import Image
|
||||
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
QUERY_DIR = HERE / "images" / "query"
|
||||
|
||||
MODEL_NAME = "Qwen/Qwen2.5-VL-3B-Instruct"
|
||||
|
||||
PROMPT = (
|
||||
"You are the recognition system inside a Pokedex app. Identify the "
|
||||
"Pokemon species shown in this image, even if it's a toy, plush, "
|
||||
"trading card, sprite, fan art, or in-game screenshot of it. "
|
||||
"Reply with ONLY the species name, nothing else. If no Pokemon is "
|
||||
"clearly depicted, reply with exactly: unknown"
|
||||
)
|
||||
|
||||
|
||||
def load_model():
|
||||
processor = AutoProcessor.from_pretrained(MODEL_NAME)
|
||||
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
|
||||
MODEL_NAME, torch_dtype=torch.bfloat16, device_map="cuda"
|
||||
)
|
||||
model.eval()
|
||||
return model, processor
|
||||
|
||||
|
||||
def identify(model, processor, path: Path) -> str:
|
||||
image = Image.open(path).convert("RGB")
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": image},
|
||||
{"type": "text", "text": PROMPT},
|
||||
],
|
||||
}
|
||||
]
|
||||
text = processor.apply_chat_template(
|
||||
messages, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
inputs = processor(text=[text], images=[image], return_tensors="pt").to("cuda")
|
||||
with torch.no_grad():
|
||||
generated = model.generate(**inputs, max_new_tokens=16)
|
||||
trimmed = generated[:, inputs["input_ids"].shape[1] :]
|
||||
output = processor.batch_decode(
|
||||
trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=True
|
||||
)[0]
|
||||
return output.strip()
|
||||
|
||||
|
||||
def main():
|
||||
print(f"Loading {MODEL_NAME} on GPU...")
|
||||
model, processor = load_model()
|
||||
|
||||
query_paths = (
|
||||
sorted(QUERY_DIR.glob("*.jpg"))
|
||||
+ sorted(QUERY_DIR.glob("*.webp"))
|
||||
+ sorted(QUERY_DIR.glob("*.png"))
|
||||
)
|
||||
|
||||
print(f"Identifying {len(query_paths)} query images...\n")
|
||||
correct = 0
|
||||
for qpath in query_paths:
|
||||
expected = qpath.stem.split("_")[0]
|
||||
answer = identify(model, processor, qpath)
|
||||
is_correct = answer.strip().lower() == expected.lower()
|
||||
correct += is_correct
|
||||
marker = "OK " if is_correct else "MISS"
|
||||
print(f"[{marker}] query={qpath.stem:<18} expected={expected:<10} model_said={answer!r}")
|
||||
|
||||
print(f"\n{correct}/{len(query_paths)} correct")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
4
saved/image_recognition_spike/requirements.txt
Normal file
|
|
@ -0,0 +1,4 @@
|
|||
torch --index-url https://download.pytorch.org/whl/cpu
|
||||
open_clip_torch
|
||||
pillow
|
||||
numpy
|
||||
57
saved/pokedex_voice_spike_scripts/README.md
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
# Pokedex voice spike — scripts
|
||||
|
||||
De-risking spike for the Pokédex app's robotic voice output. Full pipeline:
|
||||
text -> TTS -> optional DSP "robot" filter -> mp3.
|
||||
|
||||
## Files
|
||||
|
||||
- `generate_samples.sh` — end-to-end driver: installs are documented at the
|
||||
top, then runs espeak-ng and Piper TTS, then the DSP filters, then
|
||||
converts everything to mp3. Run it from this directory (`./generate_samples.sh`).
|
||||
- `robot_filter.py` — heavy/original robot filter: square-wave ring
|
||||
modulation, bitcrush (sample-and-hold + bit-depth reduction), narrow
|
||||
band-pass, hard drive. This is what "too robotic" was generated with.
|
||||
- `robot_filter_light.py` — barely-there version: sine-wave ring mod at
|
||||
low mix, wide band-pass, mild soft-clip. Meant to add a hint of
|
||||
synthesized edge to an already-natural neural voice without disturbing
|
||||
cadence.
|
||||
- `robot_filter_v2.py` — the current one: parameterized, takes an
|
||||
`intensity` argument from 0.0 (untouched) to 1.0 (full heavy filter) and
|
||||
interpolates ring-mod mix/carrier, bandpass width, bitcrush, and drive
|
||||
along that scale. Usage: `python3 robot_filter_v2.py in.wav out.wav 0.45`.
|
||||
This is the one to keep tuning going forward — texted intensity numbers
|
||||
("try 0.6") map directly to its third argument.
|
||||
|
||||
## Voice model
|
||||
|
||||
TTS engine is [Piper](https://github.com/OHF-Voice/piper1-gpl) — a small,
|
||||
fully offline neural TTS with real Android ports available, which matters
|
||||
for the app's "works with no signal" requirement.
|
||||
|
||||
The voice used is `en_US-joe-medium.onnx`, fine-tuned from Piper's stock
|
||||
"lessac" voice on a CC0-licensed dataset (see MODEL_CARD if you unpack the
|
||||
wheel) — no attribution/licensing issue to ship it.
|
||||
|
||||
**How I got the model file in this sandbox:** the sandbox's network
|
||||
allowlist didn't reach huggingface.co, where Piper's official voices are
|
||||
hosted, only package registries. I found a community-published PyPI wheel
|
||||
(`joe-us-piper-voice`) that bundles the .onnx file directly, so a plain
|
||||
`pip install` pulled it down. That's a workaround specific to *this*
|
||||
sandbox — for the real project, get voices the normal way from Piper's
|
||||
official releases/voice list, which gives you far more voice choices
|
||||
(different speakers, accents, quality tiers) than this one bundled option.
|
||||
|
||||
## What's still open
|
||||
|
||||
- Pronunciation of individual Pokémon species names hasn't been
|
||||
systematically checked — only "Pokémon," "Bulbasaur," "Charizard,"
|
||||
"Pikachu" were tested. Expect to need a pass listening to all ~1000+
|
||||
species names and hand-fixing the ones the phonemizer (espeak-ng, under
|
||||
Piper's hood) gets wrong, either by respelling in the source text or via
|
||||
espeak's phoneme-override escape syntax.
|
||||
- Robot filter intensity (`robot_filter_v2.py`'s `intensity` arg) needs to
|
||||
land wherever the actual desired "amount of robotic" ends up — currently
|
||||
parked at 0.45 pending feedback.
|
||||
- This whole pipeline is meant to run once, offline, over every dex entry
|
||||
ahead of time (batch pre-generation), not live on-device — see the
|
||||
earlier discussion in this conversation for why.
|
||||
77
saved/pokedex_voice_spike_scripts/generate_all.py
Normal file
|
|
@ -0,0 +1,77 @@
|
|||
"""
|
||||
Windows-friendly regeneration of the voice spike, extended to all 5
|
||||
Pokemon the app currently knows about. Ports generate_samples.sh's
|
||||
"fixed" (cadence-corrected) + robot_filter_v2.py (intensity 0.45) stages
|
||||
to Python, using piper's Python API directly instead of the `piper` CLI
|
||||
and skipping the ffmpeg mp3 step (WAV plays fine on Android, and ffmpeg
|
||||
isn't installed on this machine).
|
||||
|
||||
Usage: .venv\\Scripts\\python.exe generate_all.py
|
||||
"""
|
||||
|
||||
import wave
|
||||
from pathlib import Path
|
||||
|
||||
from piper import PiperVoice
|
||||
from piper.config import SynthesisConfig
|
||||
|
||||
import robot_filter_v2
|
||||
|
||||
HERE = Path(__file__).parent
|
||||
VOICE_MODEL = HERE / "voices" / "en_US-joe-medium.onnx"
|
||||
OUT_DIR = HERE / "output"
|
||||
|
||||
ENTRIES = {
|
||||
"bulbasaur": (
|
||||
"Bulbasaur, the seed Pokémon. It can be seen napping in bright "
|
||||
"sunlight. There is a seed on its back. By soaking up the sun's "
|
||||
"rays, the seed grows progressively larger."
|
||||
),
|
||||
"charizard": (
|
||||
"Charizard, the flame Pokémon. Charizard flies around the sky in "
|
||||
"search of powerful opponents. It breathes fire of such great "
|
||||
"heat that it melts anything."
|
||||
),
|
||||
"pikachu": (
|
||||
"Pikachu, the mouse Pokémon. When several of these Pokémon "
|
||||
"gather, their electricity could build and cause lightning "
|
||||
"storms."
|
||||
),
|
||||
"eevee": (
|
||||
"Eevee, the evolution Pokémon. Its genetic code is irregular. It "
|
||||
"may mutate if it is exposed to radiation from element stones."
|
||||
),
|
||||
"squirtle": (
|
||||
"Squirtle, the tiny turtle Pokémon. After birth, its back swells "
|
||||
"and hardens into a shell. It powerfully sprays foam from its "
|
||||
"mouth."
|
||||
),
|
||||
}
|
||||
|
||||
ROBOT_INTENSITY = 0.45
|
||||
|
||||
|
||||
SYN_CONFIG = SynthesisConfig(noise_scale=0.5, noise_w_scale=0.3, length_scale=0.98)
|
||||
|
||||
|
||||
def main():
|
||||
OUT_DIR.mkdir(exist_ok=True)
|
||||
print(f"Loading voice model {VOICE_MODEL.name}...")
|
||||
voice = PiperVoice.load(str(VOICE_MODEL))
|
||||
|
||||
for name, text in ENTRIES.items():
|
||||
fixed_path = OUT_DIR / f"{name}_fixed.wav"
|
||||
robot_path = OUT_DIR / f"{name}_fixed_robot.wav"
|
||||
print(f"Synthesizing {name}...")
|
||||
with wave.open(str(fixed_path), "wb") as wav_file:
|
||||
voice.synthesize_wav(text, wav_file, syn_config=SYN_CONFIG)
|
||||
sr, x = robot_filter_v2.load(str(fixed_path))
|
||||
y = robot_filter_v2.robotize(x, sr, intensity=ROBOT_INTENSITY)
|
||||
robot_filter_v2.save(str(robot_path), sr, y)
|
||||
print(f" -> {robot_path.name}")
|
||||
|
||||
print("\nDone. See output/*_fixed_robot.wav")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
74
saved/pokedex_voice_spike_scripts/generate_samples.sh
Normal file
|
|
@ -0,0 +1,74 @@
|
|||
#!/usr/bin/env bash
|
||||
# Reproduces the Pokedex voice spike end to end:
|
||||
# TTS (espeak-ng and Piper) -> robot-voice DSP filter -> mp3
|
||||
#
|
||||
# Setup (Debian/Ubuntu):
|
||||
# sudo apt-get install -y espeak-ng ffmpeg
|
||||
# pip install piper-tts --break-system-packages
|
||||
# pip install numpy scipy --break-system-packages
|
||||
#
|
||||
# Voice model:
|
||||
# Official route: download a Piper voice (.onnx + .onnx.json) from the
|
||||
# Piper voices repo and point --model at it, e.g. en_US-lessac-medium.
|
||||
# In this sandbox, huggingface.co wasn't reachable, so I instead pulled a
|
||||
# community PyPI wheel that bundles the model file directly:
|
||||
# pip download --no-deps joe-us-piper-voice
|
||||
# (unzip the wheel; the .onnx/.onnx.json live under
|
||||
# joe_us_piper_voice/data/). Not an official source -- for production,
|
||||
# get voices from Piper's own releases instead.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
|
||||
VOICE_MODEL="voices/en_US-joe-medium.onnx"
|
||||
export ALSA_CONFIG_PATH=/dev/null # silence ALSA warnings in headless envs
|
||||
|
||||
# ---- Stage 1: raw espeak-ng (formant synth, offline, inherently "robotic") ----
|
||||
espeak-ng -v en-us -s 150 -w bulbasaur_raw.wav \
|
||||
"Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
|
||||
|
||||
espeak-ng -v en-us -s 150 -w charizard_raw.wav \
|
||||
"Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
|
||||
|
||||
espeak-ng -v en-us -s 150 -w pikachu_raw.wav \
|
||||
"Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
|
||||
|
||||
# ---- Stage 2: Piper neural TTS, first pass (natural but bad cadence on rare words) ----
|
||||
gen_neural() {
|
||||
local name="$1" text="$2"
|
||||
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_neural.wav"
|
||||
}
|
||||
gen_neural bulbasaur "Bulbasaur. Seed pokemon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
|
||||
gen_neural charizard "Charizard. Flame pokemon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
|
||||
gen_neural pikachu "Pikachu. Mouse pokemon. When several of these pokemon gather, their electricity could build and cause lightning storms."
|
||||
|
||||
# ---- Stage 3: Piper neural TTS, cadence-fixed pass ----
|
||||
# Fix = (a) standard "X, the Y Pokemon." phrasing instead of two short
|
||||
# sentences, which stopped "Pokemon" from landing phrase-final where
|
||||
# duration models over-lengthen it, and (b) reduced noise-w-scale (duration
|
||||
# randomness) and noise-scale (audio variance) so rare proper nouns don't
|
||||
# get a random stretched/warped rendering.
|
||||
gen_fixed() {
|
||||
local name="$1" text="$2"
|
||||
echo "$text" | piper -m "$VOICE_MODEL" -f "${name}_fixed.wav" \
|
||||
--noise-scale 0.5 --noise-w-scale 0.3 --length-scale 0.98
|
||||
}
|
||||
gen_fixed bulbasaur "Bulbasaur, the seed Pokémon. It can be seen napping in bright sunlight. There is a seed on its back. By soaking up the sun's rays, the seed grows progressively larger."
|
||||
gen_fixed charizard "Charizard, the flame Pokémon. Charizard flies around the sky in search of powerful opponents. It breathes fire of such great heat that it melts anything."
|
||||
gen_fixed pikachu "Pikachu, the mouse Pokémon. When several of these Pokémon gather, their electricity could build and cause lightning storms."
|
||||
|
||||
# ---- Stage 4: robot-voice DSP filter passes ----
|
||||
# robot_filter.py -> heavy/original (square-wave ring mod + bitcrush + narrow bandpass)
|
||||
# robot_filter_light.py -> barely-there (sine ring mod, wide bandpass, mild drive)
|
||||
# robot_filter_v2.py -> parameterized 0..1 intensity knob (used at 0.45 = "medium")
|
||||
for name in bulbasaur charizard pikachu; do
|
||||
python3 robot_filter.py "${name}_raw.wav" "${name}_robot.wav"
|
||||
python3 robot_filter_light.py "${name}_neural.wav" "${name}_neural_light.wav"
|
||||
python3 robot_filter_v2.py "${name}_fixed.wav" "${name}_fixed_robot.wav" 0.45
|
||||
done
|
||||
|
||||
# ---- Stage 5: mp3 for easy playback/delivery ----
|
||||
for f in *_raw.wav *_robot.wav *_neural.wav *_neural_light.wav *_fixed.wav *_fixed_robot.wav; do
|
||||
[ -f "$f" ] && ffmpeg -y -loglevel error -i "$f" -codec:a libmp3lame -qscale:a 4 "${f%.wav}.mp3"
|
||||
done
|
||||
|
||||
echo "Done. See *.mp3 for output."
|
||||
BIN
saved/pokedex_voice_spike_scripts/output/bulbasaur_fixed.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/charizard_fixed.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/eevee_fixed.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/eevee_fixed_robot.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/pikachu_fixed.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/pikachu_fixed_robot.wav
Normal file
BIN
saved/pokedex_voice_spike_scripts/output/squirtle_fixed.wav
Normal file
79
saved/pokedex_voice_spike_scripts/robot_filter.py
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
"""
|
||||
Quick-and-dirty 'robot voice' post-processor for the Pokedex voice spike.
|
||||
|
||||
Pipeline: TTS wav in -> ring modulation + bitcrush + band-pass (speaker-grille
|
||||
coloring) + light hard-clip drive -> wav out.
|
||||
|
||||
Usage: python3 robot_filter.py in.wav out.wav
|
||||
"""
|
||||
import sys
|
||||
import numpy as np
|
||||
from scipy.io import wavfile
|
||||
from scipy.signal import butter, sosfilt
|
||||
|
||||
|
||||
def load(path):
|
||||
sr, data = wavfile.read(path)
|
||||
if data.dtype == np.int16:
|
||||
x = data.astype(np.float64) / 32768.0
|
||||
elif data.dtype == np.int32:
|
||||
x = data.astype(np.float64) / 2147483648.0
|
||||
elif data.dtype == np.uint8:
|
||||
x = (data.astype(np.float64) - 128) / 128.0
|
||||
else:
|
||||
x = data.astype(np.float64)
|
||||
if x.ndim > 1:
|
||||
x = x.mean(axis=1)
|
||||
return sr, x
|
||||
|
||||
|
||||
def save(path, sr, x):
|
||||
x = np.clip(x, -1.0, 1.0)
|
||||
wavfile.write(path, sr, (x * 32767).astype(np.int16))
|
||||
|
||||
|
||||
def ring_modulate(x, sr, carrier_hz=45.0, mix=0.55):
|
||||
t = np.arange(len(x)) / sr
|
||||
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) # square carrier = harsher/metallic
|
||||
modulated = x * carrier
|
||||
return (1 - mix) * x + mix * modulated
|
||||
|
||||
|
||||
def bitcrush(x, bit_depth=6, sr=None, downsample_factor=3):
|
||||
# sample-and-hold downsample (classic lo-fi/robotic stepping)
|
||||
if downsample_factor > 1:
|
||||
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
|
||||
if len(held) < len(x):
|
||||
held = np.pad(held, (0, len(x) - len(held)))
|
||||
x = held
|
||||
levels = 2 ** bit_depth
|
||||
x = np.round(x * levels) / levels
|
||||
return x
|
||||
|
||||
|
||||
def bandpass(x, sr, low=350.0, high=3400.0, order=4):
|
||||
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
|
||||
return sosfilt(sos, x)
|
||||
|
||||
|
||||
def drive(x, amount=1.8):
|
||||
return np.tanh(x * amount) / np.tanh(amount)
|
||||
|
||||
|
||||
def robotize(x, sr):
|
||||
y = ring_modulate(x, sr, carrier_hz=45.0, mix=0.5)
|
||||
y = bitcrush(y, bit_depth=7, downsample_factor=2)
|
||||
y = bandpass(y, sr, low=300, high=3800)
|
||||
y = drive(y, amount=1.6)
|
||||
# normalize
|
||||
peak = np.max(np.abs(y)) or 1.0
|
||||
y = y / peak * 0.9
|
||||
return y
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
src, dst = sys.argv[1], sys.argv[2]
|
||||
sr, x = load(src)
|
||||
y = robotize(x, sr)
|
||||
save(dst, sr, y)
|
||||
print(f"wrote {dst}")
|
||||
58
saved/pokedex_voice_spike_scripts/robot_filter_light.py
Normal file
|
|
@ -0,0 +1,58 @@
|
|||
"""
|
||||
Light-touch robot coloring for a neural (already-natural-sounding) TTS voice.
|
||||
Goal: keep natural cadence/prosody intact, just add a synthesized/device edge.
|
||||
"""
|
||||
import sys
|
||||
import numpy as np
|
||||
from scipy.io import wavfile
|
||||
from scipy.signal import butter, sosfilt
|
||||
|
||||
|
||||
def load(path):
|
||||
sr, data = wavfile.read(path)
|
||||
if data.dtype == np.int16:
|
||||
x = data.astype(np.float64) / 32768.0
|
||||
elif data.dtype == np.int32:
|
||||
x = data.astype(np.float64) / 2147483648.0
|
||||
else:
|
||||
x = data.astype(np.float64)
|
||||
if x.ndim > 1:
|
||||
x = x.mean(axis=1)
|
||||
return sr, x
|
||||
|
||||
|
||||
def save(path, sr, x):
|
||||
x = np.clip(x, -1.0, 1.0)
|
||||
wavfile.write(path, sr, (x * 32767).astype(np.int16))
|
||||
|
||||
|
||||
def ring_modulate(x, sr, carrier_hz=90.0, mix=0.12):
|
||||
t = np.arange(len(x)) / sr
|
||||
carrier = np.sin(2 * np.pi * carrier_hz * t) # sine carrier = subtler than square
|
||||
return (1 - mix) * x + mix * (x * carrier)
|
||||
|
||||
|
||||
def bandpass(x, sr, low=200.0, high=5500.0, order=2):
|
||||
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
|
||||
return sosfilt(sos, x)
|
||||
|
||||
|
||||
def drive(x, amount=1.15):
|
||||
return np.tanh(x * amount) / np.tanh(amount)
|
||||
|
||||
|
||||
def robotize_light(x, sr):
|
||||
y = ring_modulate(x, sr, carrier_hz=90.0, mix=0.12)
|
||||
y = bandpass(y, sr, low=200, high=5500)
|
||||
y = drive(y, amount=1.15)
|
||||
peak = np.max(np.abs(y)) or 1.0
|
||||
y = y / peak * 0.9
|
||||
return y
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
src, dst = sys.argv[1], sys.argv[2]
|
||||
sr, x = load(src)
|
||||
y = robotize_light(x, sr)
|
||||
save(dst, sr, y)
|
||||
print(f"wrote {dst}")
|
||||
72
saved/pokedex_voice_spike_scripts/robot_filter_v2.py
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
"""
|
||||
Parameterized robot-voice filter, v2.
|
||||
intensity in [0, 1]: 0 = untouched, 1 = full "heavy" robot from round 1.
|
||||
"""
|
||||
import sys
|
||||
import numpy as np
|
||||
from scipy.io import wavfile
|
||||
from scipy.signal import butter, sosfilt
|
||||
|
||||
|
||||
def load(path):
|
||||
sr, data = wavfile.read(path)
|
||||
if data.dtype == np.int16:
|
||||
x = data.astype(np.float64) / 32768.0
|
||||
elif data.dtype == np.int32:
|
||||
x = data.astype(np.float64) / 2147483648.0
|
||||
else:
|
||||
x = data.astype(np.float64)
|
||||
if x.ndim > 1:
|
||||
x = x.mean(axis=1)
|
||||
return sr, x
|
||||
|
||||
|
||||
def save(path, sr, x):
|
||||
x = np.clip(x, -1.0, 1.0)
|
||||
wavfile.write(path, sr, (x * 32767).astype(np.int16))
|
||||
|
||||
|
||||
def ring_modulate(x, sr, carrier_hz, mix, square=False):
|
||||
t = np.arange(len(x)) / sr
|
||||
carrier = np.sign(np.sin(2 * np.pi * carrier_hz * t)) if square else np.sin(2 * np.pi * carrier_hz * t)
|
||||
return (1 - mix) * x + mix * (x * carrier)
|
||||
|
||||
|
||||
def bitcrush(x, bit_depth, downsample_factor):
|
||||
if downsample_factor > 1:
|
||||
held = np.repeat(x[::downsample_factor], downsample_factor)[: len(x)]
|
||||
if len(held) < len(x):
|
||||
held = np.pad(held, (0, len(x) - len(held)))
|
||||
x = held
|
||||
levels = 2 ** bit_depth
|
||||
return np.round(x * levels) / levels
|
||||
|
||||
|
||||
def bandpass(x, sr, low, high, order=3):
|
||||
sos = butter(order, [low, high], btype="band", fs=sr, output="sos")
|
||||
return sosfilt(sos, x)
|
||||
|
||||
|
||||
def drive(x, amount):
|
||||
return np.tanh(x * amount) / np.tanh(amount)
|
||||
|
||||
|
||||
def robotize(x, sr, intensity=0.45):
|
||||
i = intensity
|
||||
y = ring_modulate(x, sr, carrier_hz=60 - 15 * i, mix=0.1 + 0.45 * i, square=(i > 0.6))
|
||||
if i > 0.35:
|
||||
y = bitcrush(y, bit_depth=int(12 - 6 * i), downsample_factor=1 if i < 0.7 else 2)
|
||||
y = bandpass(y, sr, low=250 - 100 * i, high=6000 - 2200 * i)
|
||||
y = drive(y, amount=1.0 + 0.9 * i)
|
||||
peak = np.max(np.abs(y)) or 1.0
|
||||
y = y / peak * 0.9
|
||||
return y
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
src, dst = sys.argv[1], sys.argv[2]
|
||||
intensity = float(sys.argv[3]) if len(sys.argv) > 3 else 0.45
|
||||
sr, x = load(src)
|
||||
y = robotize(x, sr, intensity=intensity)
|
||||
save(dst, sr, y)
|
||||
print(f"wrote {dst} (intensity={intensity})")
|
||||
BIN
saved/pokedex_voice_spike_scripts/voices/en_US-joe-medium.onnx
Normal file
|
|
@ -0,0 +1,484 @@
|
|||
{
|
||||
"audio": {
|
||||
"sample_rate": 22050,
|
||||
"quality": "medium"
|
||||
},
|
||||
"espeak": {
|
||||
"voice": "en-us"
|
||||
},
|
||||
"inference": {
|
||||
"noise_scale": 0.667,
|
||||
"length_scale": 1,
|
||||
"noise_w": 0.8
|
||||
},
|
||||
"phoneme_type": "espeak",
|
||||
"phoneme_map": {},
|
||||
"phoneme_id_map": {
|
||||
"_": [
|
||||
0
|
||||
],
|
||||
"^": [
|
||||
1
|
||||
],
|
||||
"$": [
|
||||
2
|
||||
],
|
||||
" ": [
|
||||
3
|
||||
],
|
||||
"!": [
|
||||
4
|
||||
],
|
||||
"'": [
|
||||
5
|
||||
],
|
||||
"(": [
|
||||
6
|
||||
],
|
||||
")": [
|
||||
7
|
||||
],
|
||||
",": [
|
||||
8
|
||||
],
|
||||
"-": [
|
||||
9
|
||||
],
|
||||
".": [
|
||||
10
|
||||
],
|
||||
":": [
|
||||
11
|
||||
],
|
||||
";": [
|
||||
12
|
||||
],
|
||||
"?": [
|
||||
13
|
||||
],
|
||||
"a": [
|
||||
14
|
||||
],
|
||||
"b": [
|
||||
15
|
||||
],
|
||||
"c": [
|
||||
16
|
||||
],
|
||||
"d": [
|
||||
17
|
||||
],
|
||||
"e": [
|
||||
18
|
||||
],
|
||||
"f": [
|
||||
19
|
||||
],
|
||||
"h": [
|
||||
20
|
||||
],
|
||||
"i": [
|
||||
21
|
||||
],
|
||||
"j": [
|
||||
22
|
||||
],
|
||||
"k": [
|
||||
23
|
||||
],
|
||||
"l": [
|
||||
24
|
||||
],
|
||||
"m": [
|
||||
25
|
||||
],
|
||||
"n": [
|
||||
26
|
||||
],
|
||||
"o": [
|
||||
27
|
||||
],
|
||||
"p": [
|
||||
28
|
||||
],
|
||||
"q": [
|
||||
29
|
||||
],
|
||||
"r": [
|
||||
30
|
||||
],
|
||||
"s": [
|
||||
31
|
||||
],
|
||||
"t": [
|
||||
32
|
||||
],
|
||||
"u": [
|
||||
33
|
||||
],
|
||||
"v": [
|
||||
34
|
||||
],
|
||||
"w": [
|
||||
35
|
||||
],
|
||||
"x": [
|
||||
36
|
||||
],
|
||||
"y": [
|
||||
37
|
||||
],
|
||||
"z": [
|
||||
38
|
||||
],
|
||||
"æ": [
|
||||
39
|
||||
],
|
||||
"ç": [
|
||||
40
|
||||
],
|
||||
"ð": [
|
||||
41
|
||||
],
|
||||
"ø": [
|
||||
42
|
||||
],
|
||||
"ħ": [
|
||||
43
|
||||
],
|
||||
"ŋ": [
|
||||
44
|
||||
],
|
||||
"œ": [
|
||||
45
|
||||
],
|
||||
"ǀ": [
|
||||
46
|
||||
],
|
||||
"ǁ": [
|
||||
47
|
||||
],
|
||||
"ǂ": [
|
||||
48
|
||||
],
|
||||
"ǃ": [
|
||||
49
|
||||
],
|
||||
"ɐ": [
|
||||
50
|
||||
],
|
||||
"ɑ": [
|
||||
51
|
||||
],
|
||||
"ɒ": [
|
||||
52
|
||||
],
|
||||
"ɓ": [
|
||||
53
|
||||
],
|
||||
"ɔ": [
|
||||
54
|
||||
],
|
||||
"ɕ": [
|
||||
55
|
||||
],
|
||||
"ɖ": [
|
||||
56
|
||||
],
|
||||
"ɗ": [
|
||||
57
|
||||
],
|
||||
"ɘ": [
|
||||
58
|
||||
],
|
||||
"ə": [
|
||||
59
|
||||
],
|
||||
"ɚ": [
|
||||
60
|
||||
],
|
||||
"ɛ": [
|
||||
61
|
||||
],
|
||||
"ɜ": [
|
||||
62
|
||||
],
|
||||
"ɞ": [
|
||||
63
|
||||
],
|
||||
"ɟ": [
|
||||
64
|
||||
],
|
||||
"ɠ": [
|
||||
65
|
||||
],
|
||||
"ɡ": [
|
||||
66
|
||||
],
|
||||
"ɢ": [
|
||||
67
|
||||
],
|
||||
"ɣ": [
|
||||
68
|
||||
],
|
||||
"ɤ": [
|
||||
69
|
||||
],
|
||||
"ɥ": [
|
||||
70
|
||||
],
|
||||
"ɦ": [
|
||||
71
|
||||
],
|
||||
"ɧ": [
|
||||
72
|
||||
],
|
||||
"ɨ": [
|
||||
73
|
||||
],
|
||||
"ɪ": [
|
||||
74
|
||||
],
|
||||
"ɫ": [
|
||||
75
|
||||
],
|
||||
"ɬ": [
|
||||
76
|
||||
],
|
||||
"ɭ": [
|
||||
77
|
||||
],
|
||||
"ɮ": [
|
||||
78
|
||||
],
|
||||
"ɯ": [
|
||||
79
|
||||
],
|
||||
"ɰ": [
|
||||
80
|
||||
],
|
||||
"ɱ": [
|
||||
81
|
||||
],
|
||||
"ɲ": [
|
||||
82
|
||||
],
|
||||
"ɳ": [
|
||||
83
|
||||
],
|
||||
"ɴ": [
|
||||
84
|
||||
],
|
||||
"ɵ": [
|
||||
85
|
||||
],
|
||||
"ɶ": [
|
||||
86
|
||||
],
|
||||
"ɸ": [
|
||||
87
|
||||
],
|
||||
"ɹ": [
|
||||
88
|
||||
],
|
||||
"ɺ": [
|
||||
89
|
||||
],
|
||||
"ɻ": [
|
||||
90
|
||||
],
|
||||
"ɽ": [
|
||||
91
|
||||
],
|
||||
"ɾ": [
|
||||
92
|
||||
],
|
||||
"ʀ": [
|
||||
93
|
||||
],
|
||||
"ʁ": [
|
||||
94
|
||||
],
|
||||
"ʂ": [
|
||||
95
|
||||
],
|
||||
"ʃ": [
|
||||
96
|
||||
],
|
||||
"ʄ": [
|
||||
97
|
||||
],
|
||||
"ʈ": [
|
||||
98
|
||||
],
|
||||
"ʉ": [
|
||||
99
|
||||
],
|
||||
"ʊ": [
|
||||
100
|
||||
],
|
||||
"ʋ": [
|
||||
101
|
||||
],
|
||||
"ʌ": [
|
||||
102
|
||||
],
|
||||
"ʍ": [
|
||||
103
|
||||
],
|
||||
"ʎ": [
|
||||
104
|
||||
],
|
||||
"ʏ": [
|
||||
105
|
||||
],
|
||||
"ʐ": [
|
||||
106
|
||||
],
|
||||
"ʑ": [
|
||||
107
|
||||
],
|
||||
"ʒ": [
|
||||
108
|
||||
],
|
||||
"ʔ": [
|
||||
109
|
||||
],
|
||||
"ʕ": [
|
||||
110
|
||||
],
|
||||
"ʘ": [
|
||||
111
|
||||
],
|
||||
"ʙ": [
|
||||
112
|
||||
],
|
||||
"ʛ": [
|
||||
113
|
||||
],
|
||||
"ʜ": [
|
||||
114
|
||||
],
|
||||
"ʝ": [
|
||||
115
|
||||
],
|
||||
"ʟ": [
|
||||
116
|
||||
],
|
||||
"ʡ": [
|
||||
117
|
||||
],
|
||||
"ʢ": [
|
||||
118
|
||||
],
|
||||
"ʲ": [
|
||||
119
|
||||
],
|
||||
"ˈ": [
|
||||
120
|
||||
],
|
||||
"ˌ": [
|
||||
121
|
||||
],
|
||||
"ː": [
|
||||
122
|
||||
],
|
||||
"ˑ": [
|
||||
123
|
||||
],
|
||||
"˞": [
|
||||
124
|
||||
],
|
||||
"β": [
|
||||
125
|
||||
],
|
||||
"θ": [
|
||||
126
|
||||
],
|
||||
"χ": [
|
||||
127
|
||||
],
|
||||
"ᵻ": [
|
||||
128
|
||||
],
|
||||
"ⱱ": [
|
||||
129
|
||||
],
|
||||
"0": [
|
||||
130
|
||||
],
|
||||
"1": [
|
||||
131
|
||||
],
|
||||
"2": [
|
||||
132
|
||||
],
|
||||
"3": [
|
||||
133
|
||||
],
|
||||
"4": [
|
||||
134
|
||||
],
|
||||
"5": [
|
||||
135
|
||||
],
|
||||
"6": [
|
||||
136
|
||||
],
|
||||
"7": [
|
||||
137
|
||||
],
|
||||
"8": [
|
||||
138
|
||||
],
|
||||
"9": [
|
||||
139
|
||||
],
|
||||
"̧": [
|
||||
140
|
||||
],
|
||||
"̃": [
|
||||
141
|
||||
],
|
||||
"̪": [
|
||||
142
|
||||
],
|
||||
"̯": [
|
||||
143
|
||||
],
|
||||
"̩": [
|
||||
144
|
||||
],
|
||||
"ʰ": [
|
||||
145
|
||||
],
|
||||
"ˤ": [
|
||||
146
|
||||
],
|
||||
"ε": [
|
||||
147
|
||||
],
|
||||
"↓": [
|
||||
148
|
||||
],
|
||||
"#": [
|
||||
149
|
||||
],
|
||||
"\"": [
|
||||
150
|
||||
]
|
||||
},
|
||||
"num_symbols": 256,
|
||||
"num_speakers": 1,
|
||||
"speaker_id_map": {},
|
||||
"piper_version": "1.0.0",
|
||||
"language": {
|
||||
"code": "en_US",
|
||||
"family": "en",
|
||||
"region": "US",
|
||||
"name_native": "English",
|
||||
"name_english": "English",
|
||||
"country_english": "United States"
|
||||
},
|
||||
"dataset": "joe"
|
||||
}
|
||||
BIN
saved/pokemon_images/Post-Thumbnail-Ratio-1.jpg
Normal file
|
After Width: | Height: | Size: 104 KiB |
BIN
saved/pokemon_images/SM35_EN_1.png
Normal file
|
After Width: | Height: | Size: 135 KiB |
BIN
saved/pokemon_images/art_bulbasaur.png
Normal file
|
After Width: | Height: | Size: 70 KiB |
BIN
saved/pokemon_images/art_charizard.png
Normal file
|
After Width: | Height: | Size: 54 KiB |
BIN
saved/pokemon_images/art_eevee.png
Normal file
|
After Width: | Height: | Size: 52 KiB |
BIN
saved/pokemon_images/art_pikachu.png
Normal file
|
After Width: | Height: | Size: 42 KiB |
BIN
saved/pokemon_images/art_squirtle.png
Normal file
|
After Width: | Height: | Size: 54 KiB |
BIN
saved/pokemon_images/charizard.png
Normal file
|
After Width: | Height: | Size: 1.1 KiB |
BIN
saved/pokemon_images/filters_quality(95)format(webp) (1).webp
Normal file
|
After Width: | Height: | Size: 70 KiB |
BIN
saved/pokemon_images/filters_quality(95)format(webp).webp
Normal file
|
After Width: | Height: | Size: 71 KiB |
BIN
saved/pokemon_images/query_bulbasaur.jpg
Normal file
|
After Width: | Height: | Size: 49 KiB |
BIN
saved/pokemon_images/query_charizard.jpg
Normal file
|
After Width: | Height: | Size: 14 KiB |
BIN
saved/pokemon_images/query_charizard_amazon.jpg
Normal file
|
After Width: | Height: | Size: 15 KiB |
BIN
saved/pokemon_images/query_charizard_plush.webp
Normal file
|
After Width: | Height: | Size: 45 KiB |
BIN
saved/pokemon_images/query_eevee.jpg
Normal file
|
After Width: | Height: | Size: 96 KiB |
BIN
saved/pokemon_images/query_pikachu.jpg
Normal file
|
After Width: | Height: | Size: 33 KiB |
BIN
saved/pokemon_images/query_squirtle.jpg
Normal file
|
After Width: | Height: | Size: 84 KiB |
BIN
saved/pokemon_images/ref_bulbasaur.jpg
Normal file
|
After Width: | Height: | Size: 114 KiB |
BIN
saved/pokemon_images/ref_charizard.jpg
Normal file
|
After Width: | Height: | Size: 161 KiB |
BIN
saved/pokemon_images/ref_eevee.jpg
Normal file
|
After Width: | Height: | Size: 145 KiB |
BIN
saved/pokemon_images/ref_pikachu.jpg
Normal file
|
After Width: | Height: | Size: 57 KiB |
BIN
saved/pokemon_images/ref_squirtle.jpg
Normal file
|
After Width: | Height: | Size: 318 KiB |