Initial commit
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
commit
71db2d1ab9
126 changed files with 3198 additions and 0 deletions
87
saved/image_recognition_spike/REPORT.md
Normal file
87
saved/image_recognition_spike/REPORT.md
Normal file
|
|
@ -0,0 +1,87 @@
|
|||
# Image recognition de-risking — report
|
||||
|
||||
Goal: figure out how the app recognizes a Pokemon when the phone is
|
||||
pointed at a figure/toy, before committing to an architecture.
|
||||
|
||||
## Test set
|
||||
|
||||
5 classes (Bulbasaur, Squirtle, Pikachu, Eevee, Charizard), grown over the
|
||||
course of the spike to 12 query images spanning very different visual
|
||||
styles on purpose: stock photos, Pokemon GO AR screenshots, plush toys
|
||||
(including two real photos of figures the user owns), a TCG card, a tiny
|
||||
battle sprite, and anime screencaps. Two reference sets were tried:
|
||||
scraped stock photos, and Bulbapedia's official artwork.
|
||||
|
||||
## Approaches tried, in order
|
||||
|
||||
### 1. CLIP (ViT-B-32, OpenAI weights) + nearest-neighbor
|
||||
|
||||
Embed a reference image per species, embed the query, cosine-similarity
|
||||
match. Initial run used the wrong open_clip model variant (`ViT-B-32`
|
||||
instead of `ViT-B-32-quickgelu`), which silently degrades OpenAI-weight
|
||||
embeddings — worth remembering if this comes up again.
|
||||
|
||||
- Photo references: **4/12** correct top-1.
|
||||
- Bulbapedia art references: **8/12** correct top-1.
|
||||
|
||||
### 2. DINOv2 (facebook/dinov2-base) + nearest-neighbor
|
||||
|
||||
Self-supervised, built for instance/visual similarity rather than
|
||||
text-image alignment — expected to beat CLIP at this specific task.
|
||||
|
||||
- Photo references: **6/12**.
|
||||
- Bulbapedia art references: **8/12**.
|
||||
|
||||
**Finding across both:** official art references consistently beat random
|
||||
stock-photo references, regardless of model — the illustration-vs-photo
|
||||
domain gap we worried about mattered less than material/form-factor
|
||||
mismatch (plush vs. rigid figure vs. flat art). `charizard_plush` failed
|
||||
in literally every embedding config tried (0/4) — plush toys are the
|
||||
genuinely hard case for this whole approach, not photos-vs-domain style.
|
||||
|
||||
### 3. Direct vision-LLM recognition (the "cheat")
|
||||
|
||||
Skip reference images and embeddings entirely — ask a vision-language
|
||||
model "what Pokemon is this" and let its pretrained world knowledge do
|
||||
the work.
|
||||
|
||||
- Claude (me, just looking at the images): **12/12**.
|
||||
- Self-hosted Qwen2.5-VL-3B-Instruct, run locally on the RTX 5070 Ti:
|
||||
**10/12** cold, no fine-tuning, no reference images at all. The 2
|
||||
misses were the two genuinely hardest images in the set (a tiny
|
||||
213x240 keychain thumbnail → correctly returned "unknown" rather than
|
||||
a wrong guess; and a plush the user themselves said "looks like shit,
|
||||
not even sure that's Charizard").
|
||||
|
||||
## Decision
|
||||
|
||||
Went with **self-hosted VLM recognition** (Qwen2.5-VL-3B-Instruct) over
|
||||
the embedding/nearest-neighbor approach. Reasons:
|
||||
- Meaningfully higher accuracy (10/12 vs. best embedding score of 8/12).
|
||||
- No reference-image sourcing/maintenance needed for 1000+ species —
|
||||
eliminates the "content volume" risk from the original risk assessment
|
||||
entirely.
|
||||
- Degrades safely: genuinely ambiguous images tend to get "unknown"
|
||||
rather than a confident wrong answer.
|
||||
|
||||
Trade-off accepted: requires a GPU server reachable over the network at
|
||||
recognition time (already an accepted dependency — the original plan
|
||||
always involved uploading the photo to a home server).
|
||||
|
||||
## What shipped from this
|
||||
|
||||
`server/` — FastAPI wrapper around the same Qwen2.5-VL-3B pipeline,
|
||||
running on this Windows machine (chosen over buying a GPU for Unraid).
|
||||
`POST /identify` takes a photo, returns `{recognized, species,
|
||||
raw_response}`. Verified working end-to-end over real HTTP.
|
||||
|
||||
## Open questions / not yet tested
|
||||
|
||||
- Accuracy at real scale (1000+ candidate species) is untested — only
|
||||
ever tried 5 classes. Confusion likely increases with more classes.
|
||||
- Never tested against the user's own figures except for two Charizard
|
||||
photos (both plush) — the real target (rigid painted figures) hasn't
|
||||
been tried.
|
||||
- Larger models (Qwen2.5-VL-7B+) not tried — likely closes some of the
|
||||
remaining gap to the 12/12 upper bound, at the cost of latency/VRAM.
|
||||
- No latency/throughput measurement done — only correctness.
|
||||
Loading…
Add table
Add a link
Reference in a new issue