reddex/saved/image_recognition_spike/REPORT.md
forgejoadmin 71db2d1ab9 Initial commit
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-18 03:16:42 -04:00

3.8 KiB

Image recognition de-risking — report

Goal: figure out how the app recognizes a Pokemon when the phone is pointed at a figure/toy, before committing to an architecture.

Test set

5 classes (Bulbasaur, Squirtle, Pikachu, Eevee, Charizard), grown over the course of the spike to 12 query images spanning very different visual styles on purpose: stock photos, Pokemon GO AR screenshots, plush toys (including two real photos of figures the user owns), a TCG card, a tiny battle sprite, and anime screencaps. Two reference sets were tried: scraped stock photos, and Bulbapedia's official artwork.

Approaches tried, in order

1. CLIP (ViT-B-32, OpenAI weights) + nearest-neighbor

Embed a reference image per species, embed the query, cosine-similarity match. Initial run used the wrong open_clip model variant (ViT-B-32 instead of ViT-B-32-quickgelu), which silently degrades OpenAI-weight embeddings — worth remembering if this comes up again.

  • Photo references: 4/12 correct top-1.
  • Bulbapedia art references: 8/12 correct top-1.

2. DINOv2 (facebook/dinov2-base) + nearest-neighbor

Self-supervised, built for instance/visual similarity rather than text-image alignment — expected to beat CLIP at this specific task.

  • Photo references: 6/12.
  • Bulbapedia art references: 8/12.

Finding across both: official art references consistently beat random stock-photo references, regardless of model — the illustration-vs-photo domain gap we worried about mattered less than material/form-factor mismatch (plush vs. rigid figure vs. flat art). charizard_plush failed in literally every embedding config tried (0/4) — plush toys are the genuinely hard case for this whole approach, not photos-vs-domain style.

3. Direct vision-LLM recognition (the "cheat")

Skip reference images and embeddings entirely — ask a vision-language model "what Pokemon is this" and let its pretrained world knowledge do the work.

  • Claude (me, just looking at the images): 12/12.
  • Self-hosted Qwen2.5-VL-3B-Instruct, run locally on the RTX 5070 Ti: 10/12 cold, no fine-tuning, no reference images at all. The 2 misses were the two genuinely hardest images in the set (a tiny 213x240 keychain thumbnail → correctly returned "unknown" rather than a wrong guess; and a plush the user themselves said "looks like shit, not even sure that's Charizard").

Decision

Went with self-hosted VLM recognition (Qwen2.5-VL-3B-Instruct) over the embedding/nearest-neighbor approach. Reasons:

  • Meaningfully higher accuracy (10/12 vs. best embedding score of 8/12).
  • No reference-image sourcing/maintenance needed for 1000+ species — eliminates the "content volume" risk from the original risk assessment entirely.
  • Degrades safely: genuinely ambiguous images tend to get "unknown" rather than a confident wrong answer.

Trade-off accepted: requires a GPU server reachable over the network at recognition time (already an accepted dependency — the original plan always involved uploading the photo to a home server).

What shipped from this

server/ — FastAPI wrapper around the same Qwen2.5-VL-3B pipeline, running on this Windows machine (chosen over buying a GPU for Unraid). POST /identify takes a photo, returns {recognized, species, raw_response}. Verified working end-to-end over real HTTP.

Open questions / not yet tested

  • Accuracy at real scale (1000+ candidate species) is untested — only ever tried 5 classes. Confusion likely increases with more classes.
  • Never tested against the user's own figures except for two Charizard photos (both plush) — the real target (rigid painted figures) hasn't been tried.
  • Larger models (Qwen2.5-VL-7B+) not tried — likely closes some of the remaining gap to the 12/12 upper bound, at the cost of latency/VRAM.
  • No latency/throughput measurement done — only correctness.