π Two-Stage Cross-Encoder & Multimodal Vector Playground
Enter any free-form query (try red wings and body or yellow body black wings) or click a curated chip. Compare Stage 1 Bi-Encoder MaxSim (384-d), Stage 2 Cross-Attention Reranker, and Direct 512-d Visual CLIP scores side-by-side.
βοΈ Engine Mode, Vector Telemetry & Preset Chips β Active Mode: ποΈ Multimodal (CE + CLIP) Expand / Collapse βΎ
βοΈ Chunk Weight Tuning Lab
In fine-grained MaxSim search, distinct chunks receive customized multipliers. Anatomy (e.g. head: 1.10) and Plumage (color_profile: 1.20) score higher to reward specific field marks, while generic tags: 0.65 are downweighted to prevent noise. Adjust any weight below to see how photo rankings adapt in real time!
𧬠Query Diversity & Curation Simulator
Test candidate search discovery queries against the 0.72 pairwise cosine similarity cutoff enforced by your curation pipeline. If a query is too similar to an existing chip (e.g. standing on a rock vs perched on a stone > 0.72), it is disqualified to prevent redundant clutter.
Nearest Matching Curated Queries:
| Existing Curated Query | Category | Cosine Sim | Diversity Meter |
|---|---|---|---|
| Type a candidate query above to view nearest matches. | |||
ποΈ Current Production Curated Queries (77 Total)
These are the 77 high-discrimination queries discovered by the LLM Curator (qwen3.8:27b) with zero species leakage and guaranteed embedding diversity (≤ 0.72).
π¬ Bird Photo Semantic Decomposition
Pick any bird in your catalog to see how JoyCaption + LLM refinement broke its photo down into 15 to 19 distinct visual chunks for search indexing.
Indexed Semantic Chunks:
π How Fine-Grained Section MaxSim Search Works
1. The "Whole-Caption Dilution" Problem
If you feed a 250-word bird description into an embedding model to get one single vector, that vector represents the average meaning of the whole text. Specific visual detailsβlike golden crown patch or rusty barbed wireβget diluted by the remaining 200 words talking about habitat, lighting, and wings. Searches for specific field marks end up scoring poorly.
2. Fine-Grained Semantic Chunking
Instead of 1 vector per photo, the pipeline decomposes each bird photo description into 10 to 19 distinct chunks (e.g. head, beak, wings, action, setting, lighting, color_profile). Each chunk gets its own 384-dimensional vector coordinate.
3. Section-Level MaxSim Scoring
When a search query is submitted, the engine computes the cosine similarity between the query vector and every chunk vector of a photo:
If the photo's setting chunk matches rusty barbed wire with a 0.88 similarity score, that single high match catapults the photo to the top of the search results!
4. Int8 Quantization & In-Memory Binary Buffer
With ~2,134 photos producing ~60,000 chunk vectors, storing 384 raw 32-bit floats would consume ~92 MB of RAM. The script quantizes every float vector into signed 8-bit integers (int8) and packs them into a single 7.64 MB binary buffer (chunk_vectors.bin). The Node.js server scans all 60,000 vectors in under 15 milliseconds with zero garbage collection overhead.
5. Curation & Embedding Diversity (≤ 0.72)
When mining discovery queries, raw LLM outputs often produce near-duplicate paraphrases (e.g. standing on a rock vs perched on a rock, cosine similarity 0.98). By calculating pairwise cosine distances with all-MiniLM-L6-v2 and enforcing a strict ≤ 0.72 ceiling, the engine guarantees every discovery chip reveals a truly distinct visual mark, posture, or environment.
6. Solving the Bi-Encoder "Color-Attribute Binding" Flaw (Stage 2 ONNX Cross-Encoder)
Pure bi-encoder models (all-MiniLM-L6-v2) embed the query and document chunks independently into fixed 384-d vectors. On compositional queries like red wings and body, high-frequency anatomical nouns (wings, body) dominate the vector direction. As a result, a black-and-white Bufflehead with detailed wing/body descriptions scored 0.618 in Stage 1 MaxSim, beating truly red birds like Northern Cardinal (0.565, Rank #14) and Scarlet Tanager (0.545, Rank #20).
To eliminate this flaw without sacrificing speed, the engine runs a Two-Stage Retrieval & Cross-Attention Reranking pipeline in ~15β28ms:
- Stage 1 Recall + Section Binding Check (
evaluateSectionBinding): Scans all ~60,000 Int8 chunk vectors in ~5ms and checks whether the query pairs color words (red,yellow,black, etc.) with structured VLM anatomy keys (head,beak,body,wings,tail). Photos matching all queried color-section pairs receive a recall bonus (+0.25) so they are guaranteed to enter the Top-16 rerank pool, while mismatches receive a demotion multiplier (sectionMultiplier = 0.12xor0.72x). - Stage 2 ONNX Cross-Encoder (
ms-marco-MiniLM-L-6-v2/bge-reranker-base): Concatenates[CLS] Query [SEP] Passage [SEP]across the Top-16 candidates and performs full token-by-token cross-attention, promoting Northern Cardinal (+13 ranks to #1, score0.865) and Scarlet Tanager (+18 ranks to #2, score0.860) while crushing Bufflehead to0.057.
7. Adaptive Passage Routing & Signed Int8 Fix: Why General Queries Like "prominent ear tufts" Require Dynamic Windowing
Because the ONNX Cross-Encoder runs with a compact 52-token sequence window to maintain ~15β25ms CPU latency, how the candidate passage is assembled in buildCrossEncoderPassage() is critical. Three subtle pitfalls were identified and solved:
Statically placing Body: ... and Wings: ... at the start of every passage starved non-body/wing queries like prominent ear tuftsβthe Cross-Encoder only saw brown body/wing descriptions and scored owls at 0.00. Now, buildCrossEncoderPassage() dynamically branches:
β’ Color-Anatomy Mode: Places queried anatomy sections (Body:, Wings:) first.
β’ Winning-Chunk Priority Mode: Places the Stage-1 winning chunk (winText, e.g. "Crest crown: Prominent ear tufts") and prominent_features first.
vlm_captions.json indexes 5 structured section keys: ['head', 'beak', 'body', 'wings', 'tail'], while eye rings and ear tufts reside inside head and prominent_features. Removing non-section keys (eye, legs) from ANATOMY_KEYS ensures queries like yellow eye ring are never penalized for missing a standalone sections.eye property, achieving 0.987 on Acorn Woodpecker and 0.985 on Brown Thrasher.
Int8Array vs Unsigned Buffer
Node.js fs.readFileSync() returns a Buffer (which is an unsigned Uint8Array reading 0..255). Binding both chunkVectorBuffer and chunkVectorInt8 to a signed new Int8Array(...) (-128..127) prevents negative quantized weights from wrapping around to 128..255 during SIMD dot-product scans.
8. Live Benchmark Verification Across Both Query Regimes
Click any row's Inspect Live β button below to load that query into the Playground and inspect its live Stage 1 MaxSim, Stage 2 Cross-Encoder logits, and evaluated 52-token passages:
| Query | Routing Mode | #1 Verified Result | #2 Verified Result | #3 Verified Result | Crushed False Positives | Action |
|---|---|---|---|---|---|---|
prominent ear tufts |
Winning-Chunk First | Long-eared Owl (0.936, CE 0.996) |
Great Horned Owl (0.935, CE 0.972) |
Dusky Eagle-Owl (0.919, CE 0.993) |
Red-vented Bulbul (0.55 β 0.00) |
|
yellow eye ring |
Winning-Chunk First | Acorn Woodpecker (0.987, CE 0.997) |
Brown Thrasher (0.985, CE 0.994) |
Brown Boobook (0.978, CE 0.994) |
β (No false positives) | |
standing on a rock |
Winning-Chunk First | Common Tern (0.905, CE 0.997) |
American Dipper (0.905, CE 0.996) |
Virginia Rail (0.904, CE 0.995) |
β (No false positives) | |
red wings and body |
Color-Anatomy Binding | Northern Cardinal (0.865, +13 β) |
Scarlet Tanager (0.860, +18 β) |
Red Avadavat (0.845, +42 β) |
Bufflehead (0.618 β 0.057) |
|
yellow body black wings |
Color-Anatomy Binding | Bullock's Oriole (0.959, CE 0.993) |
Black-hooded Oriole (0.953, CE 0.993) |
Lawrence's Goldfinch (0.947, CE 0.991) |
White-throated Kingfisher (0.800 β 0.209) |