From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model (PixelJev)
Authors: Xunlan Zhou, Xianliang Yang, Li Zhao
$\log P(c_k \mid I, Q) = \frac{1}{L_k^\alpha} \sum_{l=1}^{L_k} \log P(w_{k,l} \mid I, Q, w_{k,<l})$
Core mechanics of PixelJev:
1. **Direct Pixel Ingestion**: Bypasses text captioning steps, preserving fine-grained spatial and texture features directly in the latent representation.
2. **Dynamic Candidate Semantics**: Supports arbitrary, developer-defined candidate option sets at runtime without retraining or fine-tuning linear projection layers.
3. **Multi-Token Candidate Scoring**: Uses teacher-forced log-likelihood accumulation with length penalty normalization (alpha = 0.6) to prevent vocabulary bias on multi-word options.
Frequently Asked Questions
What is the key advantage of PixelJev over a fixed classification head?
A fixed classification head only recognizes the specific classes seen during training. PixelJev accepts arbitrary dynamic candidate options provided at runtime by software code.
Can PixelJev be used for real-time generative diffusion filtering?
Yes. PixelJev evaluates generated image artifacts, anatomy compliance, and prompt fidelity in under 15 milliseconds, enabling automated in-loop rejection sampling.