Peking University & Visual Intelligence Group · 2026

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model (PixelJev)

Authors: Xunlan Zhou, Xianliang Yang, Li Zhao

$\log P(c_k \mid I, Q) = \frac{1}{L_k^\alpha} \sum_{l=1}^{L_k} \log P(w_{k,l} \mid I, Q, w_{k,<l})$
Core mechanics of PixelJev: 1. **Direct Pixel Ingestion**: Bypasses text captioning steps, preserving fine-grained spatial and texture features directly in the latent representation. 2. **Dynamic Candidate Semantics**: Supports arbitrary, developer-defined candidate option sets at runtime without retraining or fine-tuning linear projection layers. 3. **Multi-Token Candidate Scoring**: Uses teacher-forced log-likelihood accumulation with length penalty normalization (alpha = 0.6) to prevent vocabulary bias on multi-word options.

Frequently Asked Questions

What is the key advantage of PixelJev over a fixed classification head?

A fixed classification head only recognizes the specific classes seen during training. PixelJev accepts arbitrary dynamic candidate options provided at runtime by software code.

Can PixelJev be used for real-time generative diffusion filtering?

Yes. PixelJev evaluates generated image artifacts, anatomy compliance, and prompt fidelity in under 15 milliseconds, enabling automated in-loop rejection sampling.