Non-Autoregressive Neural Sequence Generation: Foundational Research Behind Parallel Sampling
Authors: Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, Richard Socher
$p_{\text{NAR}}(Y|X; \theta) = p_L(T|X; \theta) \prod_{t=1}^T p(y_t|x_1, \dots, x_{T^\prime}, \theta)$
The non-autoregressive transformer architecture fundamentally alters decoding dynamics:
1. **Elimination of Causal Masking**: The decoder uses bidirectional self-attention across all target positions, allowing all output tokens to exchange information simultaneously.
2. **Fertility & Length Prediction**: Predicts sequence length $T$ and source token fertilities prior to decoding, establishing the coordinate grid for parallel token emission.
3. **Multi-Modality Resolution via Iterative Refinement**: Resolves candidate conflicts across positions through iterative parallel refinement passes rather than autoregressive left-to-right backtracking.