Decoded-output supervision
Discriminate generated and real samples after RGB decoding, making blur, texture defects, and VAE artifacts directly observable.
Radian complements on-policy distribution matching with real-data supervision in a frozen visual representation space, without changing inference cost.
Few-step autoregressive video generation enables efficient streaming synthesis, but early errors are reused as context and can accumulate into detail loss, structural drift, and unstable motion. Existing distribution matching distillation primarily aligns student and teacher distributions in diffusion latent space, without directly supervising the perceptual quality of decoded videos.
Radian sparsely decodes frames from on-policy student rollouts, maps them into a frozen visual foundation model, and trains lightweight discriminator heads against real videos. DMD anchors the student to the teacher distribution, while representation-space adversarial supervision supplies complementary perceptual and semantic signals. The feature encoder and discriminator are removed after training, so the generator architecture and sampling budget remain unchanged.
Discriminate generated and real samples after RGB decoding, making blur, texture defects, and VAE artifacts directly observable.
Keep DMD as the distribution anchor, introduce perceptual directions during training, then discard the feature encoder and discriminator at inference.
Systematically compare image, video, joint image–video, and diffusion-internal representations under a controlled adversarial distillation setup.
Radian begins from a causal ODE initialization, calibrates DMD and the discriminator, then jointly optimizes the student with on-policy DMD and a scheduled representation-space adversarial objective.
ℒDMD preserves the teacher-derived generative prior, while the scheduled ℒΦadv term reshapes outputs toward perceptually realistic and semantically coherent real-data modes.
For each sparsely decoded frame, the frozen encoder supplies local tokens augmented with its global readout. A small trainable head then scores every spatial location; logits are averaged over time, levels, and tokens.
Training only. The VAE decoder, frozen VFM, and discriminator heads are removed after training, adding no inference-time parameters or NFE.
All tables below are redrawn from the paper's unified local evaluation. Darker purple denotes stronger column-wise performance.
| Method | NFE | VBench ↑ | VideoAlign ↑ | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Total | Quality | Semantic | Subj. Cons. | Dynamic | Aesthetic | Imaging | Object | Color | VQ | MQ | TA | Total | ||
| Self Forcing | 4 | 0.8393 | 0.8471 | 0.8082 | 0.9378 | 0.6972 | 0.6745 | 0.6995 | 0.9440 | 0.8680 | 0.0695 | 0.0616 | 0.3683 | 0.4995 |
| Causal Forcing | 4 | 0.8400 | 0.8495 | 0.8019 | 0.9310 | 0.8694 | 0.6766 | 0.6993 | 0.9544 | 0.8311 | 0.0072 | -0.0951 | 0.3542 | 0.2663 |
| Salt | 4 | 0.8434 | 0.8548 | 0.7982 | 0.9383 | 0.7926 | 0.6797 | 0.7034 | 0.9494 | 0.8291 | 0.0472 | -0.0310 | 0.3774 | 0.3936 |
| DiT-GAN | 4 | 0.8411 | 0.8483 | 0.8122 | 0.9492 | 0.5954 | 0.6832 | 0.6984 | 0.9563 | 0.8713 | -0.0242 | 0.0855 | 0.3715 | 0.4327 |
| Radian (CD init) | 4 | 0.8439 | 0.8550 | 0.7993 | 0.9380 | 0.7972 | 0.6792 | 0.7037 | 0.9508 | 0.8300 | 0.0521 | -0.0296 | 0.3754 | 0.3979 |
| Radian | 4 | 0.8444 | 0.8541 | 0.8054 | 0.9504 | 0.8056 | 0.6801 | 0.7148 | 0.9570 | 0.8722 | 0.2134 | 0.2094 | 0.3805 | 0.8033 |
| Method | NFE | VBench ↑ | VideoAlign ↑ | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Total | Quality | Semantic | Object | App. Style | Temp. Style | Overall Cons. | TA | Total | ||
| Causal Forcing++ | 1† | 0.8383 | 0.8473 | 0.8025 | 0.9473 | 0.7248 | 0.6908 | 0.7255 | 0.3115 | -0.1813 |
| One Forcing | 1† | 0.8415 | 0.8533 | 0.7943 | 0.9568 | 0.7433 | 0.6907 | 0.7178 | 0.3168 | 0.0448 |
| Radian | 1† | 0.8417 | 0.8511 | 0.8040 | 0.9663 | 0.7546 | 0.7034 | 0.7267 | 0.4105 | 0.2563 |
| Method | NFE | Total ↑ | Quality ↑ | Semantic ↑ | Subject Cons. ↑ | Temp. Flicker ↑ | Motion Smooth. ↑ | Dynamic ↑ | Imaging ↑ |
|---|---|---|---|---|---|---|---|---|---|
| Rolling Forcing | 5 | 0.7805 | 0.8191 | 0.6260 | 0.9753 | 0.9871 | 0.9842 | 0.4532 | 0.7075 |
| Radian | 4 | 0.8041 | 0.8511 | 0.6160 | 0.9790 | 0.9882 | 0.9864 | 0.6741 | 0.7194 |
Controlled comparisons under the same 4-NFE chunk-wise training and inference setup. Darker purple denotes stronger column-wise performance.
| Feature source | Feature model | Total ↑ | Quality ↑ | Semantic ↑ | Multi-Obj. ↑ | Color ↑ | Consist. ↑ | Dynamic ↑ | Imaging ↑ | VQ ↑ | MQ ↑ | TA ↑ | VA Total ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Image VFM | DINOv2 | 0.8444 | 0.8541 | 0.8054 | 0.8343 | 0.8722 | 0.2662 | 0.8056 | 0.7148 | 0.2134 | 0.2094 | 0.3805 | 0.8033 |
| Image VFM | DINOv3 | 0.8401 | 0.8498 | 0.8012 | 0.8528 | 0.8729 | 0.2649 | 0.7685 | 0.7104 | 0.1462 | 0.2937 | 0.3607 | 0.8006 |
| Image VFM | SigLIP2 | 0.8391 | 0.8475 | 0.8052 | 0.8384 | 0.8662 | 0.2680 | 0.6343 | 0.6992 | 0.1156 | 0.1320 | 0.3751 | 0.6226 |
| Video VFM | V-JEPA 2.1 | 0.8360 | 0.8430 | 0.8078 | 0.8604 | 0.8627 | 0.2676 | 0.5991 | 0.7042 | 0.0498 | 0.1275 | 0.4119 | 0.5892 |
| Video VFM | VideoMAE | 0.8372 | 0.8436 | 0.8115 | 0.8417 | 0.8787 | 0.2676 | 0.5213 | 0.7026 | 0.1000 | 0.1452 | 0.3436 | 0.5888 |
| Image + Video VFM | DINOv2 + VideoMAE | 0.8434 | 0.8525 | 0.8070 | 0.8652 | 0.8545 | 0.2679 | 0.6593 | 0.7074 | 0.0831 | 0.0575 | 0.3079 | 0.4485 |
| Generative | DiT-GAN | 0.8411 | 0.8483 | 0.8122 | 0.8432 | 0.8713 | 0.2674 | 0.5954 | 0.6984 | -0.0242 | 0.0855 | 0.3715 | 0.4327 |
| Size | VBench Total ↑ | Quality ↑ | Semantic ↑ | VA Total ↑ |
|---|---|---|---|---|
| S | 0.8444 | 0.8541 | 0.8054 | 0.8033 |
| B | 0.8322 | 0.8404 | 0.7993 | 0.8910 |
| L | 0.8301 | 0.8366 | 0.8042 | 0.3686 |
| G | 0.8387 | 0.8475 | 0.8035 | 0.3257 |
| # Layers | VBench Total ↑ | Quality ↑ | Semantic ↑ | VA Total ↑ |
|---|---|---|---|---|
| 1 | 0.8182 | 0.8206 | 0.8085 | 0.6747 |
| 2 | 0.8283 | 0.8389 | 0.7859 | 0.6586 |
| 4 | 0.8444 | 0.8541 | 0.8054 | 0.8033 |
| 8 | 0.8381 | 0.8469 | 0.8026 | 0.7228 |
Choose a generation regime and case, then use the shared controller to start every method on the same frame. Videos are muted for consistent playback.
Additional prompts across natural scenes, stylized actions, and difficult human–object interactions.
Drag or shift-scrollVideos play when centered
@misc{radian,
title = {Enhancing Autoregressive Video Generation via Representation Adversarial Distillation},
author = {Lin, Fangyu and Ge, Xingtong and Zhu, Lunjie and Zhang, Yi and Liu, Zhening and Wang, Tianhang and Li, Mengfei and Zhang, Yumeng and Song, Guanglu and Liu, Yu and Zhang, Jun},
year = {2026},
note = {Manuscript}
}