Few-step autoregressive video generation

Enhancing Autoregressive Video Generation
via Representation Adversarial Distillation

Radian complements on-policy distribution matching with real-data supervision in a frozen visual representation space, without changing inference cost.

Fangyu Lin1,2 Xingtong Ge1,2 Lunjie Zhu1,2 Yi Zhang2,‡ Zhening Liu1 Tianhang Wang3 Mengfei Li1,2 Yumeng Zhang1 Guanglu Song2 Yu Liu2 Jun Zhang1,†

1 The Hong Kong University of Science and Technology 2 Vivix Group Limited 3 Zhejiang University

‡ Project Lead   ·   † Corresponding Author

4 NFE · Bicycle
1† NFE · Panda dining
4NFE · Radian
Intro

Real-data supervision, where visual quality is represented.

Few-step autoregressive video generation enables efficient streaming synthesis, but early errors are reused as context and can accumulate into detail loss, structural drift, and unstable motion. Existing distribution matching distillation primarily aligns student and teacher distributions in diffusion latent space, without directly supervising the perceptual quality of decoded videos.

Radian sparsely decodes frames from on-policy student rollouts, maps them into a frozen visual foundation model, and trains lightweight discriminator heads against real videos. DMD anchors the student to the teacher distribution, while representation-space adversarial supervision supplies complementary perceptual and semantic signals. The feature encoder and discriminator are removed after training, so the generator architecture and sampling budget remain unchanged.

01

Decoded-output supervision

Discriminate generated and real samples after RGB decoding, making blur, texture defects, and VAE artifacts directly observable.

02

Complementary signal with no inference overhead

Keep DMD as the distribution anchor, introduce perceptual directions during training, then discard the feature encoder and discriminator at inference.

03

Broad feature-space exploration

Systematically compare image, video, joint image–video, and diffusion-internal representations under a controlled adversarial distillation setup.

Method
Complementary supervision. DMD anchors the student to the teacher distribution; Radian contributes a real-data perceptual signal that corrects degraded modes in decoded output space.

A three-stage recipe for representation adversarial distillation.

Radian begins from a causal ODE initialization, calibrates DMD and the discriminator, then jointly optimizes the student with on-policy DMD and a scheduled representation-space adversarial objective.

Joint objective
ℒRadian(t) = λDMD ℒDMD + gt λadv ℒΦadv

ℒDMD preserves the teacher-derived generative prior, while the scheduled ℒΦadv term reshapes outputs toward perceptually realistic and semantically coherent real-data modes.

Training pipeline. Sparse decoded frames are processed by a frozen image or video foundation model. Lightweight heads distinguish generated features from real-data features; only the causal generator is retained at inference.
Lightweight discriminator. Independent heads convert multi-level frozen features into dense spatial realism logits.
Multi-level feature discrimination

Dense evidence at every selected feature level.

For each sparsely decoded frame, the frozen encoder supplies local tokens augmented with its global readout. A small trainable head then scores every spatial location; logits are averaged over time, levels, and tokens.

Frozen representation
F̃ℓ(ym) = Φℓ(𝒫(ym)) + 1 r⊤ℓ
Dense discriminator logits
dm,ℓ(ym) = hψ,ℓ(F̃ℓ(ym)) ∈ ℝNℓ
Real–generated hinge objective
ℒD = 𝔼x[(1 − d(x))+] + 𝔼x̂[(1 + d(x̂))+]

Training only. The VAE decoder, frozen VFM, and discriminator heads are removed after training, adding no inference-time parameters or NFE.

Quantitative Results

One method, three autoregressive regimes.

All tables below are redrawn from the paper's unified local evaluation. Darker purple denotes stronger column-wise performance.

Four-step chunk-wise generation15 shared seeds · 944 prompts per seed
4 NFE
MethodNFEVBench ↑VideoAlign ↑
TotalQualitySemanticSubj. Cons.DynamicAestheticImagingObjectColorVQMQTATotal
Self Forcing40.83930.84710.80820.93780.69720.67450.69950.94400.86800.06950.06160.36830.4995
Causal Forcing40.84000.84950.80190.93100.86940.67660.69930.95440.83110.0072-0.09510.35420.2663
Salt40.84340.85480.79820.93830.79260.67970.70340.94940.82910.0472-0.03100.37740.3936
DiT-GAN40.84110.84830.81220.94920.59540.68320.69840.95630.8713-0.02420.08550.37150.4327
Radian (CD init)40.84390.85500.79930.93800.79720.67920.70370.95080.83000.0521-0.02960.37540.3979
Radian40.84440.85410.80540.95040.80560.68010.71480.95700.87220.21340.20940.38050.8033
Ablation Study

Representation space, encoder scale, and feature depth.

Controlled comparisons under the same 4-NFE chunk-wise training and inference setup. Darker purple denotes stronger column-wise performance.

Adversarial feature spaces15 shared seeds · identical generator and inference procedure
Table 2
Feature sourceFeature modelTotal ↑Quality ↑Semantic ↑Multi-Obj. ↑Color ↑Consist. ↑Dynamic ↑Imaging ↑VQ ↑MQ ↑TA ↑VA Total ↑
Image VFMDINOv20.84440.85410.80540.83430.87220.26620.80560.71480.21340.20940.38050.8033
Image VFMDINOv30.84010.84980.80120.85280.87290.26490.76850.71040.14620.29370.36070.8006
Image VFMSigLIP20.83910.84750.80520.83840.86620.26800.63430.69920.11560.13200.37510.6226
Video VFMV-JEPA 2.10.83600.84300.80780.86040.86270.26760.59910.70420.04980.12750.41190.5892
Video VFMVideoMAE0.83720.84360.81150.84170.87870.26760.52130.70260.10000.14520.34360.5888
Image + Video VFMDINOv2 + VideoMAE0.84340.85250.80700.86520.85450.26790.65930.70740.08310.05750.30790.4485
GenerativeDiT-GAN0.84110.84830.81220.84320.87130.26740.59540.6984-0.02420.08550.37150.4327
Qualitative Results

Watch every method on the same prompt.

Choose a generation regime and case, then use the shared controller to start every method on the same frame. Videos are muted for consistent playback.

More Visualization

Radian beyond the paper selection.

Additional prompts across natural scenes, stylized actions, and difficult human–object interactions.

Drag or shift-scrollVideos play when centered

Citation

Radian

@misc{radian,
  title  = {Enhancing Autoregressive Video Generation via Representation Adversarial Distillation},
  author = {Lin, Fangyu and Ge, Xingtong and Zhu, Lunjie and Zhang, Yi and Liu, Zhening and Wang, Tianhang and Li, Mengfei and Zhang, Yumeng and Song, Guanglu and Liu, Yu and Zhang, Jun},
  year   = {2026},
  note   = {Manuscript}
}
Enlarged paper figure