LiFT: Loop Flow Transformers Cut Image Generation Costs While Improving Quality
Image via arxiv.org
Researchers from the University of Amsterdam and TNO have introduced Loop Flow Transformers (LiFT), a new family of generative model architectures that replaces stacking many distinct transformer layers with repeatedly applying a single shared Diffusion Transformer (DiT) core. Each loop step is trained against a regression target on a straight path toward the flow-matching goal, indexed by a continuous depth coordinate — allowing the model to run more loops at inference time than it was trained on, without retraining or architectural changes.
On the ImageNet 256×256 benchmark, LiFT-L/2 achieves an FID score 3.34 points lower than a dense DiT-XL/2 baseline while using roughly 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs. The ability to extend rollout depth at inference without adding parameters offers a practical compute-quality trade-off that could be relevant to production pipelines relying on diffusion-based image and video generation.