Research

LiFT: Loop Flow Transformers Cut Image Generation Costs While Improving Quality

about 24 hours ago
arXiv logo

Image via arxiv.org

Researchers from the University of Amsterdam and TNO have introduced Loop Flow Transformers (LiFT), a new family of generative model architectures that replaces stacking many distinct transformer layers with repeatedly applying a single shared Diffusion Transformer (DiT) core. Each loop step is trained against a regression target on a straight path toward the flow-matching goal, indexed by a continuous depth coordinate — allowing the model to run more loops at inference time than it was trained on, without retraining or architectural changes.

On the ImageNet 256×256 benchmark, LiFT-L/2 achieves an FID score 3.34 points lower than a dense DiT-XL/2 baseline while using roughly 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs. The ability to extend rollout depth at inference without adding parameters offers a practical compute-quality trade-off that could be relevant to production pipelines relying on diffusion-based image and video generation.