Level-of-Token Diffusion Brings Adaptive Compute Allocation to Image and Video Generation
Researchers have introduced Level-of-Token (LoT) Diffusion, a framework that replaces the uniform token grids used by standard diffusion transformers with adaptive multiresolution layouts. Instead of allocating equal computation to every image region, LoT assigns finer tokens to detail-rich areas — guided by semantic masks, bounding boxes, texture variance, or depth-of-field cues — and coarser tokens to plain backgrounds, reducing total sequence length without sacrificing full-resolution flow prediction at each denoising step.
The framework fine-tunes pretrained diffusion transformers via patch-wise asymmetric flow matching, preserving existing generative priors while enabling layout-adaptive generation for both images and video. The team demonstrates significant inference speedups that scale with the token budget, and extends the approach to agentic planning scenarios where layout decisions are made programmatically. Authors include researchers from Stanford, Google, and ETH Zurich affiliations.