We scan new podcasts and send you the top 5 insights daily.
Training a diffusion model involves repeatedly adding and removing noise from data points. This process effectively creates many augmented views of the same data, making the models potentially more data-efficient—a key advantage if data becomes a future bottleneck.
The SNR-T bias can be fixed efficiently without retraining models. At each denoising step, the image is broken into frequency bands using wavelets. Each band is then given a small correction based on its specific noise mismatch before being recombined. This surgical approach is computationally cheap and universally effective.
Diffusion models were a breakthrough for protein generation because they reframe the problem. Instead of a one-shot generation, they learn to make many small, iterative refinements ("make it slightly better"). This "time to think" approach proved more effective for complex biological structures than previous methods like VAEs.
Diffusion models work on a continuous medium like an image by adding noise until it's unrecognizable, then training a model to reverse the process. This holistic, denoising method is fundamentally different from autoregressive models like large language models, which predict data one token at a time.
While hard-coding physical symmetries (equivariance) into a model is theoretically efficient, it can fail in practice. Prof. Welling explains that these constraints can complicate the optimization landscape, making it harder to find good minima. Sometimes, abundant data augmentation with a simpler model yields superior results.
Previously, imitation learning required a single expert to collect perfectly consistent data, a major bottleneck. Diffusion models unlocked the ability to train on multi-modal data from various non-expert collectors, shifting the challenge from finding niche experts to building scalable data acquisition and processing systems.
During training, diffusion models learn a perfect relationship between noise level (SNR) and denoising step (T). During inference, this relationship breaks as the model's own predictions introduce errors, creating SNR values it never trained on for a given step. This causes compounding errors and quality loss.
Diffusion models naturally reconstruct images in layers. In early denoising stages with high noise, they focus on low-frequency information like overall composition and color. As noise decreases in later steps, they add high-frequency details like textures and sharp edges. This hierarchical process is key to understanding their behavior.
Autoregressive models like GPT are sequential at inference (one token at a time), creating a GPU bottleneck. Diffusion models process many tokens in parallel during inference, similar to how transformers parallelized training, leading to fundamental speed advantages.
Models like Stable Diffusion achieve massive compression ratios (e.g., 50,000-to-1) because they aren't just storing data; they are learning the underlying principles and concepts. The resulting model is a compact 'filter' of intelligence that can generate novel outputs based on these learned principles.
The quality of generative visuals has leaped from blurry blobs to near-photorealistic films in a few years. Yet, the core technology—a diffusion process of adding and then removing noise—has remained consistent. Progress stems from optimizations and architectural improvements, not a complete paradigm shift.