Stable Diffusion Opens a New Era of Generative Images
Published:
Stability AI released Stable Diffusion 1.4 publicly on August 22, 2022, making the model weights downloadable under the CreativeML Open RAIL-M license (allowing commercial use with behavioral restrictions, unlike prior generative image models locked behind APIs). The model had been developed by CompVis (Ludwig Maximilian University of Munich, led by Robin Rombach) with funding from Stability AI and collaboration from Runway ML, based on the “Latent Diffusion Models” paper (Rombach et al., arXiv December 2021, CVPR 2022 Best Paper). The training dataset was LAION-5B (5.85 billion image-text pairs scraped from the public web), the largest openly published training dataset at the time, curated by the LAION non-profit organization. The SD 1.x models fit in approximately 4 GB (fp32 weights) or 2 GB (fp16), runnable on consumer NVIDIA GPUs with 6–8 GB VRAM or Apple Silicon M1/M2 (via CoreML or MPS backends, slower but functional without a discrete GPU) — a stark contrast to DALL-E 2 (OpenAI, April 2022, API-only at ~$0.02/image) and Midjourney (March 2022, Discord bot, subscription, closed source).
The architecture used a three-component pipeline implementing Latent Diffusion: a VAE (Variational Autoencoder) encoder compressed 512×512 RGB images (768KB uncompressed) to 64×64 latent representations (256KB, an 8× spatial reduction in each dimension), reducing the computational space by 64× compared to pixel-space diffusion. A CLIP text encoder (ViT-L/14, from OpenAI’s CLIP model) converted the user’s text prompt into a 77-token embedding sequence. A U-Net denoiser (the core generative model, ~860M parameters for SD 1.x) iteratively removed noise from the latent representation over 20–50 steps (configurable), conditioned on the text embedding at each step via cross-attention layers. The VAE decoder then expanded the denoised 64×64 latent back to 512×512 RGB. Each sampling step on a consumer RTX 3080 took approximately 40ms, producing images in 1–2 seconds total for 20-step PLMS/DPM++ sampling.
The open weights enabled a rapid community ecosystem that distinguished Stable Diffusion from API-gated competitors. AUTOMATIC1111 (stable-diffusion-webui, September 2022) provided a feature-rich web UI with img2img (modifying existing images guided by text), inpainting (replacing masked regions using text prompts), negative prompts, LoRA (Low-Rank Adaptation) model fine-tuning support, and ESRGAN upscaling. LoRA fine-tuning (arXiv January 2022 for LLMs, adapted for SD) allowed training a small adapter (10–200MB) that modified Stable Diffusion’s outputs using 20–100 labeled images, enabling users to capture specific art styles, characters, or subjects without training the full model. DreamBooth (Google Research, arXiv August 2022) enabled subject-specific personalization from 3–5 input images. ControlNet (Stanford/Stony Brook, arXiv February 2023) added spatial conditioning on depth maps, human pose skeletons, edge maps (Canny/HED), and segmentation masks, giving artists precise compositional control that text prompting alone couldn’t achieve. Getty Images filed a copyright lawsuit against Stability AI in January 2023, citing the inclusion of Getty’s watermarked images in LAION-5B’s training data; class-action suits from artists (including Sarah Andersen, Kelly McKernan, and Karla Ortiz) followed in January 2023 and March 2023, initiating the legal debate over whether training on copyrighted images without license or compensation constitutes copyright infringement — a question unresolved at the time of Stable Diffusion XL (July 2023) and SD 3.0 (February 2024).
