Google Reveals Imagen
Published:
Google Research published the Imagen paper on May 23, 2022, alongside a website showing examples — but without releasing the model or a public API. Imagen combined a frozen 4.6-billion-parameter T5-XXL text encoder (the same language model used for NLP tasks) with a pipeline of three cascaded diffusion models: a base model generating 64×64 images from text, then two super-resolution diffusion models upscaling to 256×256 and 1024×1024. The T5-XXL encoder had been trained on text only, not image-text pairs, and Google’s finding was that using a large language model’s text representations as conditioning signals produced better text-image alignment than training a custom CLIP-style encoder specifically for image generation — suggesting that language understanding, not vision-specific training, drove Imagen’s prompt adherence. Read more
