How AI Generates Images
Diffusion isn't the only way AI makes images. Here's how the major approaches compare, and what's actually running behind a typical image generator.
Major image-generation approaches
| Approach | How it works | Known for |
|---|---|---|
| GANs | Two networks compete: a Generator tries to fool a Discriminator | Early photorealistic faces (e.g. StyleGAN) |
| Diffusion Models | Iteratively denoise random noise, guided by a prompt | Stable Diffusion, Midjourney, DALL·E 3 |
| Autoregressive / Transformer | Generates an image token-by-token, like text generation | Early DALL·E, some newer multimodal models |
Text encoder
Turns your written prompt into a signal the image model can follow.
Generation model
The GAN, diffusion, or transformer core that actually produces pixels.
Upscaler
Increases resolution and sharpens details after the base image is generated.
Safety filter
Screens prompts and outputs against content policies before returning results.
Key takeaways
- GANs, diffusion models, and autoregressive transformers are the three major families of image generators.
- Diffusion models currently dominate — powering Stable Diffusion, Midjourney, and DALL·E 3.
- A production image generator is a pipeline: text encoder → generation model → upscaler → safety filter, not a single step.
Check your understanding
0/2 answered1.Which image-generation approach uses two competing networks — a Generator and a Discriminator?
2.Most modern text-to-image products like Midjourney and DALL·E 3 are built on diffusion models.
Lesson summary
AI image generation spans GANs, diffusion models, and autoregressive transformers — with diffusion currently leading — wrapped in a pipeline of text encoding, generation, upscaling, and safety filtering.
AI-generated notes