Stable Diffusion
Text-to-image models like Stable Diffusion don't 'draw' — they start with pure noise and gradually refine it into an image, guided by your prompt.
From noise to image
- 1
Start with random noise
The process begins with a canvas of pure statistical noise — no image at all.
- 2
Encode the text prompt
A text encoder (like CLIP) converts your prompt into a numeric representation of its meaning.
- 3
Iteratively denoise
Over dozens of steps, the model removes a little noise at a time, each step nudged toward matching the prompt.
- 4
Decode the final image
The refined representation is decoded from latent space into the final pixel image.
Why 'diffusion'?
The name comes from physics: just as ink diffuses into water and becomes disordered, these models are trained by watching images get progressively noisier — then they learn to run that process in reverse, turning noise back into a coherent image.
Latent space
A compressed numeric representation of images, far smaller than raw pixels — this is what actually gets denoised.
Text encoder
Converts your prompt into a guidance signal the denoising process follows at every step.
Sampling steps
More denoising steps generally mean higher quality, at the cost of more compute time.
Key takeaways
- Diffusion models generate images by starting with random noise and iteratively removing it, guided by a text prompt.
- A text encoder converts your prompt into a signal that steers every denoising step.
- Most of the actual computation happens in a compressed 'latent space', not on raw pixels.
Check your understanding
0/2 answered1.What does a diffusion model start with before generating an image?
2.Diffusion models generate the final image in a single step.
Lesson summary
Stable Diffusion generates images by iteratively denoising random noise in latent space, guided at every step by a text encoding of your prompt.
AI-generated notes