AI and the Future

Lesson 4 of 7

AI Images

Uncover how diffusion models turn text prompts into stunning visual art by learning to reverse random noise.

🟡 Intermediate 80 XP

Learn it

Have you ever watched an artist start with a rough canvas and slowly refine it into a masterpiece? AI image generators like Midjourney or Stable Diffusion do something similar, but with a wild twist: they start with static noise, like a fuzzy broken TV screen!

To make an image, the AI follows your text prompt and gradually cleans up the digital static, turning random dots into shapes, colours, and textures until a crisp, original picture appears.

Key terms

Diffusion Model
A generative neural network that produces images by iteratively removing noise from a random starting canvas.
Latent Space
A mathematical multi-dimensional space where concepts, features, and embeddings are represented numerically.
CLIP
A neural network trained to associate natural language descriptions with visual concepts in a shared space.
Denoising
The iterative computational process of removing random static to reveal structured visual elements.

How Text Becomes an Image

Follow the step-by-step journey of an image generation pipeline from text prompt to final render.

  1. 1Prompt Interpretation: A text encoder converts your descriptive words into numerical embeddings representing concepts, styles, and lighting.
  2. 2Initialising Noise Canvas: The system generates a grid of completely random Gaussian noise in a compressed latent space.
  3. 3Iterative Denoising Steps: A U-Net or transformer neural network progressively predicts and subtracts noise over 20 to 50 iterations.
  4. 4Conditioning Guidance: At each step, the text embeddings steer the denoising trajectory towards matching your requested subject and art style.
  5. 5High-Resolution Decoding: A variational autoencoder (VAE) decodes the finished latent representation into standard pixel formats like PNG or JPEG.

Conceptual Denoising Loop

python# Conceptual illustration of a diffusion denoising loop
steps = 5
current_canvas = 'Pure Gaussian Noise'

print(f'Step 0: {current_canvas}')
for step in range(1, steps + 1):
    noise_level = 100 - (step * 20)
    print(f'Step {step}: Removing noise... {noise_level}% static remaining (adding prompt features)')

print('Final Step: Crisp render generated from prompt!')

Diffusion models execute repeated denoising steps, gradually reducing entropy and introducing structured visual features steered by the text prompt.

Try it

Evaluate these statements about AI image generation.

Diffusion models start with pure random noise and iteratively refine it into an image.

AI image generators copy and paste whole photographs directly from the internet.

CLIP helps the AI connect text descriptions with visual concepts in a shared mathematical space.

Once an AI is trained on images, it creates physical paintings using robotic arms.

Challenge

Formulate an advanced prompt for an AI image generator to create a futuristic eco-city. Specify lighting, perspective, architectural style, atmosphere, and render engine details.

Pick whichever way suits you — every mode earns the same bonus XP.

Write at least 40 more characters to submit.

Mark your own work

Guided walkthrough — 0/5 clues revealed

  1. Clue 1 locked — reveal it only if you get stuck.
  2. Clue 2 locked — reveal it only if you get stuck.
  3. Clue 3 locked — reveal it only if you get stuck.
  4. Clue 4 locked — reveal it only if you get stuck.
  5. Clue 5 locked — reveal it only if you get stuck.

Each clue costs 4 XP (never below 20 XP). You'd earn 40 XP right now.

Extension: Explain how the AI model uses the CLIP encoder to distinguish between an eco-city and a cyberpunk dystopian city.

Quiz time

Question 1 of 4Score 0

What starting input do diffusion models use to generate a new picture?