Guide
What Is an AI Image Prompt? A Complete Guide
A prompt is not a magic phrase. It is a description with a structure, and the structure is what makes it repeatable. This guide explains what a prompt is made of, why each part exists, and how to tell a vague prompt from a specific one.
What an AI image prompt is
An AI image prompt is a written description that an image model turns into a picture. The model has learned associations between language and visual features, so the words you choose act less like commands and more like weights: each phrase pulls the result toward a region of everything the model has seen. "A dog" pulls toward the average of countless dog pictures. "A wet border collie mid-stride on a wooden pier at dawn, low sun behind it" pulls toward a much smaller, much more specific neighbourhood.
That is the whole game. You are not writing instructions to a renderer that will follow them literally. You are narrowing a very large space of possible images until only the ones you want are left.
Why prompts matter
Two people can use the same model on the same day and get wildly different quality. The difference is almost never the model — it is whether the prompt resolves ambiguity. Every detail you leave out, the model fills in with whatever is statistically typical, and "typical" is exactly what makes AI images look generic: centred subject, soft even light, shallow depth of field, mid-blue sky.
A good prompt is also an asset. Once you have one that works, you can swap the subject and keep the look, which is how people produce a coherent set of images instead of a pile of unrelated ones. That reusability is the main reason it is worth writing prompts deliberately rather than typing a sentence and rerolling.
The components of a prompt
Subject
What is in the picture, and what state it is in. Name the thing, then qualify it: age, material, condition, action, clothing, expression. "An old man" and "a weathered fisherman in his seventies, squinting, three days of stubble, oilskin jacket" produce very different images. Keep one clear primary subject unless you deliberately want a crowd.
Composition
Where the subject sits in frame and how much of it we see. Useful words: close-up, medium shot, wide establishing shot, overhead, low angle, eye level, three-quarter view, centred, off-centre, framed by foreground foliage, negative space on the left. Composition is the single most under-specified part of most prompts and the fastest way to stop images looking default.
Lighting
Direction, quality and colour of light. Direction: front, side, backlit, top-down, underlit. Quality: hard, soft, diffuse, dappled, specular. Source and colour: golden hour sun, overcast daylight, single tungsten bulb, neon spill, candlelight, studio softbox with a rim light. Lighting carries mood more reliably than any adjective about mood.
Camera language
Focal length, aperture and distance change perspective in ways models genuinely respond to. A 24mm close shot exaggerates and bends; an 85mm at f/1.8 compresses and separates; a 200mm flattens layers into each other. Add film or sensor character only if you want it — "35mm film grain", "large format detail", "handheld, slight motion blur".
Style
The visual language: documentary photography, studio product shot, ink and wash, risograph print, cel animation, technical illustration, matte painting. Pick one. Two competing styles in one prompt usually produce a muddy compromise rather than a fusion.
Colour
State a palette rather than a vibe. "Restrained palette: warm sand, off-white, one deep oxblood accent" beats "nice colours". You can also describe relationships — analogous and calm, complementary and tense, near-monochrome with one saturated element.
Detail and quality
Say what should be sharp and what should not. "Crisp texture on the fabric, background falling away softly" is a useful instruction. Stacked quality words — masterpiece, 8K, ultra detailed, award winning — mostly add noise on modern models and can flatten your style into stock-photo gloss.
Examples: vague versus specific
Vague
A beautiful coffee shop, high quality, 8K, amazing lighting.
Specific
Create an image based on the following guidelines: a narrow neighbourhood coffee bar photographed from just inside the doorway, medium-wide shot at eye level, one barista in a grey apron pulling a shot at the far end of the counter. Late-morning daylight enters from the street behind camera and rakes across the zinc counter, leaving the back wall in soft shadow. Warm neutral palette — oak, cream tile, matte black fittings — with a single red enamel sign as the only saturated element. 35mm at f/4, everything from counter to barista readable, gentle film grain, no lens flare.
The second prompt is not longer for the sake of it. Every added clause removes a decision the model would otherwise make at random, which is why it will produce a similar image tomorrow and the vague one will not.
Common mistakes
Contradicting yourself. "Wide shot, extreme close-up of the eyes", or "soft dreamy light, harsh dramatic shadows". The model averages the conflict and you get mush.
Describing what you don't want in the prompt. "No text" often produces text, because you have mentioned text. Put exclusions in a negative prompt if your tool supports one, or simply describe the positive alternative.
Piling on adjectives. After a certain length, extra words dilute the ones that mattered. Cut anything that does not change the picture.
Naming a living artist to copy their style. Beyond the ethics, it is a shortcut that hides what you actually like. Describe the qualities instead — you will be able to reuse them.
Rerolling instead of editing. If the result is 70% right, change one clause and run again. Rerolling the same prompt teaches you nothing.
Where to go next
Ready to write your own? Work through how to write better AI image prompts, then read annotated examples. If you would rather start from a picture you already like, upload it to the image to prompt tool and edit what comes back.
