Prompting well is less about magic words and more about being specific in the dimensions the model actually responds to.
A Structure That Works
Most tools respond well to a prompt organized in this order:
- Subject — what it is, with specific attributes. "A ceramic coffee cup" beats "a cup".
- Action or state — what it's doing, or how it's arranged.
- Setting — where it is, and what's around it.
- Composition and framing — close-up, wide shot, overhead, portrait orientation.
- Lighting — soft window light, golden hour, hard studio flash, backlit.
- Style and medium — photograph, watercolour, 3D render, line drawing.
- Technical modifiers — shallow depth of field, 85 mm lens, high contrast.
Not every prompt needs all seven, but if a result is disappointing, the missing element is usually one of the middle three — composition, lighting, or setting.
Use Photographic Vocabulary
Models are trained on captioned images, and photographic terms appear in those captions with consistent visual meaning. So these work well:
- Framing: close-up, medium shot, wide shot, overhead / top-down, eye-level, low angle.
- Optics: 35 mm, 85 mm, macro, shallow depth of field, bokeh, tilt-shift.
- Light: golden hour, blue hour, softbox, rim light, backlit, high key, low key, chiaroscuro.
- Film and rendering: black and white, film grain, long exposure, studio product photography.
Vague quality words — "beautiful", "amazing", "high quality", "8K" — do far less than a concrete description of light and framing.
Common Mistakes
- Too many subjects. Models struggle to compose several distinct objects with precise relationships. Simplify, or generate elements separately and compose them yourself.
- Contradictions. "Minimalist and highly detailed", "wide shot close-up". The model averages the conflict into mush.
- Describing what you don't want in the main prompt. Saying "no text" can make text more likely; use the tool's negative prompt field if it has one.
- Specifying exact text. Most models still render text unreliably. Generate the image without text and add real type afterwards in an editor — which also lets you translate and edit it later.
- Expecting precise counts. "Exactly five apples" is unreliable in most models.
- Rewriting from scratch after each attempt. Change one variable at a time so you learn what actually caused the difference.
Practical Settings
- Aspect ratio is usually a parameter rather than prompt text. Set it to match the final use — 16:9 for a banner, 1:1 for a social tile, 4:5 for a portrait post.
- Seed values let you reproduce or vary a result deliberately; keep the seed and change one word to isolate its effect.
- Reference images (image-to-image, style reference) control composition and style far more reliably than words. If you know the layout you want, sketch it and use it as input.
After Generating
Generated images are a starting point, not a deliverable:
- Upscale carefully and check faces and text at 100%.
- Fix hands, edges and artifacts with inpainting.
- Add real text in a design tool rather than accepting generated lettering.
- Colour-correct to match your brand palette.
- Check the result for unintended resemblance to existing works, and label it as AI-generated where that's required or expected.