When to use it
You need to generate an image (a cover, an illustration, a concept, a product shot), but the first attempts come out "not quite it". The AI plays art director, turning your idea into a precise prompt for Midjourney/FLUX/SDXL. The result: a coherent scene description plus parameters that gives you a predictable image instead of a lottery ticket.
The prompt (copy and paste)
You are an art director. Build a prompt for the image generator <FILL IN: Midjourney v7 / FLUX / SDXL>.
MY IDEA: "<FILL IN: what I want to see>".
PURPOSE: <cover / in-article illustration / avatar / product ad>.
MOOD: <e.g. calm, premium / bold, energetic>.
Build the prompt as ONE coherent description, in this order:
1. The main subject (one, specific — who/what, pose, a defining detail).
2. Environment/background (where, time of day).
3. Light (type and direction — e.g. soft backlight at sunset).
4. Style/medium (35mm photo / watercolour / 3D render / editorial illustration) and a reference aesthetic.
5. Framing (shot size, angle, composition — e.g. close-up, rule of thirds).
6. Parameters at the end: aspect ratio, and for Midjourney --ar, --s (stylize), plus --chaos if needed.
Write in natural language (not a comma-separated tag dump). One subject in focus. At the end, add a "Avoid:" block — what to exclude (text, extra hands, watermarks).
Why it works: modern models understand a coherent scene description better than a pile of keywords; the explicit "subject → light → style → framing → parameters" order gives you control and repeatability.
Filled-in example
Idea: "a mug of coffee on a table by a window, cosy". Purpose: cover for an article about focus. Mood: warm, calm. Platform: Midjourney v7.
What the AI should return (roughly): "A single ceramic mug of black coffee steaming on a worn oak table by a tall window, quiet morning, soft warm side light angling through the glass, shallow depth of field, editorial photography, 35mm, close-up, rule of thirds, calm cozy mood --ar 16:9 --s 250. Avoid: text, people, clutter, watermarks."
Variations
- A video frame. Same skeleton, but add camera and subject movement (for Sora/Veo/Kling) — see the card on video generation.
- A series in one style. "Give me 4 prompts for the same scene from different angles, style and light unchanged" — for a consistent set (in Midjourney, lock the style with --sref).
- From a reference. "Describe the aesthetic of this image in words so I can reproduce the style on a different subject" — style extraction.
Pro tips
- One subject in focus: "a cat AND a dog AND a bird" confuses the model and produces mutants. If you need a crowd, describe it as a single object in the scene ("a group of people around a campfire").
- Light and framing decide whether an image looks expensive or cheap far more than a list of adjectives: "soft backlight, 35mm, close-up" lifts a picture more than ten epithets like "beautiful, stunning, masterpiece".
- Parameters are levers, not decoration: a low --s (stylize) means literal adherence to the text, a high one means more artistic licence. Start in the middle and move one at a time, or you won't know what worked.