AI image generation gets easier when you stop treating the prompt as a wish and start treating it as a visual production brief. The model needs to know what you are making, what should appear, how the result should look, and which details matter most.
The goal is not to write the longest prompt. It is to remove ambiguity without burying the main idea under conflicting instructions.
This guide turns the principles in OpenAI's Image Gen models prompting guide into a practical workflow you can reuse for text-to-image generation and prompt-based editing.
Quick formula: purpose + scene + subject + style + composition + lighting + constraints.
Try the structure in PhotoArtify's AI image generator.
How to use the AI image prompting playbook
Use the cover infographic as a pre-generation checklist. Not every prompt needs every field, but each field should answer a real visual question.
1. Name the intended asset
Start by telling the model what kind of deliverable you need. The same subject will be composed differently for a product page, social post, mobile wallpaper, editorial illustration, or presentation cover.
Compare these openings:
- "A ceramic coffee cup on a table."
- "A landscape product-page hero image for a handmade ceramic coffee cup."
The second version establishes the job before describing the scene. This helps the model make better decisions about polish, framing, and usable negative space.
Useful asset labels include:
- product photo;
- editorial illustration;
- cinematic concept frame;
- social media advertisement;
- app onboarding visual;
- scientific diagram;
- seamless texture.
2. Describe the scene before decorating it
Build the visual from large decisions to small ones:
- Where is the scene?
- What is the main subject?
- What is the subject doing?
- Which supporting objects are necessary?
- What should the viewer notice first?
A strong core description might be:
A handmade ceramic coffee cup sits on a dark walnut table beside an open sketchbook. Morning light enters from a window on the left. The cup is the clear focal point, with generous empty space on the right for page copy.
This gives the model a visual hierarchy. Adding fifteen style adjectives before defining the subject usually produces less control, not more.
3. Use concrete visual language
Words such as "beautiful," "amazing," and "high quality" express approval, but they do not tell the model what to render. Replace vague praise with observable choices.
Instead of "make it cinematic," specify what creates that feeling:
- low-angle medium-wide framing;
- controlled backlight through haze;
- deep foreground and background separation;
- restrained color contrast;
- realistic lens depth;
- subtle film grain.
Concrete language gives the model visual evidence to construct.
4. Choose a style or medium deliberately
Style works best when it describes a coherent visual language. Name the medium first, then add the few qualities that distinguish the treatment.
Examples:
- clean studio product photography with realistic reflections;
- hand-painted gouache illustration on textured paper;
- architectural visualization with precise materials and daylight;
- bold risograph poster with flat spot colors and visible grain;
- documentary street photography with natural available light.
Avoid stacking unrelated aesthetics such as "minimalist, maximalist, photorealistic watercolor, neon vintage corporate." When the style instructions disagree, the model has to choose which ones to ignore.
5. Control composition with camera language
Composition instructions should explain where the camera is and how the subject fits inside the frame. Useful controls include:
- shot size: close-up, medium shot, full body, wide establishing shot;
- viewpoint: eye level, low angle, top-down, three-quarter view;
- lens behavior: macro detail, shallow depth of field, compressed telephoto look;
- placement: centered, left third, lower right, symmetrical;
- space: clean area for a headline, uncropped product, visible hands;
- format: square, portrait, landscape, or a specific aspect ratio.
For example:
Three-quarter product view at eye level, framed as a wide 3:2 composition. Place the cup on the left third and preserve clean negative space on the right. Keep the handle fully visible and do not crop the saucer.
These instructions are more actionable than "good composition."
6. Treat lighting as structure
Lighting defines form, material, mood, and time. A useful lighting instruction includes direction, softness, and purpose.
Try combinations such as:
- soft window light from camera left with gentle fill;
- hard midday sunlight with crisp architectural shadows;
- broad overhead studio light with controlled highlights;
- warm practical lamps balanced by cool evening daylight;
- bright backlight revealing condensation and glass edges.
If a product must retain accurate color, say so. If a face should remain natural, avoid contradictory lighting requests that would obscure features.
7. Quote exact text and specify its role
When an image must contain words, include the copy verbatim and explain where it belongs. Keep the amount of text limited, because every extra line creates another opportunity for an error.
Use a format like this:
Create a vertical event poster. Set the headline exactly as "DESIGN AFTER DARK" in large white uppercase sans-serif type near the top. Add "SEPTEMBER 18" as a smaller second line. Keep both lines unobstructed and preserve generous spacing around them. No other text.
Check spelling, capitalization, punctuation, hierarchy, and placement in the output. If the text is wrong, revise the text instruction alone before changing the entire design.
8. Use negative constraints selectively
Negative instructions are most effective when they target likely failures. A short list of relevant constraints is easier to follow than a generic wall of defects.
Useful examples include:
- no additional people;
- no cropped hands;
- no logos or watermark;
- keep the background uncluttered;
- preserve the product's exact silhouette;
- do not place objects in the headline area.
Keep the positive visual goal dominant. The prompt should still describe what the image is, not only what it must avoid.
9. Separate changes from invariants when editing
For image editing, define two boundaries:
- What should change?
- What must stay unchanged?
Use direct language:
Change only the plain studio background into a bright modern kitchen with morning daylight. Preserve the person's facial identity, pose, hairstyle, clothing, body proportions, camera angle, and crop. Do not add other people or alter the subject.
This pattern is more reliable than "put this person in a kitchen," which leaves the model free to reinterpret the whole image.
Open the prompt-based image editor.
10. Assign a role to every reference image
When using multiple images, label what each one contributes. Do not assume the model will infer your intended relationship between them.
For example:
Image 1 is the edit target: preserve the person's identity, pose, and clothing. Image 2 is a lighting reference only: apply its soft sunset direction and warm-to-cool color balance. Do not copy the face, wardrobe, or background from Image 2.
Reference roles can include subject identity, product shape, composition, palette, texture, lighting, or style. Explicit roles reduce accidental mixing.
11. Match prompt detail to task difficulty
A simple generation may need only a few clear sentences. A packaging mockup with exact text, material behavior, and layout needs a more structured specification.
Add detail when it resolves a decision. Remove detail when it repeats the same idea, introduces a conflict, or has no visible effect.
A practical structured prompt looks like this:
Asset type: landscape website hero image
Scene: a handmade ceramic cup on a walnut table beside a sketchbook
Subject: matte ivory cup with a blue painted rim, handle fully visible
Style: refined editorial product photography
Composition: cup on the left third, negative space on the right, 3:2 ratio
Lighting: soft morning window light from the left, gentle shadow falloff
Constraints: accurate cup shape and glaze, no text, no logo, no extra objects
The labels are not magic words. They simply make the brief easier to inspect and revise.
12. Iterate one variable at a time
After each generation, classify the biggest failure:
- subject or identity;
- composition;
- style;
- lighting or color;
- text accuracy;
- unwanted objects;
- edit drift.
Then change only the instruction connected to that failure. If the composition is correct but the light is too harsh, keep the scene and framing unchanged while revising the lighting line. If the product shape changes during an edit, reduce the requested transformation and strengthen the preservation rule.
Changing the subject, style, camera, color palette, and constraints at once makes it impossible to know what fixed or damaged the result.
Complete prompt examples
Product hero image
Landscape product-page hero image. A compact silver espresso machine sits on a clean dark stone counter in a contemporary kitchen. Refined commercial product photography, three-quarter eye-level view, machine placed on the right third with clean copy space on the left. Soft daylight from the left and a narrow highlight across the metal. Preserve realistic controls, seams, scale, and reflections. No people, text, logos, clutter, or cropped product edges.
Editorial character illustration
Editorial magazine illustration of an urban gardener carrying seedlings across a rooftop greenhouse at sunrise. Hand-painted gouache with visible paper texture, simplified geometric shapes, and confident brush edges. Medium-wide composition with the gardener as the focal point and layered city depth behind them. Coral sunrise, green foliage, cobalt shadows. No text, watermark, duplicated limbs, or extra figures.
Controlled portrait edit
Change only the background into a quiet library with warm practical lamps and subtle depth of field. Preserve the subject's exact facial identity, expression, skin texture, hairstyle, glasses, navy jacket, pose, hands, framing, and camera angle. Match the new background light naturally to the existing face. Do not retouch the face, change clothing, add jewelry, or add other people.
Final checklist
Before generating, confirm that the prompt answers these questions:
- What asset are you making?
- What is the primary subject and action?
- What visual style or medium should be used?
- How should the image be framed?
- What lighting and colors define the mood?
- Does any text need to appear verbatim?
- What must remain unchanged?
- Which likely failures should be prevented?
The strongest AI image prompt is usually the shortest version that makes the intended result unambiguous. Start with a clear visual brief, review the output like a creative director, and make the next instruction solve one specific problem.
Reference: OpenAI Cookbook: Image Gen models prompting guide.




