Generating an image does not begin with the prompt. It begins by defining what would make that image correct.
That distinction changes the entire process. When purpose, format, identity and constraints remain implicit, the model has to fill the gaps with decisions that may be plausible but not useful. The more important the application — a campaign, interface, recurring character, diagram or game asset — the less the result should depend on that guesswork.
GPT Image, Higgsfield and Nano Banana all accept natural language and visual references, but they distribute control differently. In one workflow, much of the direction lives in the written brief. In another, presets, identity and palette work as persistent controls. In another, conversation and multi-image composition form the main creative space. Understanding this difference is more valuable than searching for one universal “perfect prompt.”
A prompt is only one part of the input
An input is anything that conditions the result. It can include:
- written instructions;
- an image that must be edited;
- separate identity, product, pose, composition or style references;
- aspect ratio, resolution and output format;
- a preset, moodboard, palette or trained character;
- context accumulated in the conversation;
- brand, production or accessibility constraints.
The prompt is only the textual part of this system. If a reference already defines a face, the text should not compete by redescribing it. If a preset defines the photographic language, the prompt can focus on action and framing. If nothing anchors composition, the prompt must take on that responsibility.
A useful rule is: use text to define what the other inputs have not already defined.
Turn intent into a visual brief
A reliable brief can follow this order:
- Purpose: where the image will be used and for whom.
- Deliverable: asset type, aspect ratio, resolution and transparency needs.
- Scene: environment, period, weather and background elements.
- Subject: who or what appears, essential traits and action.
- Details: clothing, objects, materials, textures and realism cues.
- Composition: shot size, angle, placement, scale, focus and negative space.
- Style: photography, illustration, 3D, editorial language and color treatment.
- Lighting and atmosphere: direction, quality, contrast, time and feeling.
- Invariants: everything that must remain unchanged.
- Exclusions: anything that would make the result incorrect.
- Exact text: literal copy, typographic hierarchy and placement.
This does not mean every prompt needs eleven blocks. The structure is a diagnostic checklist. Lines that do not change a decision can be omitted.
GPT Image: write a production brief
In Codex, built-in generation is performed by GPT Image. The agent organizes context, references and iterations; the image model synthesizes or edits the asset.
This workflow responds well to an explicit sequence: scene → subject → details → composition → style → constraints. The intended use also helps distinguish an editorial illustration from a UI mockup or a transparent production asset.
For editing, the most important instruction separates change from preservation:
Change only the lighting to an overcast late afternoon. Preserve identity, pose, framing, geometry, clothing, objects and background. Do not add text or new elements.
When several images are involved, assign a role to each:
- Image 1: base image to edit;
- Image 2: character identity;
- Image 3: product that must be preserved;
- Image 4: lighting and color reference only.
Explicit constraints are effective here. “No watermark,” “no additional text” and “do not change the background” are operational instructions, not side notes.
Higgsfield: distribute direction across controls
In Higgsfield Soul, the prompt is only one layer. Soul ID can lock a person; presets and moodboards define visual language; Soul HEX anchors the palette; a reference image can carry pose, composition, lighting and atmosphere.
The strongest brief therefore does not repeat every control. Once identity and aesthetics are selected, the prompt can focus on:
Woman holding a black bottle at chest level in a concrete hotel corridor. Medium shot, direct gaze, asymmetrical composition, restrained tension and negative space on the left.
For photography and cinema, concrete terms are more useful than abstract adjectives: wide or close, low-angle or eye-level, side or diffuse light, static or moving subject. “Beautiful” and “cinematic” provide little direction until they are translated into visible decisions.
The differentiator is persistence. In a recurring campaign, identity and style do not need to be reinvented with each generation; they can become reusable assets.
Nano Banana: compose references by responsibility
Gemini image models are natively multimodal and conversational. This favors requests that combine characters, objects, environments, real-world knowledge and successive modifications.
When several references participate, name the responsibility of each one:
Use Image 1 to preserve Lia’s identity, Image 2 as the exact product reference and Image 3 for lighting only. Create a vertical 9:16 campaign in a brutalist corridor. Preserve the face, packaging, logo and proportions. Change only pose, setting and light direction.
For complex scenes, describing the construction in steps reduces ambiguity. Instead of accumulating negative phrases, also describe the intended positive state: “a completely empty street with no traffic” communicates the composition better than “no cars” alone.
Conversation is part of the control surface. After the first image, a useful iteration changes one variable at a time: angle, lighting, action, setting or treatment. Rewriting the entire brief increases the chance of losing decisions that were already correct.
References need explicit roles
An attached image can mean several different things:
- edit target: the image itself must be modified;
- approved invariant: geometry, identity or brand elements cannot change;
- identity: face, body or character;
- product or object: shape, label and materials;
- composition: camera, pose and spatial arrangement;
- style: color, texture, grain or visual language;
- setting: architecture and recurring environment elements.
Without this classification, a model may apply the product’s aesthetic to the setting, copy a pose when it should preserve only the face or redesign an element that should remain untouched.
Define metrics before generating
“I like it” remains relevant, but it is not enough for a production workflow. Before generating, select the criteria that actually determine success:
- adherence: the subject, action and required elements are present;
- composition: framing, hierarchy, negative space and readability fit the use;
- identity: face, character, product and brand remain recognizable;
- text: literal content, spelling, count and legibility are correct;
- visual coherence: lighting, perspective, scale, materials and shadows agree;
- brand fit: palette, tone and language follow the approved system;
- technical delivery: aspect ratio, resolution, transparency and format suit the consumer;
- absence of critical faults: no extra elements, artifacts, marks or forbidden changes.
Not every asset needs to maximize every criterion. Concept exploration tolerates aesthetic variance; packaging needs label integrity; a recurring character needs consistency; a banner needs the correct content area.
When to ask before generating
Questions are necessary only when a missing detail could materially change the result or prevent evaluation. The conversation should seek the highest-impact information first and ask one question at a time.
A useful order is:
- what is the purpose and where will the image be used;
- which format or framing is mandatory;
- what must remain recognizable or unchanged;
- which visual direction should guide the image;
- which text must appear exactly;
- what would make the result unacceptable.
As soon as the brief contains enough information to generate and evaluate an output, the interview stops. Lower-impact preferences can receive reversible defaults. The goal is not to turn every request into a questionnaire, but to keep critical decisions from being made accidentally.
Iterate as if debugging a visual system
After the first output, evaluate it against the selected metrics. Identify the largest divergence and request one isolated correction:
- “keep everything and increase only the negative space on the right”;
- “restore the reference face without changing clothing or lighting”;
- “correct only the text to ‘Yours to Create.’, exactly once”;
- “preserve the product and reduce only the background contrast”.
This turns generation into an observable loop: brief → image → evaluation → isolated change → new evaluation. The value is not in writing the longest prompt, but in reducing the distance between intent and evidence.














