Nano Banana Pro Prompts for Fashion Campaigns
The model that rewards sentences over keywords. How to brief it, how to use references, and why most prompts written for Midjourney make it worse.
Nano Banana Pro is the first image model most directors have used that genuinely reads a brief. Not a keyword stack, not a bag of weighted tokens, a brief. That single difference changes how you should write for it, and it is why prompts carried over from Midjourney tend to produce weaker results here, not stronger ones.
It is the model we hold as primary inside the House, and the one we lock for previz. That is a considered choice rather than a fashionable one. When you are generating five sequential frames that have to feel like they came from the same shoot, on the same afternoon, in the same room, the model's ability to hold written instruction across a series matters more than any single frame's beauty.
This guide covers what the model actually rewards, the prompt structure we use, how to run reference images properly, and the failure modes that make its output look ordinary.
Write Sentences, Not Keywords
The instinct most directors bring from earlier models is compression. Strip the articles, stack the adjectives, weight the important terms, append the flags. That habit was correct for a generation of models that tokenised your prompt and weighted the fragments. It is actively counterproductive here.
Nano Banana Pro parses written language. It understands that “the coat is oversized, the shoulder dropped below the natural line” describes a garment's cut rather than listing two unrelated concepts. It handles spatial relationships, sequencing, and negation, all of which collapse into noise when you compress them into comma-separated tags.
The practical test: read your prompt aloud. If it sounds like something you could say to a photographer standing next to you, it is written correctly. If it sounds like a search query, rewrite it.
The same intent, written both ways
Compressed, in the old habit:
Written as direction:
The second is longer and that is the point. Every sentence resolves an ambiguity the first version left to the model: which way she faces, where her eyeline goes, where the light originates, what the shadow is allowed to do, which element carries the darkest value. The first prompt delegates all of those decisions. The second makes them.
The Structure We Use
The order matters, because the model weights early material more heavily. Our house order runs light before subject, and subject before garment. That is deliberately the opposite of how most people write, and it is the single highest-leverage change you can make.
1 — The room and the light
Open with where we are and where the light comes from. Name the source, the direction, the quality, and whether there is fill. “Cold north daylight from a high window on the left, no fill” specifies more than any number of quality adjectives.
2 — The subject and the pose
Who is in frame, how they are oriented to camera, and where they are looking. Eyeline is the most under-specified element in fashion prompting and the one that most often makes an image feel generic. “Looking back down the stairs rather than at us” is a directorial instruction. “Beautiful model” is not.
3 — The garment
Cut, weight, and behaviour rather than category. “Heavy wool, holding its own shape, barely moving” tells the model something. “Luxury coat” tells it nothing it can render.
4 — The camera
Format, focal length, and aperture behaviour. Medium format reads differently from 35mm in this model, and it responds to the distinction. Name the lens in millimetres rather than describing the effect.
5 — The finish
Palette, contrast, grain, and where the darkest and lightest values sit. Assigning value relationships — “the coat is the darkest thing in frame” — gives you compositional control that no amount of colour naming will.
Reference Images: One Job Each
Nano Banana Pro takes reference images, and this is where most of its advantage over text-only prompting lives. It is also where most directors waste it, by attaching three images of roughly the same thing and hoping the model averages them into a look.
Give each reference a distinct job, and say in the prompt what that job is. The model handles this instruction well and it is the difference between references that direct and references that muddy.
Three is the practical ceiling. Past that, references begin to compete, and the model resolves the competition by averaging — which produces something softer and more generic than any individual input. If you find yourself wanting five references, the problem is usually consistency rather than direction, and that is a different discipline with different tools.
Where It Is Genuinely Strong
Fabric and material behaviour
It renders the difference between wool and cashmere, between silk that falls and silk that floats. Name the fibre and its behaviour separately: “heavy raw silk, holding a hard fold rather than draping”.
Short in-frame typography
It is among the stronger models for a masthead, a label, or a single cover line set into the image. It is far less reliable for dense typographic layouts with several competing blocks. For a poster built around a type system, art-direct the type separately and composite it rather than asking the model to set it.
Negation and exclusion
Because it reads language rather than tokens, it handles “no visible jewellery, no makeup beyond skin” correctly, where older models would occasionally render the very thing you excluded. You can direct by subtraction, which is how most editorial styling decisions are actually made.
Four Failure Modes
It looks like a catalogue
You described a subject rather than a photograph. The model supplied the most ordinary treatment, and the most ordinary treatment of a fashion subject is e-commerce. Fix by leading with light and lens before the garment ever appears.
Everything is beautifully lit and nothing is interesting
You specified quality without specifying direction. “Beautiful soft light” averages toward flattering, frontal, and dull. Name a single source and a single direction, then let the shadow do the work.
The frames do not feel like one shoot
You rewrote the prompt between frames. Hold the room, the light, the camera and the finish constant as fixed text, and change only the subject and pose. Consistency across a series is a copy-and-paste discipline, not a creative one.
The face drifts across the campaign
This is the one problem prompting cannot solve on its own, no matter how carefully written. Identity across many frames is a reference and seed problem rather than a language problem.
Frequently asked
How is prompting Nano Banana Pro different from Midjourney?
Midjourney rewards compressed, weighted phrases and parameter flags. Nano Banana Pro rewards plain written direction in full sentences. Where Midjourney treats a prompt as a bag of weighted tokens, Nano Banana Pro reads it closer to a brief, which means spatial relationships, negation and conditional instructions land far more reliably. Write it the way you would write a note to a photographer.
How many reference images should I give it?
One to three, each with a different job: one for the subject or garment, one for the light and palette, one for the frame. Beyond three, references begin to fight and the model averages them into something softer than any of the inputs.
Can it render legible text in an image?
Yes, and it is among the stronger models for short in-frame type such as a masthead, a product label, or a single cover line. It is less reliable for dense typographic layouts. For a poster built on a type system, composite the type rather than generating it.
Why does my image look like stock photography?
Because the prompt described a subject rather than a photograph. Name only what is in the frame and the model supplies the most statistically ordinary treatment of it. Lead with light, lens and atmosphere, so the garment arrives inside a world that has already been specified.
What resolution and aspect ratio should I generate at?
Generate at the delivery ratio you actually need rather than cropping later, because composition is decided at generation time. For campaign stills, 4:5 and 1:1 cover most social placements and 16:9 covers film stills and web headers. Upscale the approved frame afterwards rather than asking for a larger generation.
Written direction, not keyword soup
Brief it once. Render it anywhere.
The Essenzi Creative Engine turns your intent into a complete brief, visual schema, shot list, and model-native prompts, then renders on Nano Banana Pro through your own connected key. Your judgement reaches the frame intact.
Try the Engine →