Illustration of a ginger cat sitting between two picture frames showing different stylized portraits of the same cat

Comparing image models through their own showcase galleries is like comparing restaurants through menu photography. Everything looks perfect, nobody shows you the soggy fries, and the winner is always whoever paid for the lighting.

So I ran a less scientific but much more useful test: I gave competing image models the same subjects, the same reference images, and the same instructions. One prompt was pure text. Two were edits where preserving the original mattered.

The result surprised me. The image I found more visually dramatic was not always the image I considered better.

The setup

I had three routes available: Google's Gemini 3.1 Flash Image Preview, better known as Nano Banana 2; Gemini 3 Pro Image Preview, or Nano Banana Pro; and OpenAI's GPT Image 2.

Google positions Flash as the fast, high-volume model and Pro as the professional asset-production model. Its current API supports 1K, 2K, and 4K output on both, with a smaller 512-pixel option on Flash. GPT Image 2 takes a different approach to edits: OpenAI's documentation says every reference image is automatically processed at high fidelity rather than exposing a fidelity slider.

Those descriptions are useful, but they are still product descriptions. I wanted to know what each model would sacrifice when the prompt forced it to choose.

Test one: make a cat look expensive

The first prompt was a Renaissance-style portrait of a tabby duke. No reference image, so neither model had an identity to preserve. This was just composition, texture, lighting, and taste.

Nano Banana 2 produced the warmer painting. It looked like something that had spent a few centuries collecting varnish in a museum: soft brushwork, rich brown tones, and a convincingly grand cat.

GPT Image 2 produced the cleaner image. The face, costume details, and edges were sharper, and the whole thing looked more polished at first glance. It also took substantially longer in my run.

Nano Banana 2’s warmer Renaissance tabby duke

GPT Image 2’s sharper Renaissance tabby duke

The original test-one outputs: Nano Banana 2 above, GPT Image 2 below.

I could happily keep either one. That is important, because it exposed the weakness of the usual “which model is best?” question. For a text-to-image prompt with no hard constraints, preference can collapse into taste. Warm and painterly is not objectively better than crisp and controlled.

The edits were much less forgiving.

Test two: do not redesign my cat

For the second test, I uploaded a real photo of Rosie, my calico, lying in her excellent sploot pose. The instruction was to turn the photo into a flat cartoon while preserving her identity, markings, pose, and proportions.

Both models clearly understood that Rosie was Rosie. Her calico patches survived, the pose remained recognizable, and neither result turned her into a generic orange mascot.

But they made different trades.

GPT Image 2 preserved the geometry better. Rosie's long sploot, silhouette, and body proportions stayed close to the photo. The result looked like a cartoon translation of the same moment.

Nano Banana Pro pushed the style harder. It made a bolder, cleaner cartoon, but it also enlarged her head, rounded her body, and shortened her legs. In other words, it chibi-fied her. Cute? Absolutely. Faithful? Less so.

GPT Image 2’s Rosie cartoon edit

Nano Banana Pro’s Rosie cartoon edit

The original Rosie edits: GPT Image 2 above, Nano Banana Pro below.

That was the moment the prettier image lost.

When I ask for an edit, creativity is not automatically a virtue. If the input contains the thing I care about, then changing it is not interpretation; it is drift. A model can produce a more attractive picture and still fail the assignment.

Test three: preserve the weird little character

The final test used a minimalist white blob character I already had. I asked both models to place it at an office computer, visibly frustrated, while preserving the character's simple shape. I also asked for small interface details and a lowercase label in the bottom-right corner.

Again, both outputs were good. Both kept the character recognizable. Both understood the scene and rendered the requested labels correctly.

GPT Image 2 stayed closer to the original character shape and produced coherent text inside the error dialogs. Nano Banana Pro made a strong stylized scene, but most of the tiny log text dissolved into convincing-looking gibberish.

GPT Image 2’s frustrated character scene

Nano Banana Pro’s frustrated character scene

The original character-preservation outputs: GPT Image 2 above, Nano Banana Pro below.

That is not a shocking failure. Google's own documentation promotes advanced text rendering, while OpenAI's documentation says text has improved but can still struggle with precise placement and clarity. Both statements can be true. A headline, logo, or menu title is one problem; a tiny wall of diagnostic text is another.

The practical lesson is simple: if readable text matters, test the actual size and density you plan to ship. “Supports text” is not a useful acceptance criterion.

The scorecard I actually care about

After three tests, I do not have a universal winner. I have a routing map.

For reference-heavy edits where pose, silhouette, or character identity matters, GPT Image 2 currently has the edge in my small sample. For stronger stylization, Nano Banana Pro is more willing to transform the source. For quick text-to-image drafts, Nano Banana 2 is fast and visually strong.

A fair cost comparison would need both models running through equivalent metered API routes, which this test did not use. So I am not scoring cost. Pretending otherwise would create a beautifully precise spreadsheet measuring two different things.

The bigger change is how I will test models from now on. I want at least one unconstrained generation, one identity-preserving edit, one text-heavy scene, and a latency check. I also want to decide before generating which failures matter. If I do not define that first, I will just pick whichever output looks coolest and call it a benchmark.

This was only a handful of images, not a laboratory result. A different prompt, style, or reference could reverse every ranking. But the small test already gave me something more valuable than a leaderboard: it showed me where each model bends.

The right question is not “Which model made the prettiest image?” It is “Which model protected the part of the prompt I could not afford to lose?”