Most "which AI image model is best" comparisons run a handful of cherry-picked prompts. We had something better: real client catalogues, where every image had to be approved by the brand before it went on the shop. One was a European furniture and art brand with 601 products. The other was a German cosmetic brand with 300 SKUs.
Across both catalogues and our own users, the answer to "which model is better" is both, for different jobs.
- Nano Banana Pro produced the look. It shot the furniture brand's entire hero catalogue and almost every lifestyle scene.
- GPT Image 2 was the model that got text right: the brand lettering, the label, the small print. It was also the only one willing to edit a logo at all.
Here is the data.
A 601-product furniture and art catalogue, reviewed by the client
Most pieces in this catalogue carry a small oval gold plaque with the maker's signature. We built the brand a studio pipeline: one reference photo per product in, a consistent studio hero shot and a set of lifestyle scenes out, with the client approving or rejecting every image in a review grid.
The hero shots: Nano Banana Pro, 9% redone, zero rejected
The final studio packshot for the whole catalogue came from Nano Banana Pro and one short prompt: a seamless #F6F6F6 backdrop, the product small and centred, a soft key light from the front left.
| Nano Banana Pro hero shots | Count |
|---|---|
| Products | 601 |
| Images generated | 663 |
| Approved by the client | 591 |
| Products that needed a second attempt | 55 (9%) |
| Rejected by the client | 0 |
GPT Image 2 is a capable studio photographer too, and its hero shots were approved as well. The difference showed in the lifestyle scenes: window-shadow rooms, Viennese salons, dappled hallways. Nano Banana made 1,028 images across 251 scene designs. GPT Image 2 was tried on 9 scenes. Nano Banana gave us the light and the mood.
The logo: where Nano Banana breaks
Then the client's complaint came in: the plaque looks wrong.
We ran a vision scan over the whole catalogue. The plaque was visible on 137 images, and its lettering was garbled on 135 of them. Nano Banana had made 135 of those 137 images, and only one of its plaques was readable. The gold oval itself was perfect every time. The signature inside it came out as letter-shaped noise.
So we tried to repair the plaques. Each model got the broken image plus a clean photo of the real plaque, with an instruction to fix only that:
| Plaque repair | Attempts | Approved | Rejected |
|---|---|---|---|
| GPT Image 2 | 41 | 20 | 14 |
| Nano Banana Pro | 12 | 0 | 12 |
Nano Banana did not get a single repair past the client. GPT Image 2 got 59% of the repairs that had been reviewed. (The other seven were still waiting for review.)
The fix that worked: take the logo out first
The lasting fix was to stop asking any model to draw the plaque. We removed it from the reference photos and put the real one back afterwards. That produced two more findings:
- Nano Banana refused to do the removal. Taking a maker's mark out of a photo reads as "removing branding" to its content filter. GPT Image 2 repainted the patch without complaint. We pasted only that patch back, so every other pixel of the photograph stayed untouched.
- Show, don't tell. Telling Nano Banana to leave the plaque out worked on 8 of 12 test images. Giving it a reference with no plaque in it worked on 10 of 10. Controlling what a model sees beats telling it what not to do.
A German cosmetic brand: 36 products, 1,056 renders
Cosmetics are the hardest case for AI text. Every compact, tube and pencil carries a wordmark, a product name, an organic badge and small print, often in German. For this brand's 300 SKUs we needed eight images for each of its 36 products, and we made them all with Nano Banana Pro through the ProductAI API.
| Nano Banana Pro, cosmetics | Count |
|---|---|
| Renders completed | 1,056 |
| Images planned (36 products × 8) | 288 |
| Images delivered | 232 (81%) |
| Renders per delivered image | about 4.5 |
| Rejected by the client in the first review | 20% |
| … of those, approved after a redo | 83% |
The client marked every image in a shared review deck: one in five was sent back in the first round. Most came back approved after one redo. The rest were either dropped from the shot list or are still open.
The look was never the problem. The label was. Our prompts had to spell out every piece of text on the pack, word for word, from the wordmark down to the product name in capitals. The slots still open are mostly the smallest items (pencils, eyeliners, a highlighter compact and the eyeshadow palettes), where the lettering is tiniest. On a catalogue like this, a model that gets small text right first time saves most of those extra renders. That is the job we now give GPT Image 2.
What we see on ProductAI
ProductAI now runs a fidelity check on subscriber renders. A few seconds after an image finishes, a vision model compares it with the seller's product photo. It flags only changes to the product: colour, label or text, pattern, shape or parts, size, and placements the real object cannot hold. The scene, the light and the angle are supposed to change, so it never flags them.
The first checked renders point the same way:
| Model | Checked | Product changed | What changed |
|---|---|---|---|
| GPT Image 2 | 11 | 0 | nothing |
| Nano Banana Pro | 19 | 7 | 6 shape or parts, 1 label |
| Seedream 5 Pro | 6 | 2 | 1 label, 1 size |
The fairest slice came from one seller who ran the same labelled bottle through all three models. GPT Image 2 got the label right 9 times out of 9, Nano Banana 4 of 5, and Seedream 5 of 6. The misses were classic AI text: "MADE WITHOUT PARABENS" came back as "WADE WITHOUT FAROASAÍNÉ" and as "MADE WITNONT PARABENS".
Nano Banana's other misses match what we saw on the furniture catalogue: it takes creative liberties with the object itself. A relaxed sweatshirt became a fitted bodysuit, a wing-shaped eye massager became a single small patch, and a zipped bag was shown open. They were good-looking images. The product just was not quite the product any more.
Which images do people actually keep?
A download is the most honest vote a user can give an image. We matched every studio download on ProductAI from the last 30 days to the generation it came from:
| Model | Downloaded |
|---|---|
| Nano Banana Pro | 37% |
| GPT Image 2 | 40% |
| Seedream 5 Pro | 58% |
| Seedance 2.0 (video) | 54% |
Two things stand out. Nano Banana Pro is by far the most used model, and people keep more than a third of what it makes. GPT Image 2's rate is in the same range. Its lead comes from one heavy user, and without that user it falls to 27%. Seedream's 58% is almost all one seller, who made 86% of its images; everyone else kept 30%.
So download rates do not crown a winner. People keep images from both models at similar rates. What differs is which images: the ones with text on the product are where GPT Image 2 earns its place. Note that "not downloaded" does not mean "rejected". Most people generate several variations and keep the best one.
Which one should you use?
| Use | When |
|---|---|
| GPT Image 2 | Lettering, a label, a logo or packaging copy is visible on the product |
| GPT Image 2 | You are fixing or removing a logo or detail in an existing image |
| GPT Image 2 | Exact shape matters: apparel fit, devices with specific parts |
| Nano Banana Pro | Styled lifestyle scenes, mood, light, interiors |
| Either | Clean studio packshots at scale |
Two practical tips from the catalogue work:
- Keep GPT Image 2 prompts short. A long prompt full of prohibitions ("keep the product exactly as it is, do not change anything…") made it return the image untouched. One clear instruction works better.
- If a logo keeps coming out garbled, stop prompting harder. Remove it from the reference and put the real one back afterwards. Neither model reliably redraws small lettering seen from a distance.
In ProductAI you can switch between GPT Image and Nano Banana on every generation. Use Nano Banana for the scene and GPT Image when the label has to be right.
How we measured this
- The catalogue numbers are client decisions, not a lab benchmark. The cosmetic catalogue was made with Nano Banana Pro only, so it shows the cost of small text on packaging, not a head-to-head. "Approved" means the brand accepted the image for its shop. The models were not always given identical jobs: Nano Banana did the bulk generation, and GPT Image 2 did more of the repair and editing work.
- The text result is a head-to-head: the same plaque and the same repair task, with 20 repairs approved against 0. The style result is our judgement, backed by what we chose to ship across 251 lifestyle scenes rather than by a scored side-by-side.
- Download rates cover ProductAI studio generations from the last 30 days. Each download event was matched to its generation by user and prompt (798 of 836 matched), because the event itself did not yet record which model made the image.
- The ProductAI fidelity sample is still small. These were the first checked renders, from a handful of accounts. The judge is a vision model with a deliberately conservative prompt that answers "ok" when in doubt. We will update the numbers as the sample grows.
%20(1).avif)
