AI & Technology

The pretty image problem. How to test whether an AI model changed the product you sell

By Bilal Azhar

Standfirst. AI image tools can produce convincing campaign scenes while quietly changing the item being advertised. A 96-run test across Flux 2, GPT Image 2, Qwen Image 3, Seedream 5.0 Pro, Nano Banana 2, and Luma Uni 1 found usable output rates from 50 to 100 percent, and the gap widened on lifestyle scenes. 

An AI-generated product image can look finished long before it is accurate. The lighting works, the composition feels expensive, and the product is recognisable at a glance. Then somebody notices that the jacket has gained a pocket, the bottle label has changed, or the kettle’s switch has moved. 

That is not a cosmetic error. It means the image may show a product the customer cannot buy. 

For ecommerce teams, the useful question is no longer whether an image model can make a polished scene. Most leading models can. The harder question is whether the model can change the scene without redesigning the product. 

A benchmark should start with the product, not the prompt 

We ran a small internal benchmark to examine that problem, using Morphed’s product image tools to send one source image to each model under identical instructions. The test used four original, photorealistic synthetic products: a serum bottle and carton, an electric kettle, a chore jacket, and a structured daypack. 

Each product had a frozen list of identity features before generation began. The serum, for example, had exact bottle and carton text, a coral stripe, a silver collar, and a clear cap. The jacket had an asymmetric chest pocket, three lower patch pockets, five visible front buttons, cream stitching, and a small orange tab. 

This feature map matters because ‘looks like the same product’ is too vague to score consistently. A merchandiser needs to know which details are allowed to change under new lighting and which details define the item a customer will receive. 

Every model received the same preservation instruction and one of two assignments. The studio task changed the background, surface, lighting, and composition. The contextual task placed the product in a plausible lifestyle setting while keeping it visible and unobstructed. 

We tested six live image-editing endpoints: Flux 2 Edit, GPT Image 2 Edit, Qwen Image 3 Edit, Seedream 5.0 Pro Edit, Nano Banana 2 Edit, and Luma Uni 1 Edit. Two independent requests were made for every product and task combination, producing 96 accepted submissions. 

A beautiful output can still fail 

We separated three judgments that creative teams often collapse into one. 

First, was the product faithful? Every required critical and major feature had to pass. The output also had to avoid invented claims, missing or duplicated components, and changes to the product’s material, colour family, geometry, seams, hardware, or function. 

Second, did the image follow the assignment? A studio transformation that cropped the product or a lifestyle scene that covered a required detail failed task suitability, even if the visible portion looked good. 

Third, was the image visually coherent? A faithful product inside a broken scene is not ready for commercial use. 

We counted an output as usable only when it passed all three checks. This is stricter than asking a reviewer which picture they prefer, but it maps more closely to the decision a marketing team has to make before publishing. 

The two reviewers agreed on usable status for 83 of 92 visible outputs, with Cohen’s kappa of 0.610, and on 820 of 828 individual feature decisions. Disagreements went to a third blinded session. 

The distinction also reflects current commerce requirements. Google Merchant Center’s image guidance says product images should accurately display the item and match attributes such as colour, pattern, and material. A generated image may be attractive and still create a mismatch between the creative, product feed, and landing page. 

The models did not fail in the same way 

Across this protocol, usable output rates ranged from 8 of 16 to 16 of 16. 

Model  Usable  Faithful  Context  Credits / usable 
GPT Image 2 Edit  16/16  16/16  8/8  12.0 
Qwen Image 3 Edit  15/16  16/16  7/8  12.8 
Seedream 5.0 Pro Edit  14/16  15/16  7/8  11.4 
Flux 2 Edit  12/16  14/16  5/8  13.3 
Nano Banana 2 Edit  10/16  11/16  3/8  12.8 
Luma Uni 1 Edit  8/16  9/16  2/8  13.0 

The contextual scene exposed the larger gap. GPT Image 2 produced eight usable contextual outputs out of eight. Qwen and Seedream each produced seven, Flux produced five, Nano Banana 2 produced three, and Luma Uni 1 produced two. 

That result should not be read as a universal ranking. There were only 16 submissions per model, one synthetic product per category, and two repetitions for each model, product, and task combination. Model endpoints also change. 

At this sample size the 95 percent Wilson intervals overlap widely. The top result spans 80.6 to 100 percent and the bottom result spans 28.0 to 72.0 percent, so the middle of the table should be treated as a cluster rather than an order. 

The operational lesson is more durable than the ordering. A model that works for a clean studio transformation may behave differently when it has to place the same product among props, people, surfaces, and stronger lighting changes. 

Test the workflow you intend to ship 

A useful internal test does not require 96 runs. It does require discipline. 

Start with a small set of products that are hard in different ways. Include packaging with exact text, an object with functional geometry, apparel with countable construction details, and an accessory with hardware or closures. 

Write the identity checklist before anyone sees generated outputs. If the team adds a new criterion only after disliking a result, the scoring will follow taste rather than product truth. 

Use the same task instructions across models. Avoid quietly improving one model’s prompt after seeing it struggle unless prompt optimisation is itself the thing being tested. 

Keep generation failures in the denominator. Four of our accepted requests returned no visible output. A production team still pays for the delay, the retry, or both, so excluding those cases makes the workflow look more reliable than it was. 

Finally, record uncertainty. The NIST AI Risk Management Framework recommends documented, repeatable testing with benchmarks and measures of uncertainty. For a small ecommerce team, that can be as simple as preserving prompts, model versions, outputs, reviewer decisions, and the reason each image passed or failed. 

Human approval needs a better checklist 

‘Keep a human in the loop’ is good advice, but it is not a quality-control method by itself. A hurried reviewer is likely to approve the overall composition and miss a changed seam, label, button count, or material. 

Review the product before the scene. Compare text, colour family, component count, geometry, seams, closures, hardware, and any claim visible on the item. Then review whether the assignment was followed and whether the surrounding image is coherent. 

The person approving the creative should also know which source image and product variant the model received. Without that reference, they are judging plausibility, not accuracy. 

This is especially important when one generated asset moves into several channels. A small alteration can spread into paid social, marketplace listings, email, and retailer materials before anybody compares it with the source product. 

Model choice should follow the failure cost 

The cheapest generation is not necessarily the cheapest usable asset. In our test, credits per usable output clustered more tightly than headline generation prices because lower pass rates consumed the apparent saving. 

Teams should price the whole approval loop. Count generation, retries, reviewer time, manual repair, and the cost of pulling an inaccurate asset after publication. 

The right model may also change by task. One endpoint can be the sensible choice for rapid studio variations while another earns its higher cost on lifestyle scenes where product drift is harder to detect. 

AI image generation is already good enough to create false confidence. The next improvement for ecommerce teams is not another adjective in the prompt. It is a test that can tell the difference between a better scene and a different product. 

Disclosure and limitations 

Morphed funded and ran this benchmark, and sells access to every model included in it, so it has a commercial interest in the subject. Morphed created the synthetic products, selected the endpoints, paid for the production runs, and prepared the results. The model providers did not sponsor, review, or approve the study. 

Two independent AI reviewer sessions scored model-free packets, and a third model-blind AI session adjudicated disagreements. The study was internally pre-specified but not externally preregistered or independently conducted. It did not test sales, conversion, advertising approval, marketplace compliance, legal compliance, or whether Morphed’s agent chooses the best model. 

Author bio 

Bilal Azhar is the founder of Morphed, an agentic creative workspace for planning and producing image and video campaigns across multiple AI models. He works on practical evaluation methods for creative teams using generative media in production. 

Related Articles

Back to top button