
Generating a convincing fashion image is no longer the hardest part of using generative AI in e-commerce. The harder problem is deciding, consistently and at scale, whether that image is accurate enough to publish.Â
A single output can look impressive in a presentation while still being wrong in ways that matter commercially. A closure may disappear, a pattern may drift, a hem may change length or a garment may fit differently from the item a customer will receive. When teams move from a pilot to thousands of stock-keeping units (SKUs), visual quality assurance becomes the real production constraint.Â
This changes the enterprise question. Instead of asking, “Can the model create a good image?”, leaders should ask, “Can our organisation detect, route and learn from bad images before they reach a customer?”Â
Photorealism is not product accuracyÂ
Traditional creative review often asks whether an image is attractive, on-brand and technically polished. Product imagery needs an additional test: does the visual preserve the commercial truth of the item?Â
Photorealism and product fidelity are different properties. A generated image may have coherent lighting, natural skin and a plausible pose while altering the garment’s colour, texture, construction or proportions. The better the image looks, the easier it can be for a reviewer to overlook a small but meaningful product error.Â
That distinction matters because a product image is not merely campaign decoration. On a product detail page it helps a customer judge silhouette, fit, material and styling. An error is therefore not just an aesthetic defect; it can create a mismatch between expectation and delivery.Â
The right unit of quality is not “a good image” in isolation. It is an image that is fit for a defined channel, use case and risk level.Â
Build an error taxonomy before automating reviewÂ
Teams cannot control what they have not defined. Before selecting automated checks or approval software, a fashion business needs a shared taxonomy of failure modes.Â
Product fidelity errors include changed logos, missing fasteners, invented seams, distorted prints, incorrect colour and altered garment length. Human realism errors include malformed hands, inconsistent anatomy, implausible contact between body and fabric, and lighting that does not match the scene. Brand and policy errors include unsuitable styling, prohibited contexts, weak representation, missing disclosure and imagery that cannot be traced to an approved source.Â
Each category should have an explicit severity. A minor background artefact might be acceptable for a short-lived social variant but not for a marketplace hero image. A changed logo, print or product shape should normally block publication regardless of channel.Â
This taxonomy turns subjective feedback into operational data. Reviewers can label a failure instead of writing “looks wrong”, engineering teams can see recurring defects, and procurement teams can compare systems using the same acceptance criteria.Â
Use risk-based review, not one universal checklistÂ
Not every image deserves the same review cost. The useful model is a risk matrix based on how visible the asset will be, how strongly it represents the product and how difficult an error would be to reverse.Â
A product-page hero image deserves stricter checks than an internal concept visual. A paid campaign scheduled across several markets carries more reputational and operational risk than an organic post that can be removed quickly. Images showing intricate prints, text, jewellery, transparent fabric or unusual construction may also require specialist review.Â
This allows a team to create review tiers. Low-risk outputs can pass automated checks and sample-based human review; medium-risk assets can require one trained approver; high-risk assets can require product, brand and legal approval. The point is not to add bureaucracy to every image, but to spend human attention where a failure would matter most.Â
Confidence scores should support this routing rather than act as automatic permission to publish. A system can be highly confident and still be wrong about the feature the retailer cares about.Â
Separate machine checks from human judgementÂ
Automated quality control is useful when the condition is measurable. Systems can test resolution, aspect ratio, file corruption, duplicate outputs, background consistency and the presence of required metadata. Computer vision can also flag possible changes in colour, logos, garment boundaries or key product features for closer inspection.Â
Human reviewers remain better placed to judge whether the garment still reads truthfully, whether styling fits the brand and whether an image could mislead a customer. Their role should be designed as exception handling and accountable judgement, not endless inspection of undifferentiated thumbnails.Â
The interface matters here. Reviewers should see the source product image beside the generated output, with zoom, variant information and the intended publication channel visible in the same view. A simple approve-or-reject button is insufficient unless rejection captures a reason code and routes the asset back to the correct step.Â
One approach we use at On-Model is to make QA a configurable stage in the image workflow rather than a separate manual hand-off. Review can be triggered automatically and routed to designated people in the customer’s organisation, or handled as a managed service by On-Model reviewers. Depending on the customer’s operating model, the process can expose review status and rejection reasons, or deliver only images that have been reviewed and approved against the customer’s criteria.Â
Sampling is also essential after automation is introduced. If a rule approves a class of outputs automatically, a random portion should still be reviewed to estimate false approvals and detect model drift. Otherwise, the organisation loses visibility precisely when throughput increases.Â
Treat provenance as production metadataÂ
An enterprise image pipeline should be able to answer basic questions after publication. Which source asset was used? Which model and configuration produced the result? What edits were made, who approved it and where was it published?Â
This information should travel with the asset or remain linked to it through a durable identifier in the digital asset management system. The Coalition for Content Provenance and Authenticity (C2PA) develops technical standards for recording the source and history of media through Content Credentials. Such credentials can support provenance, but they do not replace product review or prove that the depicted item is accurate.Â
The regulatory direction also favours traceability. The European Commission states that the AI Act’s Article 50 transparency obligations apply from 2 August 2026 and include requirements concerning identifiable AI-generated content. The exact obligation depends on the system and use, so companies should obtain legal advice for their circumstances rather than treating one generic label as universal compliance.Â
Even where a visible disclosure is not required, keeping machine-readable generation and approval records is sensible operational hygiene. Metadata is most valuable when a retailer needs to investigate a complaint, withdraw a group of assets or understand why a failure escaped review.Â
Measure escapes, not just output volumeÂ
Generative image pilots often report how quickly teams can produce assets or how many variants they can create. Those figures say little about whether the system is ready for production.Â
A stronger scorecard includes first-pass acceptance rate, rejection rate by reason, time spent reviewing each asset, regeneration rate and the number of defects found after publication. Teams should also track performance by garment category and channel, because an aggregate acceptance rate can hide a persistent weakness in prints, accessories or a particular image format.Â
The most important metric is the escape rate: the proportion of material defects discovered after an asset has been approved. Escapes should trigger a short root-cause review covering the model, input data, automated checks and human decision. The result may be a new validation rule, a revised prompt, a narrower approved use case or additional reviewer training.Â
Cost per approved asset is more informative than cost per generated image. Cheap generation followed by repeated rejection, manual correction and rework is not a cheap production system.Â
Quality assurance becomes a learning systemÂ
McKinsey has argued that generative AI in fashion should be approached as augmentation and acceleration, while warning that employees may fail to check errors and recommending processes for risk, ethics and quality assurance. That advice becomes concrete when every rejection produces structured feedback rather than disappearing into a chat thread.Â
Over time, labelled review decisions form a valuable evaluation set. Teams can use it to compare model versions, test workflow changes and identify categories that are ready for greater automation. Approval data can also reveal disagreement between reviewers, which is often a sign that the quality standard itself needs clarification.Â
This is why the operating model matters as much as the image model. Responsibility should be explicit: merchandising owns product truth, creative owns brand expression, legal interprets disclosure and rights requirements, and technology owns traceability and system performance. One named role should still be accountable for the final release policy.Â
Scale approval before scaling generationÂ
The competitive advantage in generative fashion imagery will not come from producing the largest pile of images. It will come from creating more usable, truthful assets with a known level of risk.Â
Before increasing volume, teams should be able to define a material defect, route assets by risk, compare outputs with source products, preserve provenance, audit approvals and measure failures after publication. If those controls work for one category and channel, the workflow can expand deliberately.Â
Generation is becoming abundant. Trustworthy approval is the scarce capability, and it is the part of the system enterprises now need to design.Â


