Skip to content
Los Angeles · Independent fashion intelligence
Fashion × AI

How AI Size Recommendation Systems Make Decisions

A size recommendation is a probability built from product, customer, and outcome data—not a measurement of objective truth.

A black jacket, measuring tape, body-measurement diagram, size cards, and probability ranges arranged on a technical-design table.
AI-generated educational editorial image. The garment, person diagram, and recommendation values are fictional. Created with OpenAI ImageGen for FashionMember.

An online size recommendation can look deceptively simple: “We recommend M.” Behind that sentence is rarely a direct measurement of whether a garment will fit. More often, the system estimates which available size is most likely to produce an acceptable outcome for a particular shopper and product.

That distinction matters. Size labels are inconsistent, bodies are multidimensional, garment ease is intentional, and “fit” includes personal preference. A system can reduce uncertainty; it cannot turn a letter on a tag into universal truth.

The four inputs

1. Product and variant data

The system needs a stable product identity, every sellable size variant, the size system, garment category, cut, material, and measurements when available. Google Merchant Center’s current apparel guidance requires separate size variants to be submitted as separate items grouped under the same item-group ID. It also provides size, size_type, and size_system attributes to distinguish values such as petite, tall, plus, US, UK, EU, or JP sizing.

Those fields help platforms understand the catalog, but they do not describe the full geometry of a garment. A useful internal record may also include chest, waist, hip, inseam, rise, shoulder, sleeve, garment length, stretch, intended ease, and the measurement method used by the technical team.

2. Shopper evidence

Depending on the retailer and the shopper’s consent, evidence can include:

  • measurements the shopper enters;
  • a size that fits in a reference brand or product;
  • prior purchases and retained items;
  • returns labeled too small or too large;
  • explicit fit feedback;
  • fit preference such as close, regular, or relaxed;
  • category-specific history, because a shoe history may not predict a blazer size.

This is personal data. A retailer should collect only what it can protect and explain, provide a non-personalized path, and avoid implying that the model has measured a body when it has not.

3. Outcome labels

Models learn from an operational definition of success. A purchase that was not returned may be treated as a positive signal, but “not returned” does not necessarily mean “fit well.” The customer may have missed the return window, altered the garment, given it away, or accepted a mediocre fit.

Return reasons are also noisy. “Too small” may describe the waist, shoulder, sleeve, or the customer’s preferred ease. Exchanges can be more informative than returns because they show the direction of correction, but only if the data connect the original and replacement variants.

4. Context

The same shopper may prefer a fitted knit top and an oversized coat. Recommendations should therefore account for category, silhouette, material behavior, intended styling, and sometimes season or layering. A single permanent “customer size” is usually too crude.

Three common modeling approaches

Rules and size charts

The most transparent system maps a measurement or reference size to a brand chart, then applies category rules. It is easy to audit and works for a new product, but it depends on accurate charts and cannot learn much from outcomes.

Latent or Bayesian fit models

Published research from Amazon described a latent-factor model that estimates whether a product-size combination is likely to be small, fit, or large for a customer. A separate hierarchical Bayesian paper jointly modeled purchased size and possible return outcomes such as no return, too small, or too big. These approaches learn relationships among customers, products, sizes, and outcomes while sharing evidence across sparse groups.

Deep learning systems

Deep models can combine product attributes, customer behavior, fit feedback, and learned representations. A 2019 paper on fashion ecommerce reported experiments on proprietary expert-feedback and purchase datasets. The important limitation is as notable as the method: performance on a private historical dataset does not prove that the same system works for a new retailer, category, body population, or size architecture.

What the output should mean

A responsible interface should treat the recommendation as a ranked estimate, for example:

  • M: most likely to match a regular-fit preference;
  • L: plausible if the shopper prefers more room;
  • confidence: limited because this is a new style with little outcome history.

Displaying one size without uncertainty encourages overconfidence. A shopper may benefit more from seeing the garment measurements, intended ease, and the next-best option than from a bold “perfect fit” claim.

Cold start and feedback loops

A new product has no returns or retention history. The system must begin with measurements, category, material, cut, and similarity to known styles. If those source attributes are weak, the model inherits the weakness.

Feedback can also become circular. If the system recommends M more often, M receives more purchases and outcomes, giving the model more evidence about M than neighboring sizes. Stockouts create another bias: a shopper cannot select a size that was unavailable, so purchase data reflect inventory as well as preference.

Record when a recommendation was shown, which alternatives were available, whether the shopper followed it, and what happened later. Without exposure and availability data, the learning set can misread its own influence.

How a retailer should evaluate the system

Do not use “recommendation acceptance” as the only metric. A persuasive but inaccurate widget can have a high acceptance rate.

Track:

  1. size-related return and exchange rate for recommended versus non-recommended purchases;
  2. direction of error—too small or too large;
  3. coverage—the share of sessions where enough evidence exists to recommend;
  4. abstention quality—whether the system stays quiet when evidence is weak;
  5. performance by category, size range, size system, and new versus established products;
  6. customer correction rate and preference changes;
  7. outcome gaps across groups, devices, and regions;
  8. calibration—whether stated confidence matches observed accuracy.

A return reduction is not automatically a fit improvement. Check conversion, retained units, exchanges, support contacts, and customer feedback together.

Questions to ask a vendor

  • What exact event is the model predicting?
  • Which data are required, optional, stored, or shared?
  • How does the model handle a new product and a new shopper?
  • Does it use garment measurements or only behavioral similarity?
  • How are stockouts and unavailable alternatives represented?
  • Can the system abstain or show multiple plausible sizes?
  • How is performance reported by category and size range?
  • Can a retailer export recommendations, reasons, and outcomes for audit?
  • How are customer corrections and deletion requests handled?

The better promise

The credible promise is not “AI knows your body.” It is: “Based on the product information and evidence you choose to share, this is the size most likely to match your stated fit preference—and here is what remains uncertain.”

That language is less magical. It is also more useful.

Sources and verification

Reporting notes

How this story was checked

Sources
5 linked records · View list
Last verified
Reporting desk
FashionMember AI & Retail Desk
Format
Analysis
AI assistance
Used with editorial review; disclosed above.

Editorial standards · Request a correction