Store Clustering: Why One National Assortment Loses in Every Store

Store Clustering: Why One National Assortment Loses in Every Store

A single national planogram is wrong everywhere because demand varies by store. Clustering on demand behavior, not region or volume, is the scalable path to localized assortment.

See how Ward detects assortment-to-demand mismatch

Get a demo → Take the 3-minute assessment
Contents

One national assortment is wrong in every store

A single national planogram is the default for most chains because it is simple to plan, simple to ship, and simple to audit. It is also wrong in every store, because demand is not national. It is local.

Two stores in the same chain, same banner, same square footage, will sell different things. One sits in a dense urban block with small households and high foot traffic. The other sits in a suburb with families buying in bulk. They get the same set, the same facings, the same space. Both are mis-served, just in opposite directions.

The cost is not abstract. McKinsey has estimated that localized assortment can lift category sales by 3 to 5 percent or more, which on a large chain is real money left on the floor every week. The national set wins on simplicity and loses on the thing that actually pays, which is matching what each store stocks to what its shoppers buy.

This post is about the middle path between one set and a thousand hand-tuned sets: clustering stores by how they actually sell, then localizing assortment per cluster.

Chains pick between two bad options

Faced with the problem, most chains land on one of two extremes, and both fail.

Option one is the single national assortment. Plan once, execute everywhere. It scales, and it is wrong everywhere, because it ignores that demand varies by store. You carry dead SKUs in stores that don't want them and miss proven sellers in stores that do.

Option two is hand-tuning every store. The category manager adjusts each store's set to local demand. In theory this is right. In practice it does not scale past a handful of stores. A chain with 400 stores and 200 categories cannot hand-tune 80,000 store-category sets every reset cycle. The labor isn't there, so it either doesn't happen or it happens badly.

The gap between these two is where the money sits. You want most of the accuracy of per-store tuning with something close to the scalability of one set. That is what clustering is for: a manageable number of groups, each localized, instead of one set or a thousand.

Clustering by region or volume is still wrong

Most chains that do cluster, cluster on the wrong variable. The two common choices, region and volume, both miss what matters.

Region clustering groups stores by geography: all the Northeast stores, all the Southeast stores. It feels intuitive, and it is mostly wrong. Two stores ten miles apart in the same metro can serve completely different shoppers, one urban and one suburban, while two stores in different states can sell nearly identically. Geography is a weak proxy for demand.

Volume clustering groups stores by total sales: A stores, B stores, C stores. That measures how big a store is. It says nothing about what sells inside it. A high-volume urban store and a high-volume suburban store land in the same A cluster and get the same set, even though their shoppers want different products. Volume sizes the store. It does not describe demand.

Both methods cluster on attributes that are easy to pull from a header record and have little to do with what moves off the shelf. They produce clusters that look tidy on a map or a ranking and perform like the national set, because they never looked at sales behavior.

Cluster on sales behavior

The variable that matters is demand behavior: what actually sells in each store and in what mix. Two stores belong in the same cluster when their shoppers buy the same way, regardless of where they are or how big they are.

This means clustering on the sales data itself. Category mix, SKU-level velocity patterns, the share of premium versus value, the local sellers that over-index. A store where organic and premium over-index belongs with other premium-skewing stores even if one is urban and one is suburban. A value-skewing store belongs with the stores that shop like it, wherever they happen to sit on the map.

Demographics, climate, urban versus suburban, and income are useful inputs, but only because they drive behavior. The behavior is what you cluster on. When you group stores by how they sell rather than by where they sit, the clusters start to predict demand instead of describing a map.

See how Ward detects assortment-to-demand mismatch

Get a demo →

The dead SKU and the missing SKU are the same problem

Once you cluster on behavior, the cost of the national set shows up as two mirror-image errors in every store.

The first is the dead SKU. A product sits on the shelf in a store where nobody buys it. It got the slot because the national planogram said so. It takes facings, ties up inventory, gets stocked and faced and rotated by the crew, and earns almost nothing. Every store carries a tail of these, products that sell fine chain-wide but are dead in this particular store.

The second is the missing SKU. A product sells well in stores just like this one, behaviorally, but it isn't in this store's set because the national plan didn't make room. The demand is real and proven by the store's cluster peers. You are simply not offering it. That is lost sales with no signal, because you can't sell what you don't stock.

These two errors are the same mistake viewed from both ends. The national set put a SKU where demand isn't and left out a SKU where demand is. In every store, you are carrying dead weight and missing proven sellers at the same time. Clustering finds both by comparing each store to its true behavioral peers.

The dead SKU is the easier one to spot, because it sits there with sales near zero. The missing SKU is the expensive one, because it leaves no trace. There is no slow-moving line to flag, no backstock to count, just a sale that never happened in a store that would have made it. The only way to surface a missing SKU is to compare the store against peers that sell the same way and carry it. Without the peer comparison, the gap is invisible.

Allocate space by local velocity

Localized assortment is not only which SKUs to carry. It is how much space each one gets, and the national set gets this wrong too.

Facings on the national planogram are set to chain-average velocity. In a given store, the fast local movers are under-faced and run empty between deliveries, while the national-average performers that are slow locally hold more space than their sales justify. The shelf is balanced for an average store that does not exist.

Cluster-level space allocation fixes the balance. Give the local fast movers the facings their velocity earns so they stay in stock, and pull space from the SKUs that are slow in this cluster. Same shelf, same total space, better matched to what sells. A good chunk of the assortment lift comes from here: sizing the SKUs you keep, which most teams skip entirely in favor of adding and cutting.

Too many clusters costs more than it returns

Clustering has a failure mode in the other direction: too many clusters. Every cluster you add is real operational cost, and past a point the cost outruns the benefit.

Each cluster is a separate planogram to build, ship, reset, and audit. It is a separate set of case packs and distribution lanes to manage. Twelve clusters is twelve times the planning and reset labor of one set, and the accuracy gain from cluster eleven to cluster twelve is small. You are paying linearly for a benefit that flattens out.

The right number is the smallest set of clusters that captures most of the demand variation. For most chains that lands somewhere between five and nine. Each cluster should be large enough to justify its own set and behaviorally distinct enough to need one. A cluster of three stores rarely earns its planning cost.

The discipline is to stop adding clusters when the marginal one stops paying. That requires measuring what each cluster actually returns against what it costs to run, which most chains never do. They either under-cluster with region and volume or over-cluster into a planning burden nobody can sustain.

How Ward surfaces this without running your assortment

Ward is a read-only observability platform for multi-store retailers. We do not build your planograms or set your assortment. We watch the POS, ERP, and inventory data you already generate and tell you where a store's set has drifted from its real local demand.

The model is detect, decide, execute, audit. Ward detects where a store diverges from its behavioral peers: the dead SKUs it carries that move nowhere locally, and the proven sellers it is missing that perform well in stores that sell just like it. You decide what to add, cut, or re-space, because your category team owns the set. Your team executes the change. Then Ward audits the cluster move, checking whether the added SKUs actually sold and the cut SKUs were actually dead.

SKUs, clusters, and planograms only move when a person moves them. Ward's part is the early warning: this store's assortment has drifted away from what its shoppers actually buy, and here is by how much.

You get insight cards. A card might say that a store carries fourteen SKUs in a category that are dead locally while missing six SKUs that over-index across its behavioral cluster, and that two local fast movers are under-faced against their velocity. It might also flag that a proposed cluster split adds reset cost without distinct enough demand to justify it. A category manager can act on that this cycle without interpreting a model first.

Key takeaways

  • One national assortment is wrong in every store, because demand forms locally. McKinsey has put the lift from localized assortment at 3 to 5 percent or more in category sales, which is real money the national set leaves on the floor.
  • The two default options both fail. A single national set scales but is wrong everywhere, and hand-tuning every store is right in theory but does not scale past a handful of stores.
  • Region and volume clustering miss what matters. Geography is a weak proxy for demand and volume only sizes the store, so both produce clusters that perform like the national set.
  • Cluster on sales behavior instead. Group stores by how their shoppers actually buy, using category mix and SKU velocity, so the clusters predict demand rather than describe a map.
  • The dead SKU and the missing SKU are the same error. Every store carries products that move nowhere locally while missing proven sellers from its behavioral peers, dead weight and lost sales at the same time.
  • Space allocation matters as much as SKU selection. National-average facings under-face local fast movers and over-face local slow movers, so re-spacing by cluster velocity recovers a good share of the lift.
  • Too many clusters costs more than it returns. Each cluster is separate planning, reset, and distribution labor, so the right number is the smallest set that captures most of the demand variation, usually a handful. Ward detects the drift and audits the move, read-only, lane assist not autopilot.

See how Ward detects assortment-to-demand mismatch

Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.

store clustering assortment planning localization merchandising

Not sure where AI fits in your operation? Ten questions, about three minutes. Your score out of 100 appears on screen when you finish, with no email required.

Take the 3-minute assessment

Your stores are generating data right now.

Ward turns it into decisions. First insight cards in 48 hours.

Read-only to start · your LLM keys · SOC 2 Type II underway · or book a call directly

Find out what your data has been hiding.

Tell us about your operation. We’ll show you the problems Ward catches, and the ones your current tools miss.

Step 1 of 3
What are your goals?
Step 2 of 3
About your operation
Step 3 of 3
Your contact info