Supplier OTIF: The Scorecard That Pays for Itself in Chargebacks
OTIF is binary: on time and in full or it fails. Get the formula, the four settings that decide the score, benchmark bands, and the safety stock each one costs you.
See how Ward detects supplier OTIF misses
Get a demo → Take the 3-minute assessmentContents
- OTIF is the supplier score most mid-market retailers never keep
- The "and" in on-time and in-full is the whole catch
- The gap between fill rate and OTIF
- The downstream cost of one late or short delivery
- How to calculate OTIF, and the four settings that decide the score
- OTIF benchmarks and what each band costs you
- The chargeback opportunity big retailers already use
- The scorecard changes the vendor negotiation
- How Ward surfaces this from your PO and receiving data
- Key takeaways
OTIF is the supplier score most mid-market retailers never keep
On-time in-full measures one thing: did the supplier deliver the right quantity on the agreed date. It is the single cleanest read on whether a vendor is reliable, and it is the score most mid-market retailers never bother to keep.
The reason is not that operators don't care. It is that the data needed to score OTIF is messy. The promised date lives in the PO, the actual arrival lives in receiving, the quantities live in both and rarely match cleanly, and the two systems were never built to talk. Scoring a supplier means joining records nobody owns the seam between.
So the miss goes unscored. A late, short delivery causes a stockout, an expedite, and a deeper safety-stock buffer, and the chain absorbs all of it as the cost of doing business. The supplier never hears about it, never improves, and never pays for it.
This post covers how to define and measure OTIF, what a single bad delivery actually costs downstream, how fill rate differs from OTIF, the chargeback opportunity big retailers already use, and how the scorecard changes a vendor negotiation. The cost you are eating is larger than you think.
The "and" in on-time and in-full is the whole catch
OTIF is an AND, not an OR. The delivery has to be on time AND in full to count. A shipment that arrives on the right day but short does not pass. A complete shipment that arrives two days late does not pass. Both halves have to be true on the same delivery.
That AND is what makes OTIF honest and what makes most suppliers score worse than they claim. A vendor will tell you they hit 95 percent on time and 97 percent in full and call it a strong record. Multiply the two and the combined OTIF can land in the low 90s or worse, because the misses do not always overlap. The on-time failures and the in-full failures stack.
You also have to define each half precisely or the score is meaningless. On-time needs a delivery window, not a single day, and you have to decide whether early counts as a miss, because early delivery ties up receiving labor and warehouse space you planned for later. In-full needs a tolerance, because a delivery one unit short on a thousand-unit order is not the same failure as one half short.
Set the definitions once, apply them to every supplier the same way, and the score becomes comparable across your whole vendor base. That comparability is the entire value. An OTIF number you measure differently for each vendor tells you nothing.
The gap between fill rate and OTIF
Fill rate and OTIF get used interchangeably, and they are not the same. Fill rate asks what percentage of the ordered quantity arrived. OTIF asks whether the whole order arrived complete and on time. A supplier can post a high fill rate and a poor OTIF at the same time.
Here is how. A vendor ships 98 percent of every order on average, which sounds excellent as a fill rate. But that 2 percent short is spread across nearly every delivery, so almost no delivery is truly in full. Measured as OTIF, where the order either passes whole or fails, that same vendor can score in the 60s or 70s. Fill rate averaged away the misses that OTIF counts one by one.
The distinction matters because the downstream cost tracks OTIF, not fill rate. A delivery that is 98 percent complete still leaves a gap on the shelf for the missing 2 percent, and if that 2 percent is a top seller, the stockout cost is real regardless of how good the average looks. OTIF is closer to how the failure actually hits the store.
Use fill rate to understand magnitude and OTIF to understand reliability. The two together tell you whether a vendor is occasionally short by a little or chronically short on the things that matter.
There is a timing wrinkle too. Fill rate is usually measured per line over a period, so a vendor that splits one order into two shipments can still post a clean fill rate once both arrive. OTIF judges the order on the agreed date, so the second shipment that closes the gap a week late is still a miss. That is the right way to score it, because the store felt the gap during the week the shelf was empty, not on the day the backfill finally landed.
See how Ward detects supplier OTIF misses
Get a demo →The downstream cost of one late or short delivery
Trace a single late, short delivery through the chain and the cost stacks fast.
It starts with the stockout. The shelf goes empty on the SKUs that did not arrive, and the lost sales begin immediately. The shopper who wanted that item buys a substitute, skips it, or buys it elsewhere. Two of those three cost you the sale, and the third teaches the shopper not to rely on you for it.
Then comes the expedite. Someone notices the gap and places an emergency reorder, often at a worse cost, sometimes with premium freight to get it in fast. The labor to chase the order, escalate to the vendor, and reschedule the receiving is real and it always falls on your team, not the supplier's.
The quiet, permanent cost is inflated safety stock. Once a buyer learns a supplier is unreliable, the buyer compensates by carrying more buffer inventory to cover the misses. That buffer is cash on the shelf for every SKU from that vendor, every week, forever. One unreliable supplier raises the working capital you carry across their entire line, and that cost dwarfs any single stockout.
Add it up and a chronically poor OTIF supplier costs you lost sales, expedite premiums, escalation labor, and a permanent inventory buffer. None of it shows up as a line item called "supplier failure," so none of it gets charged back, and the vendor has no reason to fix anything.
How to calculate OTIF, and the four settings that decide the score
OTIF is orders delivered both on time and in full divided by total orders, times 100. Binary per order: it either passes whole or it fails. There is no partial credit, which is the entire point.
Four configuration choices determine what score a given supplier gets, and they need to be written into the vendor agreement before anyone reports a number.
The delivery window. On time against what? A single agreed date, a two-day window, or a same-week window. Most mid-market programs use a window of zero to two days late and count early deliveries beyond three days as misses too, because early receipts consume dock labor and backroom space you did not plan for.
The in-full tolerance. Exactly 100%, or 98% and above counts as full. A tolerance band makes the score less brutal and less useful. Start at 100% for A-tier SKUs and allow a band on the tail if you need vendor buy-in to launch the program.
The unit of measure. Order-level, line-level, or case-level. Order-level is harshest and closest to how the store experiences it. Line-level is the common compromise and the one most suppliers already report internally.
Who owns the exceptions. Weather, carrier failure, and retailer-caused rescheduling all need a documented handling rule up front, or every monthly review becomes an argument about which misses count.
Write these four down and both sides are scoring the same game. Skip them and the vendor arrives at the review with a different number than yours, which is how most first-year OTIF programs stall.
OTIF benchmarks and what each band costs you
OTIF scores look alarming the first time a mid-market retailer measures them, largely because most suppliers have never been scored this way on the account.
| OTIF band | Read | Typical safety stock premium | Action |
|---|---|---|---|
| 95%+ | Best in class | None | Reduce buffer, shorten review cycle |
| 90 to 94% | Acceptable | 5 to 10% extra cover | Score monthly, share the scorecard |
| 85 to 89% | Typical mid-market reality | 10 to 20% extra cover | Formal improvement plan with dates |
| 75 to 84% | Costing real money | 20 to 35% extra cover | Chargebacks or dual-source the A-tier SKUs |
| Below 75% | Structural failure | 35%+ extra cover | Replace, or accept the buffer as a cost of goods |
The safety stock column is the number to bring to the negotiation. It converts a service score into a working capital figure, which is the only version of this conversation a CFO will act on. A supplier at 82% OTIF across a $6M annual line is quietly forcing roughly $300,000 to $500,000 of extra inventory onto your balance sheet, permanently.
That framing also decides whether to fix or replace. A vendor at 88% with a real improvement plan is worth the patience. A vendor at 78% with no plan is charging you a hidden premium larger than most price concessions you could negotiate. Their fill rate impact downstream is where the rest of the cost lands.
The chargeback opportunity big retailers already use
Large retailers do not eat OTIF failures. They charge them back. Walmart's OTIF program fines suppliers a percentage of the value of goods that arrive late or short, and the threshold has been set high enough that even strong vendors pay when they miss. Target, Kroger, and other big chains run their own versions. The message to the supplier is direct: your reliability is your cost, not ours.
Mid-market retailers almost never do this, and not because the contracts forbid it. It is because they cannot prove the miss. A chargeback requires a defensible record: this PO promised this quantity on this date, this is what receiving logged, here is the gap. Without that record joined cleanly, there is nothing to invoice the vendor for, so the cost stays with the retailer.
The opportunity is not really the chargeback dollars, though those are real. It is the behavior change. A supplier that pays for misses fixes its misses. The moment a vendor knows you are scoring OTIF and willing to charge it back, their service to your account improves, because now their unreliability costs them instead of you.
You do not need Walmart's volume to use the lever. You need the record. Once you can show a supplier their own OTIF score with the dollar cost attached, the conversation changes whether or not you ever issue a formal chargeback.
The scorecard changes the vendor negotiation
An OTIF scorecard moves a vendor negotiation from opinion to evidence. Without it, the annual review is the buyer saying service has been spotty and the vendor saying they have done their best. Nobody can prove anything, so nothing changes and the conversation turns to price.
With the scorecard, you walk in with the number. This supplier hit 71 percent OTIF over the last two quarters, here are the misses by SKU and date, and here is the downstream cost they pushed onto us. That reframes the entire discussion. Now the vendor is explaining a documented performance gap, not debating a vibe.
It also gives you a fair trade to offer. Better OTIF earns better terms, more shelf, or a bigger share of the category. Worse OTIF earns a chargeback, a tighter window, or a second source. The scorecard turns reliability into a currency both sides can negotiate with, instead of a complaint that goes nowhere.
How Ward surfaces this from your PO and receiving data
Ward is a read-only observability platform for multi-store retailers. We do not place your orders, manage your vendors, or issue your chargebacks. We watch the PO and receiving data you already generate and make the supplier scorecard that your systems were never built to produce.
The model is detect, decide, execute, audit. Ward detects OTIF misses by supplier, DC, and SKU by joining what the PO promised against what receiving logged, the join nobody owns by hand. It quantifies the downstream cost: the stockout the miss caused, the expedite it triggered, the safety stock it justifies. You decide whether to have the conversation, tighten the terms, or charge it back, because you own the vendor relationship. Your team executes. Then Ward audits whether that supplier's OTIF actually improved after the conversation or the chargeback.
Ward will not invoice a vendor and will not cancel a contract. It builds the defensible record, attaches the dollar cost, and hands both to the buyer who has to sit across the table from that supplier.
And it is not another scorecard report to assemble. You get insight cards. A card might say a specific supplier ran 68 percent OTIF to two of your DCs last quarter, driven by chronic short-shipments on three top-velocity SKUs, with an estimated downstream cost in lost sales and expedite freight, and a chargeback amount your contract supports. That is something a buyer can take into the next vendor call, not a spreadsheet to build first.
The point is to turn a cost you have been eating into a number you can act on. OTIF is the cleanest read on supplier reliability, and most chains never see it because the data is siloed. Ward joins it and puts it in front of you with the cost attached.
Key takeaways
- OTIF is the cleanest read on supplier reliability, and most mid-market chains never keep it. The PO promise and the receiving record live in separate systems that were never built to talk, so the miss goes unscored and the cost gets eaten.
- The "and" is the catch. A delivery must be on time AND in full to pass, so a vendor claiming 95 percent on time and 97 percent in full can post a combined OTIF in the low 90s or worse once the misses stack.
- Fill rate is not OTIF. A supplier short 2 percent on nearly every order posts a great fill rate and a poor OTIF, because OTIF counts each incomplete order as a failure instead of averaging the shortfall away.
- One bad delivery costs four ways. Lost sales from the stockout, expedite premiums on the emergency reorder, escalation labor, and a permanent safety-stock buffer that raises working capital across the vendor's entire line.
- Four settings decide the score, so agree them first. The delivery window, the in-full tolerance, whether you measure at order, line, or case level, and who owns weather and carrier exceptions. Skip these and the vendor arrives at the review with a different number than yours.
- Benchmark bands: 95%+ is best in class, 85 to 89% is typical mid-market, below 75% is structural failure. Each band carries a safety stock premium, from none at the top to 35%+ extra cover at the bottom.
- Convert the score into working capital before the negotiation. A supplier at 82% OTIF on a $6M line quietly forces roughly $300,000 to $500,000 of extra inventory onto your balance sheet, permanently, which is larger than most price concessions you could win.
- Big retailers charge the misses back. Walmart, Target, and Kroger fine suppliers for late and short deliveries, while mid-market chains eat the cost because they cannot prove the miss with a defensible record.
- The scorecard moves the negotiation from opinion to evidence. Walking in with a documented OTIF number and downstream cost reframes the vendor review and turns reliability into a currency both sides can trade on.
- The record is the lever. Ward joins PO against receiving to score OTIF by supplier, DC, and SKU, quantifies the downstream cost, and audits whether performance improved after the conversation, read-only, lane assist not autopilot.
See how Ward detects supplier OTIF misses
Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.