Where the Nine Months Go: An Audit of a Data Lake Timeline

Where the Nine Months Go: An Audit of a Data Lake Timeline

Thirty to sixty person-days of engineering, spread across 190 elapsed days. The six activities that eat the calendar, what each one costs, and which ones you can cut outright.

See how Ward detects answers that arrive months after the question

Get a demo → Take the 3-minute assessment
Contents

Audit the calendar, not the plan

Take any nine-month data platform project and count two numbers. Person-days of engineering actually performed. Elapsed working days on the calendar.

In mid-market retail the first number lands between 30 and 60. The second is around 190. The ratio is roughly one day of build for every four days of elapsed time, and in the worst cases one in six.

That gap is not laziness and it is rarely incompetence. It is six specific activities, each defensible on its own, that together consume five sixths of the schedule. Here they are, with what each one really costs and which ones you can cut.

1. Platform selection: 6 to 10 weeks

Demos with three vendors, a scoring matrix with 40 weighted criteria, reference calls, a bake-off with sample data, and a procurement negotiation.

The outcome of this process is Snowflake, BigQuery, or Databricks. All three will handle a 200-store retailer without straining. The decision is close to a coin flip weighted by what your team already knows and which cloud your ERP sits next to.

What the ten weeks actually buys is political cover. If the platform is chosen by a process, nobody owns the choice. That is a real organizational need, and it is worth being honest that you are buying cover rather than a better decision.

Compressible to two weeks. Write down your existing cloud, your team's SQL background, and your annual data volume. If two of the three point the same direction, the choice is made. Spend the other eight weeks on definitions.

2. The enterprise data model by committee: 8 to 12 weeks

Workshops with merchandising, finance, store operations, and supply chain to agree on a canonical model before ingestion starts. Entity definitions, hierarchies, conformed dimensions, the lot.

The impulse is right. Definitions are the thing that decides whether the platform produces trustworthy numbers, and nobody else can write them for you.

The sequencing is wrong. Defining 400 tables in the abstract, in a room, with no data to look at, produces a model that collides with reality the first week someone queries it. Then it gets revised anyway, having cost a quarter.

Compressible to three weeks, run in parallel. Land the raw data first. Define metrics against tables you can actually see, one metric at a time, in the order somebody needs them. The definitions come out better because they are argued against real rows instead of remembered ones.

3. Custom connector engineering: 2 to 3 weeks per source

Every source with a maintained connector takes under an hour. Every source without one takes two to three weeks: reverse-engineering an undocumented schema, negotiating read access, handling incremental extraction, and finding out that the vendor's export drops the timezone.

Mid-market retail stacks usually have one or two of these. A legacy WMS, a labor scheduling tool with no API, a franchise POS on an on-premise SQL Server nobody has patched.

Not compressible, but reorderable. The mistake is treating the hard source as a gate. Connect the four easy systems in week one, run on those, and let the WMS land in week six. A project that shows results in week two survives a hard connector in week six. A project that spends its first six weeks on the WMS shows nothing until month two and is negotiating for its life.

4. Security and vendor review: 4 to 8 weeks

SOC 2 review, DPA, network access, credential provisioning, penetration test evidence, and an internal architecture review board that meets every other Thursday.

This is legitimate work and mostly queueing time. The variable is not how long the review takes. It is when it starts.

Compressible to zero net calendar. Start it in week one, before the platform is chosen, using the shortlist. Read-only access is a materially easier review than write access, which is one of several reasons to start read-only. Nothing about the security queue is faster in month four than in month one.

See how Ward detects answers that arrive months after the question

Get a demo →

5. Historical backfill and its arguments: 2 to 6 weeks

Not the loading. Loading three years of transactions is an overnight job. The weeks go to deciding how much history, discovering that the POS changed schema in 2023, and reconciling a restated fiscal calendar.

Ask what the history is for. Demand forecasting wants two to three years. Operational monitoring wants 13 months, because it needs last year same week and nothing more. Most projects load five years because it felt safer, then spend three weeks fixing 2021 data that no model will ever read.

Compressible to days. Load 13 months to start. Add history later when a forecasting use case asks for it, by which time you know which tables it needs.

6. The reporting rebuild: 6 to 16 weeks

The biggest single line item, and the one nobody puts in the data lake budget because it is filed under BI.

Rebuilding every existing report on the new platform before anyone is allowed to use it. Two hundred reports, each needing a definition match against the old system, most of them opened fewer than a dozen times a year.

Cuttable outright. Instrument the current BI tool for 30 days and count opens per report. The distribution is brutal and consistent: a handful of reports carry nearly all the usage and the long tail is close to dead. Rebuild the top ten. Archive the rest and rebuild on request. Most requests never come.

What the compressed plan looks like

Same scope, same systems, same team, resequenced:

Week 1. Security review opens. Storage provisioned. Two easy sources connected and backfilling.

Weeks 2 to 3. Remaining API sources connected. Store and SKU dimensions built against real data. First metric reconciled.

Weeks 4 to 6. Five more metrics reconciled, in the order the business asks. Hard connector work begins in the background.

Weeks 7 to 10. Running in parallel with the incumbent process. Top ten reports rebuilt. Legacy source lands.

Ten weeks instead of nine months, with the same 30 to 60 person-days of engineering. Nothing was skipped. The work was overlapped instead of stacked, and two activities that produced no answers were deleted.

The tell that a project is buying calendar, not capability

One question exposes it: what is the earliest date on which a single business question gets answered from this platform?

If the answer is more than four weeks out, the plan is sequenced by activity rather than by outcome, and the schedule is mostly waiting. If the answer is "week two, and it will be wrong, and we will fix it in week three," the plan is sound.

Nine-month data lake projects rarely fail on technology. They fail because month eight arrives, the sponsor has changed, nobody has seen an answer yet, and the budget goes somewhere with a shorter feedback loop.

See how Ward detects answers that arrive months after the question

Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.

data lake data platform CIO project management

Not sure where AI fits in your operation? Ten questions, about three minutes. Your score out of 100 appears on screen when you finish, with no email required.

Take the 3-minute assessment

Questions about answers that arrive months after the question.

Six activities consume most of the schedule: platform selection at 6 to 10 weeks, an enterprise data model designed by committee at 8 to 12 weeks, custom connector engineering at 2 to 3 weeks per legacy source, security review at 4 to 8 weeks, historical backfill arguments at 2 to 6 weeks, and a reporting rebuild at 6 to 16 weeks. Actual build effort is 30 to 60 person-days.

Platform selection compresses to two weeks, because the answer is Snowflake, BigQuery, or Databricks and all three work. Committee modeling compresses to three weeks by defining metrics against landed data instead of in the abstract. Security review compresses to zero net calendar by starting it in week one. The reporting rebuild can usually be cut outright.

No. Instrument the current BI tool for 30 days and count opens per report. A handful carry nearly all the usage and the long tail is close to dead. Rebuild the top ten, archive the rest, and rebuild on request. Most requests never come, and report parity with a suite nobody trusts is not a milestone worth three months.

Ask for the earliest date on which one business question gets answered from the platform. More than four weeks out means the plan is sequenced by activity rather than outcome. "Week two, and it will be wrong, and we will fix it in week three" is a sound plan, because it is testable before the budget review.

From the article to the product.

How this topic maps to what Ward does, who it’s for, and the alternatives buyers benchmark against.

Your stores are generating data right now.

Ward turns it into decisions. First insight cards in 48 hours.

Read-only to start · your LLM keys · SOC 2 Type II underway · or book a call directly

Find out what your data has been hiding.

Tell us about your operation. We’ll show you the problems Ward catches, and the ones your current tools miss.

Step 1 of 3
What are your goals?
Step 2 of 3
About your operation
Step 3 of 3
Your contact info