Where the Nine Months Go: An Audit of a Data Lake Timeline
Thirty to sixty person-days of engineering, spread across 190 elapsed days. The six activities that eat the calendar, what each one costs, and which ones you can cut outright.
See how Ward detects answers that arrive months after the question
Get a demo → Take the 3-minute assessmentContents
- Audit the calendar, not the plan
- 1. Platform selection: 6 to 10 weeks
- 2. The enterprise data model by committee: 8 to 12 weeks
- 3. Custom connector engineering: 2 to 3 weeks per source
- 4. Security and vendor review: 4 to 8 weeks
- 5. Historical backfill and its arguments: 2 to 6 weeks
- 6. The reporting rebuild: 6 to 16 weeks
- What the compressed plan looks like
- The tell that a project is buying calendar, not capability
Audit the calendar, not the plan
Take any nine-month data platform project and count two numbers. Person-days of engineering actually performed. Elapsed working days on the calendar.
In mid-market retail the first number lands between 30 and 60. The second is around 190. The ratio is roughly one day of build for every four days of elapsed time, and in the worst cases one in six.
That gap is not laziness and it is rarely incompetence. It is six specific activities, each defensible on its own, that together consume five sixths of the schedule. Here they are, with what each one really costs and which ones you can cut.
1. Platform selection: 6 to 10 weeks
Demos with three vendors, a scoring matrix with 40 weighted criteria, reference calls, a bake-off with sample data, and a procurement negotiation.
The outcome of this process is Snowflake, BigQuery, or Databricks. All three will handle a 200-store retailer without straining. The decision is close to a coin flip weighted by what your team already knows and which cloud your ERP sits next to.
What the ten weeks actually buys is political cover. If the platform is chosen by a process, nobody owns the choice. That is a real organizational need, and it is worth being honest that you are buying cover rather than a better decision.
Compressible to two weeks. Write down your existing cloud, your team's SQL background, and your annual data volume. If two of the three point the same direction, the choice is made. Spend the other eight weeks on definitions.
2. The enterprise data model by committee: 8 to 12 weeks
Workshops with merchandising, finance, store operations, and supply chain to agree on a canonical model before ingestion starts. Entity definitions, hierarchies, conformed dimensions, the lot.
The impulse is right. Definitions are the thing that decides whether the platform produces trustworthy numbers, and nobody else can write them for you.
The sequencing is wrong. Defining 400 tables in the abstract, in a room, with no data to look at, produces a model that collides with reality the first week someone queries it. Then it gets revised anyway, having cost a quarter.
Compressible to three weeks, run in parallel. Land the raw data first. Define metrics against tables you can actually see, one metric at a time, in the order somebody needs them. The definitions come out better because they are argued against real rows instead of remembered ones.
3. Custom connector engineering: 2 to 3 weeks per source
Every source with a maintained connector takes under an hour. Every source without one takes two to three weeks: reverse-engineering an undocumented schema, negotiating read access, handling incremental extraction, and finding out that the vendor's export drops the timezone.
Mid-market retail stacks usually have one or two of these. A legacy WMS, a labor scheduling tool with no API, a franchise POS on an on-premise SQL Server nobody has patched.
Not compressible, but reorderable. The mistake is treating the hard source as a gate. Connect the four easy systems in week one, run on those, and let the WMS land in week six. A project that shows results in week two survives a hard connector in week six. A project that spends its first six weeks on the WMS shows nothing until month two and is negotiating for its life.
4. Security and vendor review: 4 to 8 weeks
SOC 2 review, DPA, network access, credential provisioning, penetration test evidence, and an internal architecture review board that meets every other Thursday.
This is legitimate work and mostly queueing time. The variable is not how long the review takes. It is when it starts.
Compressible to zero net calendar. Start it in week one, before the platform is chosen, using the shortlist. Read-only access is a materially easier review than write access, which is one of several reasons to start read-only. Nothing about the security queue is faster in month four than in month one.
See how Ward detects answers that arrive months after the question
Get a demo →5. Historical backfill and its arguments: 2 to 6 weeks
Not the loading. Loading three years of transactions is an overnight job. The weeks go to deciding how much history, discovering that the POS changed schema in 2023, and reconciling a restated fiscal calendar.
Ask what the history is for. Demand forecasting wants two to three years. Operational monitoring wants 13 months, because it needs last year same week and nothing more. Most projects load five years because it felt safer, then spend three weeks fixing 2021 data that no model will ever read.
Compressible to days. Load 13 months to start. Add history later when a forecasting use case asks for it, by which time you know which tables it needs.
6. The reporting rebuild: 6 to 16 weeks
The biggest single line item, and the one nobody puts in the data lake budget because it is filed under BI.
Rebuilding every existing report on the new platform before anyone is allowed to use it. Two hundred reports, each needing a definition match against the old system, most of them opened fewer than a dozen times a year.
Cuttable outright. Instrument the current BI tool for 30 days and count opens per report. The distribution is brutal and consistent: a handful of reports carry nearly all the usage and the long tail is close to dead. Rebuild the top ten. Archive the rest and rebuild on request. Most requests never come.
What the compressed plan looks like
Same scope, same systems, same team, resequenced:
Week 1. Security review opens. Storage provisioned. Two easy sources connected and backfilling.
Weeks 2 to 3. Remaining API sources connected. Store and SKU dimensions built against real data. First metric reconciled.
Weeks 4 to 6. Five more metrics reconciled, in the order the business asks. Hard connector work begins in the background.
Weeks 7 to 10. Running in parallel with the incumbent process. Top ten reports rebuilt. Legacy source lands.
Ten weeks instead of nine months, with the same 30 to 60 person-days of engineering. Nothing was skipped. The work was overlapped instead of stacked, and two activities that produced no answers were deleted.
The tell that a project is buying calendar, not capability
One question exposes it: what is the earliest date on which a single business question gets answered from this platform?
If the answer is more than four weeks out, the plan is sequenced by activity rather than by outcome, and the schedule is mostly waiting. If the answer is "week two, and it will be wrong, and we will fix it in week three," the plan is sound.
Nine-month data lake projects rarely fail on technology. They fail because month eight arrives, the sponsor has changed, nobody has seen an answer yet, and the budget goes somewhere with a shorter feedback loop.
See how Ward detects answers that arrive months after the question
Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.