2026-05-30 · AI Assisted Analysis

The Path to 97%: Engineering Data Platform Availability as a KPI

Your warehouse vendor sells you four nines. Your data products — the tables and metrics people actually consume — are probably running somewhere between 50% and 80%. That gap is not a vendor problem. It is a measurement problem and a design problem, and both are fixable. Here is the taxonomy that tells you who owns each fix, and the weighted metric that turns 97% from a wish into a roadmap.

$12.9M
Average annual cost of poor data quality per organization (Gartner)
41% → 57%
Practitioners citing poor data quality as a top problem, 2022 to 2024 — the gap is widening (dbt Labs, State of Analytics Engineering 2024)

For data & SLT leaders — the one-screen version

Decision requiredDefine your Tier-0 / Tier-1 data products and appoint a named owner for each.
The KPIValue-weighted availability by data product — not warehouse uptime.
First 30 daysBaseline weighted availability, tag every incident by source and cost, and name your top five failure modes.
Business riskPoor data quality costs the average org ~$12.9M/year (Gartner) and is a major blocker to scaling GenAI (Forrester) — yet SLA metrics are a key measure for only 5% of organizations (dbt Labs 2024).
The outcomeAn availability roadmap you can fund and sequence — not vendor uptime theatre.

The honest starting line: you are not at four nines

Open your cloud warehouse's status page and you will see a comfortable number. some BigQuery editions advertise up to a 99.99% uptime SLA — about 52 minutes of allowed downtime per year (the Standard edition is 99.9%). It is tempting to inherit that number and report it upward as your platform's availability.

It is also wrong. The warehouse being up is necessary but nowhere near sufficient. What your stakeholders consume is a data product: a fact table that is fresh, a metric that reconciles, a dashboard that loaded with today's numbers. When a dbt run fails at 3am, when an upstream team renames a column, when a vendor feed arrives empty — the warehouse can still be within SLA and your data product is at zero.

Measured honestly, most established data platforms sit between 50% and 80% availability. That band is not incompetence; it is the natural resting state of a platform that has incidents but does not yet manage them as a system. At that level you are absorbing at least one incident per week on average. And because most dbt projects are still effectively monolithic — one big DAG, one run, one failure domain — a single break can blank a large share of what downstream consumers depend on.

Author's field estimate. The 50–80% band, the ~one-incident-a-week cadence, and the incident-source split later in this piece are what I have seen on unmanaged platforms — not a published benchmark. The point is the shape of the problem; instrument your own platform and your real numbers will differ.

The industry data backs the felt experience. Monte Carlo and Wakefield Research found that data downtime nearly doubled year over year, with monthly incidents rising from roughly 59 to 67, detection taking four hours or more for 68% of teams, and resolution averaging around 15 hours per incident. Across 11 million-plus monitored tables, Monte Carlo observes roughly one data quality issue per ten tables every year. This is not a tail risk. It is the weather.

Here is the tell for any leadership team. The same dbt Labs survey that found 57% of practitioners naming poor data quality a top problem also found that SLA metrics are a key metric for only 5% of organizations. Almost everyone says data quality matters; almost no one runs it as a managed KPI. That gap — between concern and measurement — is the entire opportunity.

So before you set a target, set the frame. 97% is not "almost perfect." It is a specific, costed engineering goal — and you cannot manage it until you can see what each level actually buys you.

Figure 1: The availability ladder. Monthly downtime budget at each availability level, computed as (1 − availability) × 730 hours — standard SRE error-budget arithmetic (so 99.9% ≈ 44 min/month, 99% ≈ 7.3 hrs/month). 97% still permits about 22 hours of downtime per month, but getting there from a typical 50–80% platform is a 5–15x reduction. dbt Labs, "Data SLAs: Best Practices" (2024) frames data SLOs with the same SRE-style error budgets.

A taxonomy of where incidents come from

"Improve data quality" is not a plan, because incidents do not all have the same owner or the same fix. The first move toward 97% is to tag every incident by its source. In my experience running data platforms, almost every incident falls into one of four ownership buckets.

External telemetry points in the same direction, but with an important caveat. In Monte Carlo's 2025 analysis of 11 million-plus monitored tables, detected data-quality events were spread across pipeline execution faults (26.2%), real-world variation (20%), ingestion disruptions (16.6%), platform instability (15.2%), intentional changes such as backfills (14.2%), and schema drift (7.8%). That taxonomy is detection-oriented, not ownership-oriented: pipeline execution can mean missed schedules, broken dependencies, permissions, compute limits, or orchestration configuration — not simply "Airflow is unreliable" — while real-world variation and intentional changes may be expected data movement rather than true incidents. For remediation, you still need the ownership view below.

a. Your own releases — the easiest owned source to fix

A change your data team shipped broke something: a model refactor, a bad backfill, a logic regression that passed review because no one could see its downstream blast radius. This is the best category to be in, because you own all of it. The remediation roadmap is well understood and entirely within your control: CI/CD on the dbt project, data-diff and impact analysis on every pull request, a validation gate, and — since dbt 1.8 shipped native unit testing in May 2024 — actual unit tests on your SQL logic. The same unit-testing machinery doubles as acceptance testing: encode the business rule as a test, and a release that violates it never merges.

b. Infra and vendor issues — solved through relationships

A SaaS source went down, a Fivetran connector silently stalled, a warehouse region degraded. You do not own the fix, so engineering rigor alone will not move it. This category is solved through relationships and alignment: making the business value of the feed explicit to the vendor, negotiating real SLAs with consequences, and pushing toward self-serve enablement so a degraded source surfaces fast and reroutes rather than failing silently.

c. Upstream data changes — the hard boundary

This is the one that keeps data teams reactive. A software engineer changes a database schema or — worse — the semantics of a field, with no idea that three dashboards and a revenue model depend on it. You cannot fix this source by working harder inside your own walls, because the cause lives in another team's codebase and roadmap.

The structural answer is data contracts — explicit, version-controlled agreements between producers and consumers — backed by explicit data measurement that is co-owned by the data team and the engineering team (and the data stewards, if your org has them). The honest caveat: as Tobiko has argued, dbt contracts validate structure — column names and types — not logic or meaning, so a semantic change can still slip through. Which is exactly why this layer is the one that most needs to become self-serve, and the one where AI changes the economics. More on that below.

d. Pipeline and data-product design — the multiplier

The fourth source is different in kind. It is not a separate slice of incidents; it is the design decision that determines how badly any of a, b, or c hurts. A monolithic pipeline converts a single broken model into a total platform outage. A decomposed one contains the same failure to a single product. The easiest way to raise availability is often not to have fewer incidents — it is to make each incident matter less.

Figure 2: The four-source taxonomy. The three classic ownership sources and their owners; pipeline design (d) is the cross-cutting multiplier. Practitioner framework — author field estimate, illustrative. Monte Carlo's telemetry shows a different, detection-oriented root-cause breakdown; this chart is a remediation lens, not a measured industry distribution.
SourceOwnerPrimary remediationDifficulty
a. Your own releasesData teamCI/CD, data-diff, dbt unit + acceptance testsLowest — full ownership
b. Infra / vendorVendor + platformRelationships, value alignment, self-serve enablementMedium — influence, not control
c. Upstream data changesCo-owned: data + engineeringData contracts + explicit, AI-assisted measurementHighest — cause lives in another team
d. Pipeline / product designData platformDecompose the monolith into independent data productsStructural — the multiplier on a/b/c

The unlock: weighted availability across data products

Here is the measurement reframe that makes the whole thing tractable. If you score availability as a single platform-wide flag, every incident is binary: the platform is down, the day is a zero. Under that accounting a team that runs flawlessly for 29 days and has one contained incident reports the same as a team in permanent crisis. The metric punishes you for honesty and tells you nothing about where to invest.

Instead, weight availability across your data assets. If you have five data products and an incident takes one of them down for a day, four-fifths of the value still flowed. That day was 80% available, not 0%. Aggregate weighted availability across the period and you have a number that actually moves when you do the right things — and barely moves when an incident lands somewhere that does not matter.

Figure 3: The same incident, scored two ways. Monolithic accounting marks the whole platform down (0%); weighted accounting recognises that four of five products kept delivering (80%). Practitioner framework — illustrative. Weighting by business value, not a flat count, is the more defensible version: a Tier-0 revenue table should carry more weight than an experimental sandbox model.

This is not just an accounting trick to make the dashboard greener. It is a design directive in disguise. The only way to genuinely raise weighted availability is to ensure failures stay contained — which means decomposing the monolith into independent data products that each deliver value and fail independently. This is the through-line from Zhamak Dehghani's data mesh to Monzo's "data as products" refactor: ownership and blast-radius go together.

Practitioners are already measuring this way. Mikkel Dengsøe describes measuring the percentage of models meeting their SLA and segmenting by criticality to expose where the real variation lives. And Petr Janda makes the complementary point that availability should be built on human-declared incidents, not raw alerts — otherwise noisy tests put you in a permanent false "down" state. Weighted availability needs both: weight by value, count only real incidents.

The metric in one line

Weighted availability = the value-weighted share of your data products that met their freshness and correctness SLA over the period. It is the single number that makes the path to 97% improvable — because it rewards containment and honest incident counting, not heroics.

Coming soon: a hands-on guide to instrumenting this

This piece is the why and the what. The how — defining per-product freshness and correctness SLOs, computing the value-weighted metric, and wiring incident attribution and cost into your warehouse and observability stack — is enough of its own subject that we're publishing it separately. A practical "how to measure data platform availability" walkthrough is next in this series.

Putting a number on the incident

Weighted availability tells you what share of your data products were unavailable. It does not tell you what that unavailability cost — and the cost is what decides whether a fix is a data-team ticket or a company project. Every incident drains value through one of two channels, and they differ by an order of magnitude.

Productivity loss — people couldn't use the product

The common, diffuse one. An executive dashboard, a finance dataset, a self-serve table goes stale, and a few dozen analysts and decision-makers are blocked — or worse, quietly working from numbers they have stopped trusting. The cost is real but rarely lands on a P&L: headcount blocked × hours × fully-loaded cost × the fraction of their work that depends on that asset. Because it never shows up as a line item, this channel is chronically under-prioritised.

Direct operational loss — the product stopped moving money

The one that gets a CFO's attention. Here the data product isn't read by a human; it feeds an automated system that spends or makes money — or carries a hard deadline. Your model stops scoring, so the bidding service can't update Google Ads bids and the campaign keeps spending at default efficiency, or fails to deploy budget it should have. A stale feature table mis-prices checkout. A broken reverse-ETL sync stops pushing leads to sales. A late regulatory filing or a board pack built on numbers that don't reconcile carries a cost that has nothing to do with headcount. The loss is direct and often an order of magnitude larger — value flowing through the product per hour × duration × the efficiency or revenue you forfeit — and it accrues whether or not anyone notices the dashboard is stale.

A back-of-envelope incident cost

Incident cost ≈ Σ (across affected data products) of productivity loss + direct operational loss over the outage window — where productivity loss = people blocked × hours × loaded hourly cost × dependency fraction, and direct loss = value flowing through the product per hour × hours × efficiency forfeited. Weight both by decision criticality: a trading signal, a regulatory filing, or a board-pack metric costs far more per hour stale than an internal exploration table. You will not get this to two significant figures. You do not need to — an order of magnitude is enough to prioritise.

Figure 4: The same incident, very different bills. A 15-hour outage costs several times more when it lands on a product that moves money than when it blocks an internal dashboard. Illustrative cost model — author's framework. The input assumptions (people blocked, loaded cost, daily ad spend, efficiency drop) are worked examples, not measured values; the 15-hour duration reflects the ~15-hour average resolution time reported by Monte Carlo / Wakefield (2024). For scale: New Relic (2024) puts high-impact outages above $1M/hour, and Splunk / Oxford Economics (2024) estimates Global 2000 downtime at $400B/year.

That gap is the whole argument for weighting availability by value, not by a flat count: the cost-per-hour of an outage is the weight. And it changes what kind of initiative the fix becomes. A dashboard that costs ~$22k per incident in lost analyst time is a backlog item. A model feed with a ~$95k-per-incident exposure is a cross-functional project — it justifies a data contract with the producing team, an on-call rotation, and sometimes an org change: a data-platform team with funded SLAs, or a steward embedded alongside the engineers who own the upstream schema. The cost number is what unlocks the budget and the mandate.

The roadmap to 97%: sequence the work by source

Put the taxonomy and the weighting together and the roadmap writes itself. You attack the sources in order of ownership and cost — cheapest-and-most-owned first — and you measure progress on one number that respects containment.

Step 0 · Instrument

Measure weighted availability, and tag every incident by source and estimated cost. You cannot manage what you do not attribute. Source tells you who fixes it; cost tells you whether it is worth fixing first. This is the foundation; everything downstream is prioritised against it.

Step 1 · Source (a)

Kill your own-release incidents. CI/CD, data-diff on PRs, dbt unit and acceptance tests on every model. Highest ROI because you own 100% of the fix and do not need another team's roadmap to move.

Step 2 · Source (d)

Contain the blast radius. Split the monolithic DAG into independent data products so any single failure caps its damage. This is what makes the weighted number structurally improvable.

Step 3 · Source (c)

Co-own the upstream boundary. Data contracts plus explicit, AI-assisted measurement of schema and semantics, jointly owned with engineering. The hardest ownership source — tackled once the cheap wins are banked.

Step 4 · Source (b)

Partner on infra and vendors. Real SLAs, escalation paths, and self-serve detection so a degraded feed is loud and reroutable rather than silent.

Each step moves the weighted number, and the sequence matters: banking the cheap, fully-owned wins from sources (a) and (d) first buys you the credibility and the headroom to take on the cross-team work in (c). Do nothing, by contrast, and the trend runs against you — the share of practitioners naming data quality their top problem climbed from 41% to 57% in just two years.

Who owns what

Availability is a shared problem, and the fastest way to stall a 97% programme is to leave ownership implicit. A workable split:

RoleOwns
CDO / Data LeaderThe KPI itself, data-product tiering, and governance — sets the target and the weights.
Data Platform teamMeasurement, release gates, observability, and per-incident source + cost attribution.
Engineering leadershipUpstream contract compliance — schema and semantic stability at the producer boundary.
Business product ownersCriticality weighting and SLA sign-off — they decide what a stale product actually costs.
Procurement / vendor ownersExternal SLA escalation and self-serve enablement for third-party feeds.
Figure 5: The climb to 97%. The cheap, fully-owned steps — release gates and pipeline decomposition — do most of the heavy lifting early; the cross-team upstream work comes once you have headroom. Illustrative — author's framework. The 65% baseline and the per-step lifts are example figures chosen to show shape and sequence, not a forecast. Your baseline and gains will differ; the order of operations is the point.
Figure 6: Doing nothing is not neutral. The share of practitioners reporting poor data quality as a top problem rose 16 points in two years. Source: dbt Labs, State of Analytics Engineering 2024.

Where AI changes the math

The executive read: the headline is not "AI writes tests." It is reduced coordination cost. AI lets you catch upstream semantic drift before it becomes a business incident — without first having to get another team to adopt and maintain contracts they have no incentive to own. That is what makes source (c), historically the hardest boundary to move, finally tractable.

For years, source (c) was the immovable one. Catching upstream schema and semantic changes meant either getting another team to adopt contracts they had no incentive to maintain, or hand-writing brittle monitors for every field you depended on. The work did not scale, so most teams simply absorbed the incidents.

This is the layer AI actually changes — and the change is structural, not cosmetic. AI-enabled engineering lets a data team self-serve work that used to require another team's roadmap: infer a contract from the schema and semantics a producer is already emitting, watch for drift at the moment of change rather than the moment of failure, and draft the data-diff and the candidate fix before anyone is paged. The same shift applies right across the taxonomy — each source has tasks that AI can now automate or self-serve, collapsing the coordination cost that kept them stuck.

SourceWhat AI-enabled engineering can self-serve or automateWhat it changes
(a) Your releasesGenerate validation and acceptance checks from a plain-language business rule; auto-draft the data-diff and a blast-radius summary on every change.Fewer self-inflicted breaks — without hand-writing every test.
(c) Upstream changesInfer and maintain data contracts from the schemas producers already emit; flag semantic drift at change-time and route it to the producing team.The hardest boundary becomes self-served — no waiting on another team's backlog.
(b) Infra / vendorDetect a stalled or degraded feed, triage the likely cause, and draft the vendor escalation.Loud, fast detection instead of silent failure.
(d) Pipeline designMap lineage and blast radius; recommend where to split the monolith into independent products.Decomposition driven by impact, not guesswork.
Cross-cuttingRoot-cause an incident, attribute it to a source, estimate its cost from the affected products, and route it to the owner.Less firefighting coordination; your incident tags fill themselves in.

The point is not the automation itself. It is that AI makes it economical to put explicit checks at the one boundary you used to leave implicit — the producer-consumer handoff — and to self-serve work that previously stalled waiting on another team. The taxonomy does not change; AI just makes its hardest category tractable for the first time. The tooling market is already converging here: Datadog acquired Metaplane in April 2025, and anomaly detection plus agentic root-cause analysis are now table stakes for the observability layer.


The Bottom Line

97% data platform availability is not a heroics number, and it is not your vendor's number to give you. It is what you get when you do three unglamorous things in order. First, attribute every incident to a source — your releases, infra and vendors, upstream changes, or pipeline design — because each has a different owner and a different fix. Second, measure availability weighted per data product, so one contained failure reads as 80%, not 0%, and your metric finally rewards the right behaviour. Third, sequence remediation cheapest-owned-first: bank your own releases and pipeline decomposition before you take on the cross-team work upstream.

The taxonomy tells you who fixes it. The weighting tells you whether it mattered. The decomposition makes the math improvable. And AI, for the first time, makes the hardest boundary — the upstream producer-consumer handoff you used to leave implicit — something a data team can self-serve. None of these steps requires a moonshot. They require deciding to manage availability as a KPI instead of inheriting it as a vibe.

So here is the question worth taking into your next planning cycle: what is your platform's weighted availability this month — and do you even measure it per data product yet? If the answer is "we report the warehouse's uptime," you have just found your Step 0.

The concrete next step: run a 30-day weighted-availability audit across your top ten data products. For each one, can you name the owner, the SLA, the dominant failure source, and the cost of its last incident? Every blank in that table is a line item in your Step 0 — and the fastest way to turn "improve data quality" into a roadmap you can actually fund.

How aictrl.dev helps

Disclosure: aictrl.dev builds workflow orchestration and knowledge-graph tooling for engineering and data teams, so we have a commercial interest in the practices described here. That said, the research cited is independent and the playbook above runs on any stack — dbt, your warehouse of choice, and the observability tool you already own. Where we help is the cross-team boundary: making upstream changes and their downstream blast radius explicit and traceable, which is exactly source (c) in this article. If that is the layer you are fighting, see how aictrl.dev can help.


Sources and Further Reading

Published 2026-05-30 · Analysis by Bulat at aictrl.dev, co-authored with AI