ForgeShift
← Insights
How-To · 4 min read

Your First Data Product Is the Downtime Log.

Six steps, no new sensors, no platform -- how a plant turns an existing spreadsheet into something a named person is accountable for.

By

Principal, ForgeShift Advisory

Key takeaways
  • A dataset becomes a data product when it gains a named owner, a published interface, numeric service levels and governed access.
  • Name one person, not a department -- ownership by team is the most common way a service level quietly decays.
  • Build the first product from data the plant already produces; the platform earns its scope from products that already work.

The downtime log exists. It lives on a clipboard by line 3 and gets typed into a spreadsheet on somebody's laptop most Fridays. When the plant manager asks why the line lost four hours in June, the answer takes two days and three people -- and the number maintenance quotes is not the number that comes back from the office.

The reflex at that point is to scope a platform. A data lake, a warehouse, a BI layer, eighteen months. That sequence is the most reliable way to spend the budget before proving any value. The log is not failing because it lives in a spreadsheet. It is failing because nobody owns it, nothing is promised about it, and there is no fixed place to find it.

Fixing those three things turns a dataset into a data product. None of them require new sensors.

What separates the two

A dataset is a file or a table. A data product adds four things a dataset does not have: a named owner accountable for it, a published interface consumers can build on, measurable service levels for freshness and completeness, and governed access. Documentation alone does not close the gap -- a wiki page attached to an export is still an export.

The framing is not new. It comes from data mesh, where Zhamak Dehghani argued that the domain which creates analytical data should own and serve it, and that "the consumers of that data should be treated as customers." The same work sets out six qualities that separate a product from a dataset: discoverable, addressable, trustworthy, self-describing, interoperable, and secure. What follows is that list turned into a build order.

Six steps, in this order

Pick the domain that creates it. Not the one that consumes it. Maintenance generates the downtime log, and maintenance is the only group that knows what it means when a four-hour entry looks wrong. Finance can tell you the number is odd. Maintenance can tell you the line ran on a borrowed motor that week.

Name the owner. One person, by name. owner: "Operations" is nobody. owner: "M. Chen, maintenance lead" is someone. This is the step that gets skipped, and skipping it is why most of the rest quietly collapses -- a service level with no owner has nobody for whom its failure is personally a problem.

Define the interface. What is published, in what shape, on what schedule. Agree it with the two or three people who will actually consume it, before it is built rather than after they complain. A table refreshed nightly with a fixed column set is an interface. A CSV emailed on request is not.

Set the promise as numbers. Completeness, freshness, accuracy -- the values the owner will be held to. Dehghani's term for this is a service level objective around the truthfulness of the data. On a plant floor it reads plainer: 99% complete, less than four hours late is a promise. "Usually good" is not a service level. A promise without a number is a hope with a schedule.

Publish it. Permanent address, discoverable without knowing who to ask. If the route to the data is a person's memory, the product has a single point of failure and it is a human being.

Operate it. Monitor against the promise, respond when it breaks, and change it deliberately with the consumers who depend on it. This is the step that never ends and the one budgets forget.

The two objections worth answering

We need the platform first. Build platform, then find use case, and the effort stalls against requirements nobody has stated. Note the sequence in data mesh itself: self-serve infrastructure is the third principle, not the first. Ship the first product with what already exists, then generalise the parts that repeat, and the second one costs less than the first. The platform earns its scope from products that already work.

We named an owner last year. Then check whether that is still true. Ownership drifts: in Q1 the owner is named and the service level is met; by Q3 the owner has moved on and the service level is still advertised. Reorganisations move people; the published promise stays published. Ownership is a fact that has to be maintained, not a box that was ticked.

Monday

Pick the one dataset three different people ask for by email -- downtime events, quality records, changeover times. Write a person's name on it, not a department's. Ask the people who ask for it what shape they want it in and how often. Publish two numbers you are willing to be held to. Put the address in one place everyone can find without asking.

That is a data product, built out of what the plant already produces. The platform conversation gets easier afterwards, because by then there is something real to generalise from.

Keep reading

Get the next brief before your competitors do.

Weekly field notes on downtime, OEE and the AI pilots that actually scale — or talk to us directly about your operation.

We use only essential cookies. We don't use advertising or tracking cookies. See our Privacy Policy.