The Data & AI Foundation

Data foundations for AI

AI applications run on data that already sits in your ERP, MES, CRM, and the spreadsheets around them. Before anything is built, we work out how that data is structured, where each field comes from, what it means, and whether it holds up under use. That assessment usually changes the plan, because some use cases turn out to be ready and others have to wait for the source systems to catch up.

Foundations

The six layers underneath an AI application

Before development starts, we map how the work moves through the operation and where the data behind it lives. Weakness in the lower layers shows up during the build as integration work nobody scoped, or as output the business will not sign off on. We assess all six at the start of an engagement and fix what a specific use case depends on, rather than rebuilding the whole stack first.

AI & analyticsdepends on the layers beneath it06Master data managementone authoritative version of each core entity05Metadatadefinitions, ownership, freshness, quality04Data lineagewhere each field comes from and how it changes03Data modelthe entities and how they relate02Data architecture & mappingwhat each system holds and how they connect01Process mappinghow the work moves through the operationeach layer is a precondition for the one above
The six layers, with AI and analytics built on top of them.
01
Process mapping

How the work moves through the operation: the steps, the handoffs, and the decisions people make at each one. We document this before automating any part of it.

02
Data architecture & mapping

Which systems hold which data and how they connect. We scope the integration work here, before a build starts depending on it.

03
Data model

The entities and their relationships: what counts as a customer, an asset, or a shipment in your systems, and how those records tie to each other.

04
Data lineage

Where each field comes from and how it is transformed on the way, so a number in a model output can be traced back to its source system.

05
Metadata

Definitions, ownership, freshness, and quality for each field. Without it, two teams can pull the same measure from the same table and report different numbers.

06
Master data management

One authoritative version of each core entity. Where customer or item lists have diverged across systems, they get reconciled before anything is built on them.

Architecture

One modeled data layer

Source systems feed one operating data layer where the data is cleaned, joined, and tracked for lineage and metadata. Dashboards, models, and the applications we build read from that layer instead of from separate extracts, so a definition changed in one place changes everywhere it is used.

SOURCE SYSTEMSOPERATING DATA LAYERCONSUMED BYERPMESSensors & telemetrySpreadsheetsIngest & cleanModel & joinTrack lineage & metadataDashboardsAI modelsOperating decisions
Source systems feed one modeled layer with lineage and metadata tracked. Dashboards, models, and operating decisions read from it.
Assessment

How we measure data quality

A field can be modeled correctly and still be unusable. Before anything is built on it, we profile the source data against six dimensions and report what we find, including the cases where a use case should wait until the source system improves.

Completeness
the required rows and fields are present
Accuracy
the value matches what is true in the operation
Consistency
the field agrees across the systems that hold it
Timeliness
the data is current enough for the decision it supports
Validity
values fall inside the expected format and range
Uniqueness
one record per real-world entity, duplicates resolved
Technique selection

Matching the technique to the problem

A recurring check against fixed criteria is usually a rules or SPC problem. Allocating constrained capacity is an optimization problem, and a number that drifts is a forecasting one. Learned and generative models go where there is no stable rule to write down and the pattern has to come from the data. Simpler techniques cost less to run and are easier for your team to maintain, so we start with the simplest one that moves the number.

Rules & SPCStatistical modelsClassical MLOptimizationLearned / generative AISIMPLERMORE CAPABLE
Illustrative. Techniques ordered from simpler to more capable, from rules and SPC through optimization to learned and generative models.
Choosing what to build

Value against data readiness

Every candidate gets two scores: what the improvement is worth, and whether the source data can support it. Value is estimated with the business owner and finance, so the number is one they will recognize when the work is reviewed later. Readiness comes out of the data quality work above. Candidates then sort into four groups.

BUILD NOWPILOT / DE-RISKWATCHSKIPValue at stake →Data readiness →Downtime predictionBUILDInvoice matchingBUILDDemand forecastPILOTVision QA · Line 4HOLDOps-manual chatbotSKIP
Illustrative. Candidates plotted by value at stake against data readiness, then sorted into build, pilot, hold, and skip.
BUILD

High value, and the data is ready. Fund it now.

PILOT

High value, with data gaps. Prove it on a narrow slice before committing.

HOLD

Worth doing, not yet feasible. Revisit when the data or readiness improves.

SKIP

Low value, or handled without a model.

Bring us one use case you want to build.

We will look at the source data behind it and tell you what it would take to make that data usable.

Book a discovery call