SAMPLING BIAS
You Get What You Checked For
Inspection data does not come from a random sample. It comes from decisions about where inspectors had access, which areas were flagged as risky, what regulations required, and what equipment was available at the time. All of that shapes the dataset long before anyone trains a model on it.
WHY IT MATTERS
The dataset inherits the history behind it.
A lot of that history is not sitting in one clean system. Plenty of integrity data still lives in binders in a field office, or in spreadsheets saved across 30-plus different folders, drives, and formats, each with its own habits for what got recorded, how consistently, and by whom. Consolidating that into a single dataset does not erase those old habits. It inherits them.
Here is why that is easy to miss.
- “High findings” can mean “looked at closely,” not necessarily “worse condition.” An area inspected repeatedly may produce more recorded findings simply because it received more scrutiny than an area inspected once.
- A quiet record is not the same as a safe asset. Few or no findings can mean limited inspection or limited evidence, not necessarily better underlying condition.
- Older data was often collected differently, sometimes by hand. What was measured, how often, and with what tools has changed over time, and years of records may have started as paper before ever becoming digital.
- Scattered spreadsheets carry scattered standards. When integrity data has lived in dozens of separate files over the years, those files may not use the same fields, units, naming conventions, or judgment calls even before a model sees the data.
- Some locations get extra attention because something already looked off. That is normal, sensible engineering judgment, but it also means those records may represent targeted investigation rather than typical asset conditions.
WHAT THE MODEL CAN LEARN
Sometimes the model learns the inspection pattern too.
A model trained on this kind of data can end up learning the inspection pattern itself: where people tended to look, what they tended to record, and how they recorded it, instead of learning only what is true about the equipment.
This does not mean the data is bad or the model is wrong. It means the first question worth asking is not “What does the model predict?” It is “How was this data collected, and does that shape the answer we are getting?”
THE QUESTION BEFORE THE PREDICTION
How was this data collected, and does that shape the answer we are getting?
That question belongs in model validation whenever historical inspection patterns, uneven coverage, legacy records, or changing inspection methods shape the training data.
