In brief
- There is a simple test for whether a maintenance database is producing information or merely storing records: ask it a specific, decision relevant question, and see whether answering it needs a query or a manual read through of free text notes.
- Quality is not a feeling about your data. It is five specific, testable properties: completeness, compliance with definitions, accurate handling, sufficient population and relevance.
- Free text is the fastest way to close a work order and the slowest way to ever use the information again. It belongs as a supplement to a coded record, not as the primary record.
- A failure code without the right equipment reference underneath it is not wrong. It is meaningless.
- ISO 55001:2024 does not leave data quality as a matter of preference: Clause 7.5 (Documented information) and Clause 7.6 (Data and information), read together, make it an auditable requirement of the asset management system.

Ask your CMMS a simple question: what are the top five failure modes on your critical pump fleet this year, ranked by downtime hours. If answering that honestly means opening a spreadsheet, reading through two thousand free text work order descriptions, and making a judgement call on what each one actually meant, the problem was never the question. It was CMMS data quality.
A failure modes and effects analysis identifies what can fail and how severely. A reliability centred maintenance study decides what to do about each failure mode. Both are only as good as the evidence that validates or challenges them once the resulting maintenance strategy goes live, and that evidence lives in the computerised maintenance management system, the CMMS, as work orders, notifications and failure records. When those records are free text descriptions and inconsistent codes, the feedback loop that is supposed to keep a maintenance strategy honest stops working. The organisation is still generating data. It has simply stopped generating information.
What follows sets out what genuinely high quality failure data looks like, why structured codes and free text are complementary tools rather than competing ones, how a failure code hierarchy is actually built and governed, and why ISO 55001 now treats data quality as a certifiable requirement rather than a CMMS configuration preference. None of this demands replacing your CMMS. It demands treating the data inside it as the governed asset it already is.
The Test Every CMMS Should Pass
There is a simple test for whether a maintenance database is producing information or merely storing records.
Ask it a specific, decision relevant question. Not how many work orders were raised on this pump fleet, which any system can answer, but what are the three most frequent failure mechanisms on this pump fleet, and what did they cost in downtime and parts. If answering requires a manual read through of free text notes rather than a query, the system has been recording activity without recording the reliability information the organisation actually needs.
This gap rarely shows up as a lack of data. Technicians log a real event every time a work order closes. What is missing is structure: a description sitting in a free text field, bearing was making noise and was replaced, rather than being captured against discrete fields for failure mode, failure mechanism, failure cause and the correct equipment reference. The consequence compounds quietly. An FMECA assumption about how often a component fails cannot be checked against field evidence. An RCM task interval cannot be validated or challenged by what is actually happening in the field. A criticality rating drifts because nobody can see the failure pattern changing beneath it. The cost of a specific failure mechanism across a fleet stays invisible, because nothing links the dollars to the cause.
ISO 14224 describes the standardised recording of reliability and maintenance data as a shared reliability language, one organisation's failure record meaning the same thing as another's. That is exactly what a CMMS needs internally as well as externally: a failure record from one site, one shift and one technician meaning the same thing as a failure record from any other.
What High Quality Failure Data Actually Means
Quality is not a feeling about your data. It is five specific, testable properties.
The five properties at a glance
- Completeness
- Every mandatory field is populated for every record, not only the fields a busy technician had time for during a night shift breakdown.
- Compliance with definitions
- A failure means the same thing everywhere it is recorded: the same threshold for what counts as a failure, the same understanding of what leakage or control failure refers to.
- Accurate handling
- What happened on the plant floor is what lands in the field, with no material loss of fidelity between the event and the record, whether the entry is manual or system generated.
- Sufficient population
- Enough records, over a long enough surveillance period, for a genuine pattern to emerge rather than statistical noise being mistaken for a trend.
- Relevance
- Every field configured in a CMMS should trace back to a decision someone actually makes.
Completeness means every mandatory field is populated for every record, not only the fields a busy technician had time for during a night shift breakdown. Compliance with definitions means a failure means the same thing everywhere it is recorded: the same threshold for what counts as a failure, the same understanding of what leakage or control failure refers to, whether the record was created on a Monday morning planned job or a Sunday night emergency callout. Accurate handling means what happened on the plant floor is what lands in the field, with no material loss of fidelity between the event and the record, whether the entry is manual or system generated.
Sufficient population is the property most organisations overlook. Even accurate individual records are not useful in small numbers. A meaningful failure pattern requires enough records, over a long enough surveillance period, for a genuine pattern to emerge rather than statistical noise being mistaken for a trend. This is exactly why worst actors, equipment experiencing failures significantly more often than is normal for its class, are worth the cost of a full root cause analysis: the pattern has already demonstrated it is real.
The fifth property, relevance, is a discipline rather than a technical requirement. Collecting data nobody uses for a decision is waste dressed up as diligence. Failing to collect data a real decision depends on is a gap dressed up as efficiency. Every field configured in a CMMS should trace back to a decision someone actually makes.

Free Text Feels Efficient. It Isn't.
Free text is the fastest way to close a work order and the slowest way to ever use the information again.
Structured, coded fields exist because of what they enable that free text cannot: queries and analysis across hundreds or thousands of records in seconds rather than hours, a consistency check applied the moment data is entered rather than discovered as an error months later, and a database that stays fast and manageable rather than bloating with unstructured description text that has to be read individually to mean anything. A well designed dropdown of failure mode and mechanism codes, once built properly, is also faster to complete at the point of entry than composing a free text sentence under time pressure.
None of this makes free text the enemy. It remains genuinely necessary as a supplement, capturing the narrative detail no code list can anticipate: the unusual circumstance, the near miss that accompanied the failure, the detail a future root cause analysis will need. The mistake organisations make is treating free text as the primary record instead of the supplement to a coded one.
The code list itself has to be calibrated. Too few codes and everything gets crammed into a catch all category that tells you nothing. Too many codes and data collectors default to guessing under time pressure, defeating the purpose entirely. Codes should be mutually exclusive and pitched at the level of precision the organisation will genuinely maintain in the field, not the level that looks most impressive in a specification document.
The Hierarchy That Makes a Failure Code Mean Something
A failure code without the right equipment reference underneath it is not wrong. It is meaningless.
Reliability data is only as trustworthy as the equipment hierarchy it sits inside: industry, business category, installation, plant or unit, system, equipment class, subunit, and down to the specific component or maintainable item that actually failed. Every failure record has to be registered against the correct reference object, the specific tag number, at the correct level of that hierarchy for equipment level metrics to aggregate correctly. Get the reference object wrong and the consequence is not a rounding error. It is invisible failure. A substantial share of failure records are commonly not correctly linked to their equipment tag, which quietly invalidates every fleet level failure count, every mean time between failures figure, and every reliability conclusion drawn from that data, without anyone realising the numbers were never real.
This is where nomenclature becomes the difference between a record that means something and one that does not. A failure record built from component, failure mechanism and failure cause is queryable and comparable in a way free text never will be. Bearing was making noise and was replaced tells a reader something, once, if they happen to open that specific record. Bearing, worn, lack of lubrication is a different kind of statement entirely: it names the component, the mechanism, wear, one of a defined set that also includes fatigue, corrosion, looseness, cavitation and control failure among others, and the root cause. That second version can be counted, ranked against every other bearing failure on the same equipment class, and traced directly back to the failure mode line in the original FMECA that predicted it.

A Governance Obligation, Not a Configuration Task
ISO 55001 does not leave data quality as a matter of preference.
In ISO 55001:2024, Clause 7.5, Documented information, requires the information an organisation relies on to be controlled: available and suitable for use when needed, protected from loss of integrity, retained and version controlled over its life. Clause 7.6, Data and information, then applies specifically to asset management, requiring the organisation to determine the quality requirements of that information and to maintain the alignment, consistency and traceability of information and terminology between the financial and non-financial functions. Read together, the two clauses convert getting the failure data right from good practice into an auditable, certifiable requirement of the asset management system, not a CMMS administrator's private concern.
The GFMAM Asset Management Landscape reinforces the same point from a different angle, describing data and information as something to be managed as an asset in its own right, with ownership and stewardship responsibilities defined across its life cycle, the same discipline the organisation already applies to physical assets and to financial records. Configuration management exists precisely to keep the asset register and the physical asset in agreement as both change over time, and a dedicated international standard, ISO 8000, exists solely to define what data quality means. Treated this way, a failure code register has an accountable owner and a periodic verification routine, the same way a financial ledger is reconciled rather than assumed correct because it exists.
What Clean Data Actually Buys the Organisation
This is where the investment case for structured failure coding gets made in dollars, not principles.
- Reliability engineering gets a genuine Pareto ranking of failure modes by frequency, cost and downtime, turning which equipment to study next from a guess into a calculation, and clearly identifying the worst actors that justify a full root cause analysis rather than another routine repair.
- Cost management gets a real driver breakdown instead of a single lagging number. Maintenance cost as a percentage of replacement asset value, or RAV, is a widely used benchmark, but on its own it only tells you the number is high or low. Correctly coded failure data tells you which failure mechanism, on which equipment class, is actually driving that number, turning a scorecard metric into a diagnostic one.
- Operations gets metrics it can trust. Mean time between failures (MTBF), mean time to repair (MTTR) and, by extension, overall equipment effectiveness (OEE) all depend entirely on accurately time stamped failure events correctly attributed to the right equipment. Build those calculations on unreliable source records and every metric derived from them inherits the same unreliability, however precise the resulting number looks on a dashboard.
- Commercial recovery becomes defensible rather than disputed. A correctly coded, evidenced failure record is what makes an OEM warranty claim stand up rather than get argued down, and that is real, recoverable value that an untraceable free text ticket simply forfeits.
- Capital and spares planning move from estimate to calculation. Failure mechanism and frequency data, correctly linked to criticality, feed directly into stocking policy and into the timing of renewal decisions, replacing a provisioning guess with a number someone can actually defend to a capital committee.

Building It Without Stopping the Plant
None of this requires a CMMS replacement project. It requires sequencing and a small number of governing habits, applied consistently.
- Lock the equipment hierarchy and tag structure first. Every other improvement depends on failure records landing against the correct reference object, so this step has to be settled before configuring anything else, not attempted in parallel with it.
- Configure a discrete, coded field set for failure mode, failure mechanism, failure cause and detection method, aligned to a recognised taxonomy so the codes mean the same thing across every site and every shift, with free text retained as a mandatory narrative supplement rather than an alternative path around the coded fields.
- Prioritise coverage by criticality rather than chasing full completeness everywhere at once. Close to full, compulsory coverage on the equipment that matters most, a solid majority on the next tier down, and a lighter but still present standard for the remainder, is a genuinely defensible starting position rather than a compromise.
- Name a data owner and a steward for the failure code register, with a periodic verification step built into the routine rather than left to be discovered during an audit, the same discipline already applied to a financial ledger.
- Use root cause analysis deliberately on the worst actors the data identifies, both to fix the failure and to backfill the mechanism and cause detail the original work order could not capture under time pressure.
This sequence sits inside the wider data standardisation and AI readiness framework, and for a worked example of failure coding rebuilt in an operating environment, see the SAP failure codes case study.

Closing the Loop
The CMMS is not an administrative record keeper sitting beside the asset management system. It is the sensor network for its evidence base. Every FMECA assumption, every RCM interval and every criticality rating is only as current as the failure data validating it, and that data is only as useful as the structure it was captured against.
A structured failure code hierarchy is what converts routine record keeping into a genuine decision support asset: the difference between a database that stores what happened and one that can tell you what to do next. Pick one critical equipment class. Pull twelve months of its failure records. Run the test from the opening of this article against it. What it reveals will tell you exactly where to start.
Frequently asked questions
What does high quality CMMS failure data actually mean?
It is five specific, testable properties rather than a feeling about your data: completeness, meaning every mandatory field is populated for every record; compliance with definitions, meaning a failure means the same thing everywhere it is recorded; accurate handling, meaning what happened on the plant floor is what lands in the field; sufficient population, meaning enough records over a long enough surveillance period for a genuine pattern to emerge; and relevance, meaning every field configured in a CMMS traces back to a decision someone actually makes.
Is free text in a work order a problem?
Free text is the fastest way to close a work order and the slowest way to ever use the information again. It remains genuinely necessary as a supplement, capturing the narrative detail no code list can anticipate, such as the unusual circumstance or the near miss that accompanied the failure. The mistake organisations make is treating free text as the primary record instead of the supplement to a coded one, because only coded fields can be queried and consistency checked across hundreds or thousands of records.
Does ISO 55001 require the organisation to manage data quality?
Yes. Clause 7.5, Information requirements, requires the organisation to determine the quality requirements of the information it relies on for asset management decisions, and to maintain consistency and traceability between technical data and the financial and non-financial data connected to it. Clause 7.6, Documented information, then requires that information to be controlled: available and suitable for use when needed, protected from loss of integrity, retained and version controlled over its life. Read together, the two clauses make getting the failure data right an auditable, certifiable requirement of the asset management system.
Why does the equipment hierarchy matter to a failure code?
A failure code without the right equipment reference underneath it is not wrong, it is meaningless. Every failure record has to be registered against the correct reference object, the specific tag number, at the correct level of the hierarchy for equipment level metrics to aggregate correctly. Get the reference object wrong and the consequence is not a rounding error, it is invisible failure: fleet level failure counts, mean time between failures figures and the reliability conclusions drawn from them are quietly invalidated.
Do we have to replace the CMMS to fix this?
No. It requires sequencing and a small number of governing habits, applied consistently: lock the equipment hierarchy and tag structure first, configure a discrete coded field set for failure mode, mechanism, cause and detection method aligned to a recognised taxonomy, prioritise coverage by criticality rather than chasing full completeness everywhere at once, name a data owner and steward for the failure code register with a periodic verification step, and use root cause analysis deliberately on the worst actors the data identifies.

