Predictive maintenance: how AI predicts machine failures in a smaller plant

Predictive maintenance means planning a repair from measurements of a machine’s condition, such as vibration, temperature, motor current and spindle load (how hard the motor turning the cutting tool is working), before the part fails. Software learns what each machine looks like when it runs normally and flags a change early enough to book the work at a planned stop. A smaller manufacturer can start on the one or two machines whose breakdown stops shipments, sometimes with data the machine’s CNC control, the computer that runs a machine tool, already records. This guide covers the signals, the software, the first steps and how to measure the result.

A row of identical stamping presses along a clean, polished factory floor

What predictive maintenance is

The US Department of Energy’s operations and maintenance guide defines predictive maintenance as “measurements that detect the onset of system degradation,” taken so the cause can be dealt with “prior to any significant deterioration in the component physical state.” In plain words, you measure a machine while it runs, watch for the change that wear produces, and plan the repair before the part breaks. Measuring alone is called condition monitoring. Predictive maintenance adds the decision: when a reading drifts, someone books an inspection or a repair at a planned stop.

The US National Institute of Standards and Technology (NIST) surveyed US manufacturers on machinery maintenance and published the results in 2020. On average, 45.7% of the maintenance they reported was reactive, meaning repairs after a breakdown. Among plants that relied mainly on preventive and predictive maintenance, the half that used the most predictive maintenance “was associated with 15 % less downtime, an 87 % lower defect rate, and 66 % less inventory increases due to unplanned maintenance.” The survey received 85 responses and reports an association. It does not measure what any one product delivers.

Preventive vs predictive maintenance

Reactive maintenance, also called run to failure, repairs a machine after it breaks. Preventive maintenance does the work on a schedule, by the calendar or by running hours. Predictive maintenance does the work when a measurement shows wear. The table compares the three on a machining centre, a CNC milling machine that changes its own tools.

ApproachWhat triggers the workExample on a machining centreWhat it needsEstimate in the DOE guide
Reactive (run to failure)A breakdownA spindle bearing fails in the middle of a jobSpare parts on the shelf and people free for overtimeNo savings figure. The advantages it lists are “Low cost” and “Less staff”
PreventiveThe calendar or running hoursFilters and lubrication done every set number of hoursA schedule and a record of completed work12% to 18% cost savings over reactive maintenance
PredictiveA measured change in conditionA rising vibration trend on the spindle, so the bearing is replaced at the next planned stopSensors or control data, a baseline of normal running, and a person who acts on alerts8% to 12% cost savings over preventive maintenance

The DOE guide was published in August 2010, so treat its percentages as a rough direction. It lists “Increased investment in diagnostic equipment” and “Increased investment in staff training” among the disadvantages of predictive maintenance, and it describes top-performing facilities as running under 10% reactive, 25% to 35% preventive and 45% to 55% predictive maintenance. For a smaller plant, that points to a mix: schedules for lubrication and filters, run to failure for cheap parts, and monitoring on the few machines whose breakdown stops shipments.

What a machine tells you: vibration, temperature, current and spindle load

Each signal shows a different kind of wear. In the DOE guide, vibration, temperature and motor current apply to motors, pumps and other rotating equipment. Spindle load comes from a machine tool’s own control.

SignalWhat it can revealHow it is collected
VibrationUnbalance, misalignment, looseness, and bearing, gear and belt problems, among the faults the DOE guide listsAn accelerometer, a small sensor that measures vibration, mounted on a bearing housing or held against it during a route of handheld readings
TemperatureFriction in rotating parts, and loose or corroded electrical connections, which run hot before they failA temperature sensor on the machine, or an infrared camera that shows heat as an image, a method called thermography
Motor currentChanges in the mechanical load on a motor, which show up as changes in its current. The DOE guide calls this motor current signature analysisA current sensor on the motor’s power cables. The DOE guide describes the method as non-intrusive
Spindle loadHow hard the spindle motor works compared with its rating. Changes in spindle power can point to tool wear or a broken toolThe CNC control, read through MTConnect or the control maker’s own software

A CNC machine already produces some of this data. MTConnect, an open, royalty-free standard for reading data from manufacturing equipment, defines load as “the measurement of the actual versus the standard rating of a piece of equipment,” in percent, and its interface is read-only. FANUC’s AI Servo Monitor uses data from the servo and spindle drives on machines with FANUC controls, with “No additional sensors necessary.” Caron Engineering’s TMAC measures “true spindle motor power” to detect tool wear and breakage.

The DOE guide also covers ultrasonic analysis, which picks up sound above human hearing from wear and leaks, and oil analysis, which tests lubricant for wear particles.

How AI predicts a machine failure

Software turns readings into a warning in three ways, and many products combine them.

Limits a person sets

The simplest method raises an alert when a reading passes a set value. The DOE guide says published vibration charts “are not absolute vibration limits above which the machine will fail,” and still advises setting limits “above which action will be taken.”

Anomaly detection against a baseline

Anomaly detection learns what normal running looks like for one machine, called its baseline, and scores how far new readings drift from it. FANUC’s European product page describes “Automatic creation of a baseline model through machine learning while the machine is in production,” with the drift shown as an “anomaly score.” Because it learns from normal running, it needs no record of past failures.

Failure prediction and remaining useful life

Failure prediction estimates remaining useful life, the running time a part has left before it stops doing its job. Siemens says Senseye highlights “risks, remaining useful life and priorities,” and IBM says Maximo can “estimate when assets will no longer perform optimally” from inputs that include “failure patterns.” A plant with few recorded failures gives this method less to learn from.

Where a language model fits

The models that read sensor data are trained on numbers. A large language model, the kind behind ChatGPT and Claude, works on text: work orders, technician notes and machine manuals. It can answer “when was this spindle last rebuilt?” with the work order named, and draft the work order when an alert arrives, for a planner to approve. Human in the loop explains that approval step.

What a smaller plant needs to start

Start with one or two machines. A small pilot is cheaper to stop, and it shows whether your team acts on the alerts.

  1. Pick the machines. Choose the ones whose breakdown stops shipments and has no backup, such as the bottleneck machining centre or the main air compressor.
  2. Pull the history. Gather two years of breakdowns, repairs and parts from the work orders, the paper log or the ERP, the business system that holds orders, inventory and accounting. A CMMS, or computerized maintenance management system, is software built for this history. A spreadsheet will do for a start.
  3. Check what the machine already reports. A CNC control may publish spindle load and alarms through MTConnect or its maker’s software. A PLC, the programmable logic controller that runs a machine, may already hold motor readings. Ask the machine builder before you buy sensors.
  4. Add sensors where nothing is measured. Motors, pumps, fans and gearboxes are the natural places for a wireless vibration and temperature sensor.
  5. Record a baseline. Let the software learn normal running across your usual mix of jobs. A job it has never seen can look like a fault.
  6. Name who acts on an alert. Decide who reads alerts, how fast, and what each can lead to: an inspection, a repair at the next planned stop, or a closed alert with a note. MaintainX, a work order system, says it can “automatically trigger work orders when meter or sensor readings detect a potential failure.”
  7. Set a review date. Six months on the same machines gives you something to compare with the history from step 2.
If nobody has time to run it

The DOE guide says monitoring equipment “should not be purchased for in-house use if there is not a serious commitment to proper implementation, operator training, and equipment monitoring and repair.” Without that commitment, it suggests contracting the work to an outside vendor that brings its own equipment and expertise.

Ask where the readings will live. Siemens says Senseye Cloud “works with your existing data from legacy machines, historians, IoT platforms or new sensors,” where a historian is a database of time-stamped machine readings. If readings must stay on a server you control, ask each vendor where they are stored and processed. Private AI for business lists the questions to put in writing. Spare parts belong in the same plan: the DOE guide notes that predictive maintenance lets a plant “order parts, as required, well ahead of time,” and AI for inventory management covers reorder points.

IoT sensors for predictive maintenance

An IoT sensor, short for Internet of Things, sends its readings over a network without anyone collecting them by hand. It is fixed to the machine and reports to a receiver, called a gateway, which passes the readings to the vendor’s software. Tractian lists its sensor’s measurements as “Vibration, Trends, and Spectrum up to 64kHz, Temperature, RPM, and Ultrasound,” over a 4G/LTE cellular connection. RPM is shaft speed in revolutions per minute. Augury’s sensors capture “vibration, temperature, and magnetic data.” Before you mount one, ask these questions:

Predictive maintenance software, and what each vendor says it does

The products below come from a control maker, two industrial software companies, two sellers of sensors with software, and a work order system. The descriptions are the vendors’ own, checked September 28, 2026.

ProductWhat the vendor says it doesWhat it reads
FANUC AI Servo Monitor“can predict possible failures of the drive systems for FANUC servo and spindle motors”The servo and spindle drives on machines with FANUC controls, with “No additional sensors necessary.” It works with FANUC’s MT-LINKi data collection software
Siemens SenseyeHighlights “risks, remaining useful life and priorities”“existing data from legacy machines, historians, IoT platforms or new sensors”
IBM Maximo Asset Performance ManagementForecasts emerging issues and estimates “when assets will no longer perform optimally”“historical asset health data, operating conditions, failure patterns”, as an extension of Maximo asset management
Augury Machine HealthDiagnostics from its AI, which it says trained vibration analysts backIts own sensors, which capture “vibration, temperature, and magnetic data”
Tractian“Tractian AI pinpoints root causes automatically using vibration signatures,” with work order management in the same productIts own sensors: vibration up to 64 kHz, temperature, RPM and ultrasound
MaintainXA work order system that can “automatically trigger work orders when meter or sensor readings detect a potential failure”Meter and sensor readings from connected equipment and other systems

If the machines that matter have FANUC controls, the control maker’s software starts from data they already produce. For pumps, fans and compressors with no data, a sensor vendor fits better. If you run a work order system, ask what it connects to before adding another screen. Ask each vendor how it charges, for example per sensor, per monitored machine, per user or by quote, and check its own pricing page. Camera inspection is the other plant-floor use of AI that needs its own hardware, covered in computer vision in manufacturing. Industrial AI sorts the uses that need sensors from those that run on records a plant already keeps.

Start from the records you already keep

Tell Derik which machines stop your plant when they fail and how your maintenance history is kept today. He will tell you which part of this ThriveAI can build on those records, and which part needs a sensor vendor.

Start a conversation

How to measure whether predictive maintenance works

Compare the pilot machines with their own history. The DOE guide lists the first three measures below, with benchmarks for a whole maintenance program that it takes from a NASA source published in 2000.

MeasureHow to calculate itBenchmark in the DOE guide
Equipment availabilityHours the machine was available to run, divided by the total hours in the periodOver 95%
Emergency maintenance percentageHours worked on emergency jobs, divided by all maintenance hours workedUnder 10%
Maintenance overtime percentageOvertime hours, divided by regular maintenance hours in the periodUnder 5%
Mean time between failuresRunning hours, divided by the number of breakdowns in the periodNone listed
Alert outcomesWhat each alert led to: a real finding, a false alarm or no action, plus any breakdown that came with no alertNone listed

Have the technician note what each alert led to. After six months, count real findings, false alarms and breakdowns that came with no warning. When most alerts turn out false, people stop trusting them, the same way approval steps turn into rubber stamps, so tune the limits or the baseline before you add machines.

How ThriveAI helps

ThriveAI is an AI engineering company in Ottawa that builds private AI systems on a client’s own data, for manufacturers and distributors in Ontario and Quebec. It does not sell or install sensors. Its part of predictive maintenance is the written record: work orders, the maintenance log, parts history and the ERP, with answers that name their source and work orders drafted for a named person to approve. The platform is designed to keep each client’s data on its own server in Canada. The client chooses a model on that server or a hosted model under a written zero data retention agreement, a contract under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada. Derik Lawlis, the founder, leads every project and stays close to the build. See AI for manufacturing and About ThriveAI.

Questions people ask

What is predictive maintenance?
Predictive maintenance is planning a repair from measurements of a machine's condition, such as vibration, temperature, motor current or spindle load, before the part fails. Software compares new readings with the machine's normal running and flags a change early enough to book the work at a planned stop.
What is the difference between preventive and predictive maintenance?
Preventive maintenance does work on a schedule, by calendar or running hours. Predictive maintenance does work when a measurement shows wear. The US Department of Energy's 2010 guide estimates 12% to 18% cost savings for preventive over run-to-failure maintenance, and 8% to 12% for predictive over preventive. A plant can run both, with schedules for lubrication and monitoring on critical machines.
What sensors are used for predictive maintenance?
Vibration, temperature and motor current sensors are the common ones. Accelerometers on bearing housings pick up unbalance, misalignment, looseness and bearing, gear or belt problems. Infrared cameras show hot spots such as loose electrical connections. On CNC machines the control already measures spindle and servo load, and FANUC sells software that reads it with no added sensors.
How does AI predict machine failure?
In three ways, often combined. A limit raises an alert when a reading passes a set value. Anomaly detection learns a baseline of normal running and scores how far new readings drift from it, with no need for past failures. Failure prediction estimates remaining useful life and learns from recorded failure patterns.
Can a small manufacturer do predictive maintenance?
Yes. Start on the one or two machines whose breakdown stops shipments, pull their repair history, check what their controls already report, add sensors only where nothing is measured, and name who acts on each alert. The US Department of Energy suggests contracting the work to an outside vendor if a plant cannot commit to training and follow-through.
How do you measure whether predictive maintenance is working?
Compare the pilot machines with their own history. Track equipment availability, the share of maintenance hours spent on emergency jobs, and maintenance overtime, for which the US Department of Energy's guide lists benchmarks of over 95%, under 10% and under 5%. Record what each alert led to, so you can count real findings, false alarms and missed failures.

Contact

Start with the machine that stops shipments

Tell Derik which machines stop your plant when they fail and how breakdowns are logged today. He will tell you what your current records can support and what would need sensors.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.