Predicting equipment failure 72 hours out
False-positive rate dropped to 9%, downtime reduced 34%, and OEE improved by 6 percentage points.
- Industry
- Industrial manufacturing
- Size
- 14 plants, 47 production lines
- Engagement
- 16 weeks
- Team
- 5 engineers, 1 reliability SME
- False-positive rate
- 9%False-positive rate
- Downtime reduction
- 34%Downtime reduction
- OEE lift
- +6 ptsOEE lift
Unplanned downtime cost the company an estimated $42M annually across 14 production lines.
Unplanned downtime cost Vertex an estimated $42M annually across 14 plants. The telemetry needed to predict most of those failures already sat in their historian, unused. A previous vendor pilot had produced alerts at roughly a 40% false-positive rate, and the floor had responded the way any team would — they muted it, and were sceptical of anything that followed.
The engagement
- Industry
- Industrial manufacturing
- Size
- 14 plants, 47 production lines
- Engagement
- 16 weeks
- Team
- 5 engineers, 1 reliability SME
A multimodal model combining SCADA telemetry, vibration sensors, and floor-camera vision to forecast failures.
Four phases, each with a written definition of done.
Reconstruct the labels
Usable failure labels did not exist. We derived them from work orders and operator logs, then had Vertex's reliability engineers validate every one before it entered a training set.
Optimize for precision, explicitly
Given the history, recall was worth less than credibility. We tuned against false positives first and accepted a shorter lead time on marginal asset classes to get there.
Go multimodal
SCADA telemetry alone plateaued well short of a useful lead time. Adding vibration spectra and floor-camera vision extended it to 72 hours across the main asset classes.
Deploy at the edge
The OT network is air-gapped from IT. Inference runs on edge hardware inside the plant boundary, with only aggregated predictions crossing into corporate systems.
What changed.
The false-positive rate settled at 9% and has held there through a full year of operation, which is what earned the system a standing place in shutdown planning. Downtime on monitored lines fell 34% and OEE improved by six percentage points. Vertex has since extended coverage to three additional asset classes using the same pipeline.
- False-positive rate
- 9%False-positive rate
- Downtime reduction
- 34%Downtime reduction
- OEE lift
- +6 ptsOEE lift
“The last vendor gave us a model. This team gave us something the floor supervisors actually check before they schedule a shutdown.”
What we built it with.
- PyTorch
- Python
- Kafka
- TimescaleDB
- Docker
- Azure
The capabilities behind this engagement.
More of our work in manufacturing.
Predictive maintenance, quality vision, and OEE-lifting systems that live on the line, not in the dashboard.
Other engagements worth reading.
Cutting prior-auth cycle time by 71%
Average cycle time fell from 4.2 days to 1.2 days. Denials dropped 38% in the first quarter post-launch.
Read moreGrounding an AI research analyst on a decade of data
Analyst throughput doubled on coverage tasks and onboarding time for new hires dropped by half.
Read moreLet's talk about what this would look like for you.
Thirty minutes with the senior team. Bring the problem and we will give you an honest read on scope, timeline, and whether we are the right fit.
