Case study 02 / Industrial
Industrial Predictive Maintenance
Vibration and thermal telemetry processed at the edge, modeled in the cloud and surfaced to maintenance planners as ranked, explainable work recommendations.
Illustrative reference architecture, not a client engagement. Outcomes described here are design goals, not measured results.
Reference flow
05 stages
01Challenge
Why this problem is hard.
01
Collected but unused
Condition data lives in historians and local HMIs, with no path into the tools planners use.
02
High-frequency signals
Vibration diagnosis needs kilohertz sampling. Streaming raw waveforms from every asset is costly and often conflicts with plant network policy.
03
Alerts without context
Threshold alarms fire without saying which component, which failure mode or how urgent. Planners learn to ignore them.
04
OT and IT separation
Plant networks are segmented for good reason. The design has to respect that boundary, not flatten it.
02Business context
Assumptions and constraints.
Constraints the design must respect
- 01One-way data flow from OT
- Telemetry leaves the plant network through a controlled conduit. Nothing in the cloud can write to control systems.
- 02Intermittent connectivity
- Edge nodes buffer locally and keep computing features through network outages.
- 03CMMS stays the system of record
- Recommendations become work requests in the existing maintenance system, reviewed by a planner.
- 04Explainability
- Reliability engineers must see why an asset was ranked: which features moved, and against which baseline.
03Architecture
Reference architecture.
Accelerometers and temperature sensors on priority assets, plus existing PLC tags such as load, speed and run state. Operating context matters: vibration at partial load means something different from vibration at full load.
- Tri-axial accelerometers
- Thermal sensors
- PLC tags via OPC UA
- Asset registry mapping
Concept architecture
01 / 07
Sensors & acquisition
Accelerometers and temperature sensors on priority assets, plus existing PLC tags such as load, speed and run state. Operating context matters: vibration at partial load means something different from vibration at full load.
- Tri-axial accelerometers
- Thermal sensors
- PLC tags via OPC UA
- Asset registry mapping
Key patterns
- Edge feature extraction
- Outbound-only OT conduit
- Regime-aware anomaly detection
- Asset digital twin
04Technology
Representative technology.
01
05 items
Edge & OT
- Industrial gateways
- OPC UA
- MQTT
- Containerized edge runtime
- Local buffering
02
04 items
Data platform
- Time-series database
- Lakehouse with open table format
- Stream processing
- Data quality checks
03
04 items
Machine learning
- Signal processing (FFT, envelope)
- Anomaly detection
- Survival analysis
- Model registry and monitoring
04
04 items
Cloud & operations
- Kubernetes
- Terraform
- Device identity and certificate management
- OpenTelemetry
05Engineering approach
Principles behind the design.
- 01
Failure modes first, then ML
Start from known failure modes and their signatures. Models extend reliability engineering knowledge rather than replace it.
- 02
Compute where the signal is
High-frequency processing stays at the edge. The cloud handles fleet-level learning, history and planning.
- 03
Respect the OT boundary
Outbound-only data flow and no write path to control systems, enforced by architecture rather than policy alone.
- 04
Close the label loop
Planner decisions and inspection findings return as labels. That history is what turns anomaly scores into credible predictions.
06Implementation
A phased delivery path.
Delivery path
04 phases
- 01Criticality & instrumentation
- 02Edge & data foundation
- 03Models & twin
- 04Planner integration & scale
- Phase 0101
Criticality & instrumentation
Choose a pilot asset class by criticality and failure history; agree sensor placement and the network path with OT engineering.
Deliverables
- Criticality ranking
- Sensor plan
- Conduit design
- Baseline data capture
- Phase 0202
Edge & data foundation
Deploy gateways, feature extraction and the time-series platform for the pilot line.
Deliverables
- Edge runtime and features
- Secure transport
- Time-series and lakehouse tables
- Data quality dashboard
- Phase 0303
Models & twin
Train regime-aware models, build the asset hierarchy and review rankings with reliability engineers.
Deliverables
- Anomaly models
- Asset twin for pilot class
- Explainability views
- Engineering review sessions
- Phase 0404
Planner integration & scale
Connect recommendations to the CMMS, capture feedback and extend to further asset classes and sites.
Deliverables
- CMMS integration
- Feedback labeling
- Model monitoring
- Rollout playbook
07Intended outcomes
Design goals, not results.
Design goal 01Intended
Earlier, explained warnings
Degradation is surfaced with the component, the signal that changed and the baseline it changed from.
Design goal 02Intended
Planner-ready output
Recommendations arrive as ranked work requests in the existing CMMS, not as another dashboard to check.
Design goal 03Intended
A defensible OT boundary
The architecture adds visibility without adding any inbound path into plant networks.
Design goal 04Intended
Models that improve with use
Captured feedback builds the labeled history needed to move from anomaly detection toward failure prediction.
No figures are attached to these goals. Actual results depend on the environment, data and delivery, and would be measured against baselines agreed at the start of a real engagement.
08Lessons & risks
Engineering lessons and risks to manage.
Risk 01
Failure data is sparse
Well-maintained assets rarely fail, so supervised models lack examples. Expect to rely on anomaly detection for longer than planned.
Risk 02
Operating context is easy to miss
Without load and speed context, normal regime changes look like anomalies and erode trust quickly.
Risk 03
Sensor installation is engineering work
Mounting, cabling and calibration shape signal quality more than any modeling choice.
Risk 04
Alert fatigue returns if unmanaged
Ranking, deduplication and a clear owner for each recommendation matter as much as model accuracy.
Start an engineering conversation
Facing a similar engineering problem?
Bring us the problem and the constraints around it. We'll help architect the system.