End-to-end machine-learning project that predicts machine failures from sensor readings, using the AI4I 2020 Predictive Maintenance dataset. The focus is a correct, leakage-free pipeline and honest evaluation on imbalanced data.
Given sensor measurements (temperatures, rotational speed, torque, tool wear) and the product quality type, the goal is to predict whether a machine will fail. Because failures are rare (~3.4% of records), the project treats this as an imbalanced binary classification problem and evaluates models with precision / recall, not accuracy.
- AI4I 2020 Predictive Maintenance Dataset — UCI Machine Learning Repository (10,000 rows, 14 columns).
- Target:
Machine failure(0 = healthy, 1 = failure). Failure rate ≈ 3.4%. - Link: https://archive.ics.uci.edu/dataset/601/ai4i+2020+predictive+maintenance+dataset
- Load & inspect — no missing values; checked class balance.
- Clean — dropped identifier columns (
UDI,Product ID) and encoded the categoricalType(L/M/H) with one-hot encoding. - Prevent data leakage — dropped the individual failure-mode columns
(
TWF,HDF,PWF,OSF,RNF). The target is derived from these, so using them as features would leak the answer. Only true sensor readings, known before a failure, are kept as features. - Split — stratified train / validation / test split (60 / 20 / 20) to keep the ~3.4% failure ratio in every partition.
- Scale —
StandardScalerfit on the training set only, then applied to validation and test. - Handle imbalance —
RandomOverSamplerapplied to the training set only, so validation and test still reflect the real-world failure rate. - Model & compare — Logistic Regression (baseline), KNN, and a small neural network (Keras). Model selection on the validation set.
- Final evaluation — the chosen model evaluated once on the held-out test set.
Validation set:
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Logistic Regression | 0.15 | 0.82 | 0.26 |
| KNN | 0.36 | 0.62 | 0.45 |
| Neural Network | 0.31 | 0.85 | 0.46 |
Held-out test set (Neural Network):
| Metric (failure class) | Value |
|---|---|
| Recall | 0.88 (caught 60 of 68 failures) |
| Precision | 0.34 |
| F1 | 0.49 |
Test performance closely matches validation, indicating the model generalizes rather than overfits.
This is predictive maintenance, where a missed failure is far more costly than a false alarm. Recall was therefore prioritized, and the neural network was chosen for the highest failure recall (0.88 on test, missing only 8 of 68 failures). If false alarms were the dominant cost, the higher-precision KNN would be preferable — the right model depends on the cost trade-off.
Python, pandas, NumPy, scikit-learn, imbalanced-learn, TensorFlow/Keras, Matplotlib, Seaborn.
- Download
ai4i2020.csvfrom the UCI link above. - Open the notebook (Google Colab or Jupyter) and upload the CSV.
- Run the cells top to bottom.
- The neural network still produces many false alarms (low precision); tuning the decision threshold or trying class-weighted loss could improve the balance.
- Only three models were compared; tree-based models (Random Forest, XGBoost) are common strong baselines for tabular data and would be a natural next step.
- Hyperparameters were kept simple; a systematic search could improve results.