Industrial Machine Learning Under Limited Data
Predicting ore-grinding output size and throughput where the available data was far thinner than the modelling problem deserved. CRISP-DM kept the work anchored to a business question and honest about what the data could support.
RoleEnd-to-end data science: framing, data understanding, modelling, evaluation

Context and problem
Ore grinding sits upstream of everything else in a mineral processing circuit. Output particle size and throughput drive downstream recovery and plant economics, and both respond to inputs that are difficult to control and expensive to measure.
The question was whether a predictive model could usefully estimate output size and throughput. First, we needed to establish whether the available data could support such a model at all.
What made it difficult
- Limited data relative to the complexity of the physical process
- Industrial measurements that are noisy and irregularly sampled
- Operating conditions that shift over time, so historical data is not uniformly representative
- A need for results interpretable by process engineers, not only by data scientists
Role and responsibilities
- Framed the business question and translated it into a modelling problem
- Led data understanding and preparation under real-world quality constraints
- Developed Python-based machine-learning models
- Evaluated results honestly against what the data could support
- Helped advance the work toward execution
Generalized architecture
011
Business understanding
Establish what decision the prediction would inform, and what accuracy would make it useful.
022
Data understanding
Assess coverage, quality, and sampling of the available process data before committing to an approach.
033
Data preparation
Clean, align, and engineer features from noisy, irregularly sampled industrial measurements.
044
Modelling
Python-based machine-learning models targeting output size and throughput.
055
Evaluation
Test against the decision the model was meant to support, not only against a held-out score.
Flow: The work followed the CRISP-DM cycle: business understanding, data understanding, data preparation, modelling, and evaluation, iterating between stages as the data revealed what was feasible.
Key decisions and tradeoffs
CRISP-DM as the frame, not a formality
With limited data, a model can score well without answering the operational question. Starting with business and data understanding gave the modelling stage a clear objective and helped explain what the data could not support.
Interpretability weighted against raw accuracy
A model that process engineers can reason about gets used; a marginally more accurate black box gets ignored. Under limited data the accuracy gap was small enough that interpretability was the better trade.
Treating data scarcity as a finding, not an obstacle
Reporting clearly on where the data was too thin to support a conclusion was part of the deliverable. It is less satisfying than a strong result, and considerably more useful to the people deciding what to instrument next.
Implementation
- Worked through the CRISP-DM cycle end to end, iterating between data understanding and modelling as constraints surfaced
- Prepared and engineered features from noisy industrial process measurements
- Developed Python-based machine-learning models to predict output size and throughput for the ore-grinding context
- Evaluated results against the operational question and reported honestly on the limits of the available data
- Helped advance the work toward execution
Results
- Python-based machine-learning models developed for ore-grinding output size and throughput prediction
- End-to-end data-science work delivered under limited-data constraints using CRISP-DM
- Work advanced toward execution
Lessons and next iteration
- Under limited data, the most valuable output is often a clear account of what the data can and cannot support.
- Process engineers need to understand and trust a model before they will use it on the plant floor.
- Next iteration: specify the instrumentation and sampling needed up front, so the data-understanding stage constrains the project less on the following attempt.
Technologies
Python · Scikit-learn · Pandas · NumPy · CRISP-DM