From measurements to patients: data aggregation in supervised classification of X-ray diffraction datasets
Date published
Free to read from
Supervisor/s
Industry supervisor/s
Journal Title
Journal ISSN
Volume Title
Publisher
Department
Course name
Type
ISSN
Format
Citation
Abstract
Background/Objectives: Machine learning approaches are widely used in modern medical diagnostics, including cancer detection. The results can be significantly improved by aggregating individual measurements, and appropriate aggregation methods should be established. Methods: We applied various measurement aggregation strategies both before and after machine learning modeling to two datasets of X-ray diffraction images: human breast biopsy samples and canine claw samples. Two classifiers, Random Forest and Logistic Regression, were used to determine classification metrics: the area under the receiver operating characteristic curve (ROC-AUC) and balanced accuracy. Results: We found that all aggregation types improve classification metrics, with aggregation after modeling yielding better performance. Depending on the dataset and approach, either classifier can produce better results. For human breast samples, Random Forest with the logit aggregation strategy provides an ROC-AUC exceeding 0.9. For the canine dataset, both Random Forest with the logit aggregation strategy and Logistic Regression with the median of cancer probabilities achieve an ROC-AUC of about 0.85. Conclusions: We examined several simple, straightforward aggregation methods for patient diagnosis based on multiple measurements per patient and achieved significant improvements in classification metrics.
