Enhancing safety of construction workers in Korea : an integrated text mining and machine learning framework for predicting accident types

Construction workers face a high risk of various occupational accidents, many of which can result in fatalities. This study aims to develop a prediction model for nine prevalent types of construction accidents, utilizing construction tasks, activities, and tools/materials as input features, through the application of machine learning-based multi-class classification algorithms. 152,867 construction accident summary reports, composed of both structured (construction task, construction activity, accident type) and unstructured data (tools/materials) were used for the study. The study employed several data processing techniques, including keyword extraction through text mining, Boruta feature selection, and SMOTE data resampling enhance model accuracy. Three performance metrics (Multi-class area under the receiver operating characteristic curve (MAUC), Multi-class Matthews Correlation Coefficient (MMCC), Geometric-mean (G-mean)) were used to compare the predictive performance of four machine learning algorithms, including Decision tree, Random forest, Naïve bayes, and XGBoost. Of the four algorithms, XGBoost showed the highest performance in predicting accident type (MAUC: 0.8603, MMCC: 0.3523, G-mean: 0.5009). Furthermore, a Shapley additive explanation (SHAP) analysis was conducted to visualize feature importance. The findings of this study make a valuable contribution to improving construction safety by presenting a prediction model for accident types derived from real-world big data.

Medienart:

E-Artikel

Erscheinungsjahr:

2024

Erschienen:

2024

Enthalten in:

Zur Gesamtaufnahme - year:2024

Enthalten in:

International journal of injury control and safety promotion - (2024) vom: 02. Jan., Seite 1-13

Sprache:

Englisch

Beteiligte Personen:

Yoo, Joon Woo [VerfasserIn]
Park, Junsung [VerfasserIn]
Park, Heejun [VerfasserIn]

Links:

Volltext

Themen:

Accident prediction
Journal Article
Machine learning
Multi-class classification
SHAP value
Text mining

Anmerkungen:

Date Revised 02.01.2024

published: Print-Electronic

Citation Status Publisher

doi:

10.1080/17457300.2023.2300424

funding:

Förderinstitution / Projekttitel:

PPN (Katalog-ID):

NLM36655154X