공공데이터를 활용한 건설사고 텍스트 분석: 내러티브 단계와 사고 중증도에 따른 위험 표현 탐색

Construction Accident Text Analysis Using Public Data: Exploring Risk Expressions Based on Narrative Stages and Accident Severity

초록

This study explores risk patterns in construction safety narratives by analyzing unstructured text of 30,546 accident reports from the Korea Authority of Land & Infrastructure Safety (KALIS) database. Focusing on three textual fields such as accident process, damage description, and post-accident measures, we conduct a field-separated analysis to capture how risks are narrated across stages of the accident. Using Python and KoNLPy (Okt), we perform data preprocessing that includes term normalization, stop-word removal, and stemming. Then, we compute TF-IDF weights to identify salient terms and summarize them via word clouds. In addition, we apply Latent Dirichlet Allocation (LDA) with six topics and bag-of-words input to each ‘field × severity-label (injury or fatality)’ combination and then try to uncover latent co-occurrence-based topic structures. The results show that injury cases are characterized by everyday hazards such as tripping, slipping, and material handling, along with medical descriptions related to fractures and joint injuries. In contrast, fatal cases are dominated by high-risk work contexts, i.e., falls from height, demolition, heavy equipment, collapse, and burial, and by institutional response narratives that involve reporting, investigation, settlement, and work stoppage. The contributions of this study are threefold: i) it proposes an exploratory analytic framework that structures public construction-accident text data by narrative stage (process damage measures) and by severity (injury-fatality); ii) it visualizes how key risk terms and topic structures differ across narrative stages and severity levels using word clouds and LDA; and iii) it offers the extracted risk terms and topic groups as a domain-specific risk lexicon and a set of text feature candidates for future predictive models and the design of training and warning systems.

키워드

Construction SafetyNarrative AnalysisWord CloudTopic ModelingLatent Dirichlet Allocation
제목
공공데이터를 활용한 건설사고 텍스트 분석: 내러티브 단계와 사고 중증도에 따른 위험 표현 탐색
제목 (타언어)
Construction Accident Text Analysis Using Public Data: Exploring Risk Expressions Based on Narrative Stages and Accident Severity
저자
박수연박세지유동희서종환
발행일
2026-06
유형
Y
저널명
인터넷전자상거래연구
26
3
페이지
103 ~ 121