Journal of Information Technology in Construction
ITcon Vol. 31, pg. 1067-1095, http://www.itcon.org/2026/45
Toward trustworthy construction safety hazard detection with visual-language models
| DOI: | 10.36680/j.itcon.2026.045 | |
| submitted: | May 2026 | |
| published: | September 2026 | |
| editor(s): | Bosché F | |
| authors: | Mina Sadat Orooje, PhD candidate
Department of Architecture, Construction Engineering and Built Environment, Politecnico di Milano, Italy https://orcid.org/0009-0002-4131-982X minasadat.orooje@polimi.it Fulvio Re Cecconi, Associate Professor Department of Architecture, Construction Engineering and Built Environment, Politecnico di Milano, Italy https://orcid.org/0000-0001-7716-8854 fulvio.rececconi@polimi.it Gaang Lee, Dr.-Ing., Assistant Professor Department of Civil and Environmental Engineering, University of Alberta, Edmonton, Alberta, Canada https://orcid.org/0000-0002-6341-2585 gaang@ualberta.ca Qipei (Gavin) Mei, PhD, PEng, Assistant Professor Department of Civil and Environmental Engineering, University of Alberta, Edmonton, Alberta, Canada https://orcid.org/0000-0003-1409-3562 qipei@ualberta.ca Muhammad Adil, M.Sc., Research Assistant Department of Civil and Environmental Engineering, University of Alberta, Canada https://orcid.org/0009-0009-1346-6479 madil2@ualberta.ca | |
| summary: | The construction industry continues to experience high rates of accidents and fatalities, underscoring the need for proactive and reliable safety management. Accurate hazard identification is essential for effective image-based monitoring and decision support on construction sites. Vision–language models (VLMs) have shown strong potential for interpreting complex visual environments; however, their deployment in safety-critical applications is limited by hallucination, where hazards are inferred without sufficient evidence. To address this limitation, this study proposes a retrieval-augmented hazard detection framework that grounds VLM outputs in visual evidence through image-to-image retrieval of visually similar hazard scenarios and structured domain knowledge within a retrieval-augmented generation (RAG) pipeline. The framework integrates image-to-image retrieval of real-world hazard scenarios, supported by a curated knowledge base of 1,040 annotated construction site images, structured hazard taxonomies, and regulation-aware verification to enforce condition-constrained hazard reasoning before compliance assessment. The proposed approach is evaluated on a separate set of 260 construction site images spanning ten construction-safety hazard categories aligned with OSHA safety domains and the construction-safety literature. At the end-to-end report level, the complete pipeline achieves an F1-score of 0.917 (precision 0.939, recall 0.896) and a hallucination rate of 6.1%, substantially lower than the VLM-only baseline. At the hazard-verification gate, adding taxonomy-constrained reasoning to RAG-1 yields a precision of 0.987 and an F1-score of 0.941, while RAG-1 alone achieves an F1-score of 0.923 with 89.8% retrieval coverage. These results indicate that visual grounding provides the largest reliability gain, with structured condition verification adding further precision and gated regulatory retrieval providing traceable compliance support after hazard confirmation. | |
| keywords: | Vision–Language Models (VLMs), Retrieval-Augmented Generation (RAG), Construction Safety, Hallucination Mitigation, Hazard Detection | |
| full text: | (PDF file, 1.241 MB) | |
| citation: | Orooje, M. S., Re Cecconi, F., Lee, G., Mei, Q., & Adil, M. (2026). Toward trustworthy construction safety hazard detection with visual-language models. Journal of Information Technology in Construction (ITcon), 31, 1067-1095. https://doi.org/10.36680/j.itcon.2026.045 | |
| statistics: |



