
Unsafe conditions of objects are major contributors to construction accidents such as falls, strikes, and collapses. Although traditional inspections and existing vision methods can detect objects, they often fail to infer their potential risks. To overcome this limitation, this paper proposes a deep identification method integrating visual attribute representations. We introduce the high-level concept of visual unsafe attributes to describe hazards arising from an object’s shape, state, and environmental context. Based on accident text analysis, a visual attribute system covering four risk types—fall, strike, collapse, and rollover—is established, and the VALD dataset is built by extending SODA dataset. An integrated detection framework based on YOLOv11 is then developed to achieve end-to-end joint inference of object locations and unsafe attributes. The model is trained and evaluated on VALD. Experiments show mAP@50 of 84.8%, recall of 91%, and F1-score of 0.82, with improvements of 21.1%, 33.6%, and 0.175 over models without attention mechanisms. These results demonstrate that visual attribute representation and attention significantly enhance risk feature extraction. The resulting dynamic risk assessment system quantifies object-strike scenarios spatiotemporally, providing a real-time and interpretable safety monitoring solution for construction sites.
construction safety management; visual attribute; computer vision; deep learning