TY - GEN
T1 - ASTFormer: A Spatio-Temporal Transformer with Dual Attention and CNN Fusion for Video Anomaly Detection
AU - Sun, Yi
AU - Scotney, Bryan W.
AU - Nie, Xiushan
AU - Zhang, Shuai
AU - Liu, Xingbo
AU - Qiu, Lanting
PY - 2025/1/23
Y1 - 2025/1/23
N2 - With the rapid advancement of video surveillance technology, achieving robust anomaly detection remains a critical challenge. Traditional approaches often focus on local spatial features while overlooking temporal dynamics, which are essential for recognizing sudden or complex anomalies. To address this limitation, we propose a novel attention-guided spatiotemporal prediction-based reconstruction named AST Former, specifically designed for video anomaly detection. AST Former integrates a deep convolutional encoder based on Wider Res Net with a unified fusion module that leverages multi-scale temporal self-attention, spatial self-attention across frames, and context-aware cross-attention to comprehensively capture spatiotemporal dependencies. In addition, lightweight transformer blocks are incorporated at each resolution stage of the decoder to enable end-to-end spatial structure modeling and to enhance the representation and discrimination of the up sampled features. Extensive experiments on three public datasets-UeSD Ped2, Avenue and Shanghai Tech-show that AST Former achieves AUC values of 98.02%, 86.80%, and 74.33%, respectively, outperforming many existing approaches and confirming the effectiveness of our spatiotemporal attention mechanism in detecting anomalies.
AB - With the rapid advancement of video surveillance technology, achieving robust anomaly detection remains a critical challenge. Traditional approaches often focus on local spatial features while overlooking temporal dynamics, which are essential for recognizing sudden or complex anomalies. To address this limitation, we propose a novel attention-guided spatiotemporal prediction-based reconstruction named AST Former, specifically designed for video anomaly detection. AST Former integrates a deep convolutional encoder based on Wider Res Net with a unified fusion module that leverages multi-scale temporal self-attention, spatial self-attention across frames, and context-aware cross-attention to comprehensively capture spatiotemporal dependencies. In addition, lightweight transformer blocks are incorporated at each resolution stage of the decoder to enable end-to-end spatial structure modeling and to enhance the representation and discrimination of the up sampled features. Extensive experiments on three public datasets-UeSD Ped2, Avenue and Shanghai Tech-show that AST Former achieves AUC values of 98.02%, 86.80%, and 74.33%, respectively, outperforming many existing approaches and confirming the effectiveness of our spatiotemporal attention mechanism in detecting anomalies.
KW - video anomaly detection
KW - spatio-temporal modeling
KW - self-attention
KW - feature reconstruction
U2 - 10.1109/cpsi66656.2025.11343933
DO - 10.1109/cpsi66656.2025.11343933
M3 - Conference contribution
SN - 979-8-3315-9962-1
T3 - 2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
SP - 1
EP - 6
BT - 2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
PB - IEEE
T2 - 2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
Y2 - 7 November 2025 through 10 November 2025
ER -