Skip to main navigation Skip to search Skip to main content

ASTFormer: A Spatio-Temporal Transformer with Dual Attention and CNN Fusion for Video Anomaly Detection

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With the rapid advancement of video surveillance technology, achieving robust anomaly detection remains a critical challenge. Traditional approaches often focus on local spatial features while overlooking temporal dynamics, which are essential for recognizing sudden or complex anomalies. To address this limitation, we propose a novel attention-guided spatiotemporal prediction-based reconstruction named AST Former, specifically designed for video anomaly detection. AST Former integrates a deep convolutional encoder based on Wider Res Net with a unified fusion module that leverages multi-scale temporal self-attention, spatial self-attention across frames, and context-aware cross-attention to comprehensively capture spatiotemporal dependencies. In addition, lightweight transformer blocks are incorporated at each resolution stage of the decoder to enable end-to-end spatial structure modeling and to enhance the representation and discrimination of the up sampled features. Extensive experiments on three public datasets-UeSD Ped2, Avenue and Shanghai Tech-show that AST Former achieves AUC values of 98.02%, 86.80%, and 74.33%, respectively, outperforming many existing approaches and confirming the effectiveness of our spatiotemporal attention mechanism in detecting anomalies.
Original languageEnglish
Title of host publication2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
PublisherIEEE
Pages1-6
Number of pages6
ISBN (Electronic)979-8-3315-9961-4
ISBN (Print)979-8-3315-9962-1
DOIs
Publication statusPublished online - 23 Jan 2025
Event2025 International Conference on Cyber-Physical Social Intelligence (CPSI) - Macau, China
Duration: 7 Nov 202510 Nov 2025

Publication series

Name2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
PublisherIEEE Control Society

Conference

Conference2025 International Conference on Cyber-Physical Social Intelligence (CPSI)
Country/TerritoryChina
CityMacau
Period7/11/2510/11/25

Keywords

  • video anomaly detection
  • spatio-temporal modeling
  • self-attention
  • feature reconstruction

Fingerprint

Dive into the research topics of 'ASTFormer: A Spatio-Temporal Transformer with Dual Attention and CNN Fusion for Video Anomaly Detection'. Together they form a unique fingerprint.

Cite this