Multi-level violence recognition via hybrid convolutional-attention and recurrent architectures

Nature作者:Manoj Kumar2026年8月12日正文已收录本站
  • Article
  • Open access
  • Published:
  • Birendra Kumar Verma2,
  • Sukhendra Singh2,
  • Abhay Kumar3,
  • Kumar Abhishek3 &
  • …
  • B. M. Ahamed Shafeeq4 

Scientific Reports (2026) Cite this article

We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.

Abstract

Violence recognition is an urgent need for real-time human activity identification in surveillance video streams, which is becoming more important in areas including public safety, law enforcement, and security monitoring. Even while visual understanding has come a long way, finding a balance between accuracy and speed of identification is still a big problem. To address these constraints, we suggest a hybrid deep learning architecture that combines a Time-Distributed Convolutional Neural Network (TD-CNN) with a Long Short-Term Memory (LSTM) network augmented by a spatial attention mechanism. The attention module, which is located between convolutional layers, adaptively highlights important spatial areas, which makes it easier to tell the features apart. The LSTM part of the model looks at how frames are related to each other across time to capture motion dynamics that are important to violent behaviors. The suggested architecture was tested on the Hockey Fight Dataset and got an accuracy of 93%, which is better than many other baselines. Experimental findings indicate that the hierarchical integration of convolutional and recurrent layers with attention significantly improves recognition accuracy while maintaining a computationally lightweight architecture suitable for near-real-time deployment.

Subjects

Funding

Open access funding provided by Manipal Academy of Higher Education, Manipal.

Author information

Authors and Affiliations

  1. School of Computer Science Engineering and Technology, Bennett University, Greater Noida, India

    Manoj Kumar

  2. JSS Academy of Technical Education Noida, Noida, UP, India

    Birendra Kumar Verma & Sukhendra Singh

  3. Dept of CSE, NIT Patna, Patna, India

    Abhay Kumar & Kumar Abhishek

  4. Manipal institute of technology, Manipal Academy of Higher Education, Manipal, India

    B. M. Ahamed Shafeeq

Authors

  1. Manoj Kumar
  2. Birendra Kumar Verma
  3. Sukhendra Singh
  4. Abhay Kumar
  5. Kumar Abhishek
  6. B. M. Ahamed Shafeeq

Corresponding author

Correspondence to B. M. Ahamed Shafeeq.

Ethics declarations

Competing interests

The authors declare no competing interests.

Ethics declaration

Not applicable.

Consent to participate

Not applicable.

Consent to publish

Not applicable.

Additional information

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Kumar, M., Verma, B.K., Singh, S. et al. Multi-level violence recognition via hybrid convolutional-attention and recurrent architectures. Sci Rep (2026). https://doi.org/10.1038/s41598-026-66068-1

Download citation

  • Received:

  • Accepted:

  • Published:

  • DOI: https://doi.org/10.1038/s41598-026-66068-1

Keywords