Informatica logo


Login Register

  1. Home
  2. To appear
  3. An Efficient Cascade Neural Network for ...

Informatica

Information Submit your article For Referees Help ATTENTION!
  • Article info
  • Full article
  • Related articles
  • More
    Article info Full article Related articles

An Efficient Cascade Neural Network for Human Action Recognition
Jianying Xiong   Zikang Fan   Ning Liu   Keyun Xiong   Leiyue Yao  

Authors

 
Placeholder
https://doi.org/10.15388/26-INFOR645
Pub. online: 2 September 2026      Type: Research Article      Open accessOpen Access

Received
1 January 2026
Accepted
1 August 2026
Published
2 September 2026

Abstract

Calculating motion features frame by frame and organizing them into a 3D matrix is a typical CNN-based solution for human action recognition (HAR). With the widespread use of consumer electronics, reducing computational costs and enabling efficient edge-side action recognition have become a research hotspot. In this paper, we extract action key frames via a well-designed algorithm to reduce computational overhead, so that the proposed method can be deployed on mobile electronic devices. Then we construct local and global motion features from these key frames and feed them into a cascade neural network for action recognition. The primary contributions include three aspects. First, the strategic adoption of key frames is introduced to greatly reduce the number of input parameters. The number of key frames can be adjusted to adapt to the temporal scales of different actions. Second, multiple origin points are adopted to construct motion matrices with larger dimensions than those constructed using a single origin point. Thus, deeper neural networks can be employed to achieve higher recognition accuracy. Third, a cascade neural network is proposed for action prediction, which leverages global and local information to achieve better efficiency and accuracy. Experimental results on UTKinect-Action3D, Florence-3D and our self-built HanYue-3D datasets demonstrate that our method achieves accuracy and efficiency competitive with state-of-the-art (SOTA) approaches. Moreover, the flexibility of the proposed method enables users to readily balance effectiveness and efficiency, making it well-suited for resource-constrained mobile devices.

References

 
Abdelbaky, A., Aly, S. (2020). Human action recognition using short-time motion energy template images and PCANet features. Neural Computing and Applications, 32(16), 12561–12574.
 
Arif, S., Wang, J., Siddiqui, A.A., Hussain, R., Hussain, F. (2021). Bidirectional LSTM with saliency-aware 3D-CNN features for human action recognition. Journal of Engineering Research, 9(3), 115–133.
 
Cheng, K., Zhang, Y., He, X., Chen, W., Cheng, J., Lu, H. (2020). Skeleton-based action recognition with shift graph convolutional network. In: Conference on Computer Vision and Pattern Recognition, pp. 180–189.
 
Dalal, N., Triggs, B. (2005). Histograms of oriented gradients for human detection. In: IEEE Conference on Computer Vision and Pattern Recognition, San Diego, 2005, pp. 886–893.
 
Dong, W., Zhang, Z., Song, C., Tan, T. (2022). Identifying the key frames: an attention-aware sampling method for action recognition. Pattern Recognition, 130, 108797.
 
Du, Y., Fu, Y., Wang, L. (2015). Skeleton based action recognition with convolutional neural network. In: IEEE Asian Conference on Pattern Recognition, pp. 579–583.
 
Giveki, D. (2024). Human action recognition using an optical flow-gated recurrent neural network. International Journal of Multimedia Information Retrieval, 13(3), 29.
 
Herath, S., Harandi, M., Porikli, F. (2017). Going deeper into action recognition: a survey. Image and Vision Computing, 60, 4–21.
 
Hu, Z., Xiao, J., Li, L., Liu, C., Ji, G. (2024). Human-centric multimodal fusion network for robust action recognition. Expert Systems with Applications, 239, 122314.
 
Hussain, A., Khan, S.U., Khan, N., Bhatt, M.W., Farouk, A., Bhola, J., Baik, S.W. (2024). A hybrid transformer framework for efficient activity recognition using consumer electronics. IEEE Transactions on Consumer Electronics, 70(4), 6800–6807.
 
Jain, V., Gupta, G., Gupta, M., Sharma, D.K., Ghosh, U. (2023). Ambient intelligence-based multimodal human action recognition for autonomous systems. ISA Transactions, 132, 94–108.
 
Jeyanthi, A., Visumathi, J., Genitha, C.H. (2024). Enhanced two-stream Bayesian hyper parameter optimized 3D-CNN inception-v3 based drop-convLSTM2D deep learning model for human action recognition. Information Technology and Control, 53(1), 53–70.
 
Karim, M., Khalid, S., Aleryani, A., Khan, J., Ullah, I., Ali, Z. (2024). Human action recognition systems: a review of the trends and state-of-the-art. IEEE Access, 12, 36372–36390.
 
Kurban, O.C., Yildirim, T. (2024). A comparative analysis of multi-biometrics performance in human and action recognition using silhouette thermal-face and skeletal data. Neural Networks: The Official Journal of the International Neural Network Society, 170, 1–17.
 
Laptev, I. (2025). On space-time interest points. Computer Vision, 64, 107–123.
 
Le, T.M., Inoue, N., Shinoda, K. (2018). A fine-to-coarse convolutional neural network for 3D human action recognition. arXiv preprint. arXiv:1805.11790.
 
Li, C., Hou, Y., Wang, P., Li, W. (2018). Multiview-based 3-D action recognition using deep networks. IEEE Transactions on Human-Machine Systems, 49(1), 95–104.
 
Li, X., Kang, J., Yang, Y., Zhao, F. (2023). A lightweight attentional shift graph convolutional network for skeleton-based action recognition. International Journal of Computers Communications & Control, 18(3).
 
Majd, M., Safabakhsh, R. (2020). Correlational convolutional LSTM for human action recognition. Neurocomputing, 396, 224–229.
 
Phyo, C.N., Zin, T.T., Tin, P. (2019). Deep learning for recognizing human activities using motions of skeletal joints. IEEE Transactions on Consumer Electronics, 65(2), 243–252.
 
Seidenari, L., Varano, V., Berretti, S., Del Bimbo, A., Pala, P. (2013). Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses. In: IEEE Conference on Computer Vision and Pattern Recognition, pp. 479–485.
 
Shi, L., Zhang, Y., Cheng, J., Lu, H. (2019). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In: Conference on Computer Vision and Pattern Recognition, pp. 12018–12027.
 
Sun, Z., Ke, Q., Rahmani, H., Bennamoun, M., Wang, G., Liu, J. (2022). Human action recognition from various data modalities: a review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3), 3200–3225.
 
Tan, K.S., Lim, K.M., Lee, C.P., Kwek, L.C. (2022). Bidirectional long short-term memory with temporal dense sampling for human action recognition. Expert Systems with Applications, 210, 118484.
 
Ullah, A., Muhammad, K., Del Ser, J., Baik, S.W., de Albuquerque, V.H.C. (2019). Activity recognition using temporal optical flow convolutional features and multilayer LSTM. IEEE Transactions on Industrial Electronics, 66, 9692–9702.
 
Wang, L., Yan, Y., Huang, D., Pan, Y., Cheang, C.F., Luo, K., Li, J. (2025). AHARNet: adaptive human activity recognition model for multimodal consumer electronics with different computation resources. IEEE Transactions on Consumer Electronics, 71(2), 5847–5855.
 
Wang, Z., Shen, K., Wang, D., Shen, H., Huang, K. (2024). Human body parsing in thermal InfraRed domain. IEEE Transactions on Consumer Electronics, 70(4), 6420–6429.
 
Wang, Y., Feng, T., Zheng, Y. (2022). Human action recognition using a depth sequence key-frames based on discriminative collaborative representation classifier for healthcare analytics. Computer Science and Information Systems, 19(3), 1445–1462.
 
Xia, L., Xin, W. (2024). Multi-stream network with key frame sampling for human action recognition. Journal of Supercomputing, 80, 11958–11988.
 
Xia, L., Chen, C.C., Aggarwal, J.K. (2012). View invariant human action recognition using histograms of 3D joints. In: IEEE Conference on Computer Vision and Pattern Recognition, pp. 20–27.
 
Xie, Q.L., Lu, W., Yang, W., Xiong, K., Zhang, L., Yao, L. (2025). Recognizing a complex human behaviour via a shallow neural network with zero video training sample. International Journal of Computers Communications & Control, 20(5), 1–11.
 
Xin, C., Kim, S., Cho, Y., Park, K.S. (2024). Enhancing human action recognition with 3D skeleton data: a comprehensive study of deep learning and data augmentation. Electronics, 13(4), 747.
 
Yan, S., Xiong, Y., Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. In: AAAI Conference on Artificial Intelligence, pp. 7444–7452.
 
Yang, W., Zhou, Y.T., Xiong, J.Y., Zhang, S., Zhang, L., Yao, L. (2025). Human action recognition using explainable features and sparse motion history images. Technical Gazette Tehnički Vjesnik, 32(5), 1614–1623.
 
Yang, C., Mei, F., Zang, T., Tu, J., Jiang, N., Liu, L. (2023). Human action recognition using key-frame attention-based LSTM networks. Electronics, 12(12), 2622.
 
Yang, Z., Li, Y., Yang, J., Luo, J. (2018). Action recognition with spatio-temporal visual attention on skeleton image sequences. IEEE Transactions on Circuits and Systems for Video Technology, 29(8), 2405–2415.
 
Yao, L.Y., Yang, W., Huang, W. (2020). A data augmentation method for human action recognition using dense joint motion images. Applied Soft Computing, 97, 106713–106723.
 
Yao, L.Y., Yang, W., Huang, W., Jiang, N., Zhou, B.B. (2021). Multi-scale feature learning and temporal probing strategy for one-stage temporal action localization. International Journal of Intelligent Systems, 12(1), 1–21.
 
Yu, J., Cheng, X., Chen, H., Xu, Y. (2024). Pose-guided robust action recognition for outdoor internet of things. IEEE Transactions on Consumer Electronics, 17(4), 7032–7043.
 
Zhang, D. (2019). ATSN: attention-based temporal segment network for action recognition. Technical Gazette Tehnički Vjesnik, 2(26), 1664–1669.
 
Zhang, S., Chen, E., Qi, C., Liang, C. (2016). Action recognition based on sub-action motion history image and static history image. MATEC Web of Conferences, 56, 02006.
 
Zhang, Y., You, S., Karaoglu, S., Gevers, T. (2025). 3D human pose estimation and action recognition using fisheye cameras: a survey and benchmark. Pattern Recognition, 162, 111334.
 
Zhou, S., Xu, H., Bai, Z., Du, Z., Zeng, J., Wang, Y., Wang, Y., Li, S., Wang, M., Li, Y., Li, J., Xu, J. (2023). A multidimensional feature fusion network based on MGSE and TAAC for video-based human action recognition. Neural Networks, 168, 496–507.
 
Zhou, Y., Cheng, Z.Q., Li, C., Fang, Y., Geng, Y., Xie, X., Keuper, M. (2023). Hypergraph transformer for skeleton-based action recognition. arXiv preprint. arXiv:2211.09590.
 
Zhang, Y., You, S., Karaoglu, S., Gevers, T. (2025). 3D human pose estimation and action recognition using fisheye cameras: a survey and benchmark. Pattern Recognition, 162, 111334.
 
Zhang, Y., Zhao, B., Wang, Y. (2026). HML-STN: high-middle-low spatio-temporal network for RGB-D based human action recognition. Signal, Image and Video Processing, 20(3), 179.

Biographies

Xiong Jianying

J. Xiong is an associate professor in the field of computer science. She received the ME degree from Zhejiang University of Technology, China, in 2006, and the PhD degree from Jiangxi University of Finance and Economics, China, in 2013. Her research interests include information systems, information management, and service computing.

Fan Zikang

Z. Fan received the bachelor’s degree from East China University of Technology in 2025, and is currently pursuing the master’s degree at Jiangxi University of Chinese Medicine. His research interests include computer vision, deep learning, multi-modal medical image analysis, and intelligent diagnosis of traditional Chinese medicine tongue.

Liu Ning

N. Liu is the director, general manager and secretary of the board, Beijing Hanlin Hangyu Technology Development Inc. He has long been engaged in the R&D, engineering and industrialization of intelligent manufacturing equipment for traditional Chinese medicine solid preparations. His research focuses on intelligent production lines, digital factory construction and real-time monitoring systems for pharmaceutical manufacturing. He leads industrial research projects on the upgrading of domestic pharmaceutical equipment and promotes industry-university-research cooperation for intelligent transformation of Chinese medicine manufacturing.

Xiong Keyun

K. Xiong was born in September 1980 in Nanchang, China. He received the BE degree and is currently a lecturer. His research interests mainly include big data architecture and data mining.

Yao Leiyue
leiyue_yao@163.com

L. Yao received the BE, ME, and PhD degrees in computer science from Nanchang University, China. He is currently a professor at the School of Intelligent Medicine and Information Engineering, Jiangxi University of Chinese Medicine. He has published several papers in international journals and conferences. His current research interests include vision-based human action recognition, image and video processing, massive data processing, distributed systems, and software engineering.


Full article Related articles PDF XML
Full article Related articles PDF XML

Copyright
© 2026 Vilnius University
by logo by logo
Open access article under the CC BY license.

Keywords
key frame extraction cascade neural network multiple origin points skeleton-based action recognition CNN-based action recognition.

Funding
This research was supported by the National Natural Science Foundation of China under Grant 62366023.

Metrics
since January 2020
131

Article info
views

43

Full article
views

37

PDF
downloads

18

XML
downloads

Export citation

Copy and paste formatted citation
Placeholder

Download citation in file


Share


RSS

INFORMATICA

  • Online ISSN: 1822-8844
  • Print ISSN: 0868-4952
  • Copyright © 2023 Vilnius University

About

  • About journal

For contributors

  • OA Policy
  • Submit your article
  • Instructions for Referees
    •  

    •  

Contact us

  • Institute of Data Science and Digital Technologies
  • Vilnius University

    Akademijos St. 4

    08412 Vilnius, Lithuania

    Phone: (+370 5) 2109 338

    E-mail: informatica@mii.vu.lt

    https://informatica.vu.lt/journal/INFORMATICA
Powered by PubliMill  •  Privacy policy