Hierarchical Real-Time Gesture Interpretation of Bharatanatyam Mudras Using Bidirectional LSTM Networks

Authors:
Advika Seetharaman, S. Vaishnavii, Anika Seetharaman, Prasanna Ranjith Christodoss, Premanand Jothilingam, Karthikeyan Sivanandi

Addresses:
Department of Computer Science and Engineering, SRM Institute of Science and Technology, Ramapuram, Chennai, Tamil Nadu, India. Department of Computing, Mathematics and Physics, Messiah University, One University Ave, Mechanicsburg, Pennsylvania, United States of America. Department of Life Cycle Service, Yokogawa Corporation of America, West Valley City, Utah, United States of America. Department of Research and Development, InfraSecure AI LLC, Casper, Wyoming, United States of America.

Abstract:

Mudras, symbolic hand gestures, convey narrative and emotional content in Bharatanatyam, one of India's oldest classical dance genres. Sign language and human-computer interaction recognize gestures, but Bharatanatyam mudras are difficult to decipher. This work introduces the first hierarchical gesture recognition system that recognizes mudras and their meanings from video data. The MediaPipe Holistic architecture represents gestures as body and hand skeleton landmarks. The joint prediction uses a Bidirectional Long Short-Term Memory (BiLSTM) network with a common temporal backbone and two independent classification heads: a seven-class mudra head (Level 1) and a three-class semantic meaning head. Understanding that the Level 2 head anticipates one of up to three local indices of meaning defined within a mudra vocabulary rather than open-vocabulary semantic translation is crucial. The model was trained and tested on 2,400 gesture sequences from two practitioners, recorded under different settings, for 20 mudra–meaning pairs. On the held-out test set, BiLSTM scored 99.72% weighted F1 Score for mudra classification and 100.00% for semantic meaning prediction. The three-class meaning head distribution, tiny dataset, and two-performer constraint should be considered when interpreting the latter finding. These results beat a unidirectional LSTM model by 0.83 and 0.56 percentage points in mudra and meaning F1-scores, respectively. Hierarchical, context-aware gesture interpretation on common hardware supports dance teaching, cultural preservation, and human-computer interaction.

Keywords: Zero-Trust Architecture; Artificial Intelligence; Network Traffic Optimization; Deep Reinforcement Learning; Device Posture; End-User Behavior; App Sensitivity.

Received on: 29/05/2025, Revised on: 10/08/2025, Accepted on: 03/10/2025, Published on: 15/08/2026

DOI: 10.69888/FTSIN.2026.000734

FMDB Transactions on Sustainable Intelligent Networks, 2026 Vol. 3 No. 3, Pages: 137-150

  • Views : 35
  • Downloads : 6
Download PDF