A Multimodal Transformer-Based Digital Companion for Emotion-Aware Human–Computer Interaction Using Real-Time 3D Avatars

Authors:
A. Shyam Sundhar, J. Angelin Jeba, Varsha Ramkumar, M. Rehena Sulthana, Rejwan Bin Sulaiman, R. P. P. Ranaweera

Addresses:
Department of Information and Communication, Tampere University, Tampere, Pirkanmaa, Finland. Department of Electronics and Communication Engineering, S.A. Engineering College, Chennai, Tamil Nadu, India. Department of Business Analytics, University of Galway, Galway, Ireland. School of Information Technology and Engineering, Melbourne Institute of Technology, Melbourne, Victoria, Australia. Department of Computer Science and Technology, Northumbria University, London, England, United Kingdom. National Institute of Library and Information Sciences, University of Colombo, Colombo, Sri Lanka.

Abstract:

Human-computer interaction has progressed significantly in recent years, yet existing digital assistants continue to suffer from limitations in emotional engagement, personalisation, and expressive communication. Most current AI systems operate primarily through text or voice, lacking a visual presence that allows users to form meaningful connections with technology. To address this gap, our paper introduces JARWIN (Just A Real-World Intelligent Network). This next-generation AI companion integrates a large language model, speech processing, and real-time 3D avatar animation to create a more natural, interactive, and intelligent user experience. The architecture features multiple operational tiers, including speech-to-text transcription, sentiment-aware conversational modelling, text-to-speech synthesis, and avatar-based emotional expression. Linear conversational modelling serves as the initial baseline for contextual understanding, while more advanced transformer-based models are employed to handle nuanced dialogue, sentiment variations, and continuous memory retention. Additionally, emotion recognition mechanisms dynamically map conversational tone to facial expressions and gestures, enabling human-like responsiveness. The result is a system capable not only of executing functional tasks, such as opening applications and retrieving information, but also of maintaining adaptive, emotionally aware dialogue. By bridging cognitive intelligence with visual embodiment, this paper demonstrates a scalable, innovative approach to immersive digital companionship, offering new opportunities for personalised assistance, interactive learning, mental wellness support, and human-AI relational computing.

Keywords: Machine Learning (ML); Resilient Food Chain; Economic Loss; Linear Regression; Random Forest; Digital Companionship; Prediction Accuracy; Interactive Learning.

Received on: 30/07/2025, Revised on: 25/10/2025, Accepted on: 02/11/2025, Published on: 09/08/2026

DOI: 10.69888/FTSCL.2026.000752

FMDB Transactions on Sustainable Computer Letters, 2026 Vol. 4 No. 3, Pages: 178-193

  • Views : 26
  • Downloads : 9
Download PDF