Impact Factor
Call For Paper
Volume 12 Issue 09
September 2026
Author(s)
Abstract
Speech Emotion Recognition (SER) Allows Machines To Understand Human Emotional States From Their Speech, Enabling Applications In Human-computer Interaction, Call Center Analytics, And Mental Health Monitoring. Traditional Deep CNNs Achieve High Accuracy; However, They Are Costly When It Comes To Computation, Which Means They Cannot Be Used In Real Time On Edge Devices With Limited Computational Power. This Paper Proposes A Lightweight CNN-based Architecture For Real-time SER Using Log-Mel Spectrograms As Input Representations. The Proposed Model Incorporates Conv2D Layers, Batch Normalization, Max Pooling, And Global Average Pooling To Learn Discriminative Time-frequency Features While Reducing Model Complexity And Inference Overhead. Unlike The Existing Artificial Neural Network (ANN) Baseline, Which Was Trained Using Only The RAVDESS Dataset, The Proposed CNN Is Trained Using A Combined RAVDESS And TESS Dataset Covering Eight Emotion Classes: Neutral, Calm, Happy, Sad, Angry, Fearful, Disgust, And Surprised. Experimental Evaluation Shows That TheproposedLightweightCNNattained75.70%testaccuracywithaweightedF1-scoreof0.78, Compared To The ANN Baseline Model’s 54.86% Accuracyand0.55F1-score. Additionally, Compared To The ANN Baseline, The Proposed Lightweight CNN Architecture Demonstrated Faster Real-time Performance Using Microphone Input While Providing Confidences Cores For Each Emotion Detected. Overall, The Proposed Lightweight CNN Model Provides Faster Inference With Good Accuracy For SER Applications.
Keywords
Paper ID
IJSARTV12I9105855
Publication Date
September 3, 2026
Research Area
Electronics And Communication Engineering