About Research Stack Experience Education Work Contact contact@enricogaraiman.com
← BACK / ACADEMIC

Integrated emotion recognition system from multimodal data

Multimodal fusion of facial video and speech into a single emotion recognition pipeline, using a window-based approach on the audio signal and a limited number of video frames.

4 images · click any to open the gallery

People express and read emotion through several channels at once, so a system that only listens, or only looks, is working with half the signal. That matters for anything from monitoring behavioural and emotional disorders to public safety.

The system recognises emotion from audio and video together. The audio branch uses a ResNet architecture with an x-vector built on self-attention, working on Mel-frequency cepstral coefficients. The visual branch is a convolutional network over frames containing faces detected with MTCNN. Fusion combines both channels across several temporal segments, rather than a single frame or a single window.

Trained on CREMA-D and RAVDESS. On CREMA-D the accuracy reached 63.05% for audio, 64.09% for visual and 76.29% for the two combined; on RAVDESS, 62%, 55.67% and 65.67%. The gap between single-channel and multimodal is the whole point of the work.

One finding worth stating plainly: past a certain number of segments, response time keeps growing and accuracy does not. Finding that optimum is part of the design, not an afterthought.

The best model was deployed on an NVIDIA Jetson Nano for inference, to check that the approach survives outside a workstation.

CONTACT

Let's get
in touch.

Full-time roles, freelance builds or research collaboration. Write and I will answer, usually within a day.

contact@enricogaraiman.com BUCHAREST · ROMANIA · EET (UTC+2)
01 / 01
100%