3 papers
cs.CV2026
GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos
Deepak Kumar, Abhishek Pratap Singh, Puneet Kumar +2
Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from i…
cs.MM2025
Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data
Puneet Kumar, Sarthak Malik, Balasubramanian Raman +1
The ability to generate sentiment-controlled feedback in response to multimodal inputs comprising text and images addresses a critical gap in human-computer interaction. This capab…
cs.CV2025
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
Puneet Kumar, Sarthak Malik, Balasubramanian Raman +1
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and te…