2 papers
cs.CV2026
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Junho Kim, Xu Cao, Houze Yang +6
Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with who…
cs.CV2026
How to Choose Your Teacher for Fine Grained Image Recognition
Oswin Gosal, Edwin Arkel Rios, Augusto Christian Surya +3
Fine-grained image recognition classifies subcategories such as bird species or car models. While state-of-the-art (SOTA) models are accurate, they are often too resource-intensive…