3 papers
cs.CV2026
CVT-Bench: Probing Spatial-State Integrity through Counterfactual Viewpoint Transformations
Shanmukha Vellamcheti, Uday Kiran Kothapalli, Disharee Bhowmick +1
Multimodal large language models (MLLMs) perform strongly on isolated spatial tasks, but whether their predictions remain persistent and mutually coherent across viewpoints and com…
cs.CV2025
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
Shanmukha Vellamcheti, Sanjoy Kundu, Sathyanarayanan N. Aakur
Understanding relationships between objects is central to visual intelligence, with applications in embodied AI, assistive systems, and scene understanding. Yet, most visual relati…
cs.CV2025
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
Sanjoy Kundu, Shanmukha Vellamcheti, Sathyanarayanan N. Aakur
Open-world egocentric activity recognition poses a fundamental challenge due to its unconstrained nature, requiring models to infer unseen activities from an expansive, partially o…