9 papers
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
Yue Zhang, Liqiang Jing, Jia Li +4
Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approa…
Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
Rohith Peddi, Saurabh, Shravan Shanmugam +4
Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason onl…
Learning to Guide Local Search for MPE Inference in Probabilistic Graphical Models
Brij Malhotra, Shivvrat Arya, Tahrima Rahman +1
Most Probable Explanation (MPE) inference in Probabilistic Graphical Models (PGMs) is a fundamental yet computationally challenging problem arising in domains such as diagnosis, pl…
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
Yue Zhang, Jilei Sun, Yunhui Guo +1
Video Large Multimodal Models (VLMMs) have made impressive strides in understanding video content, but they often struggle with abstract and adaptive reasoning-the ability to revis…
LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning
Navapat Nananukul, Yue Zhang, Ryan Lee +5
High-assurance reasoning, particularly in critical domains such as law and medicine, requires conclusions that are accurate, verifiable, and explicitly grounded in evidence. This r…
Learning to Condition: A Neural Heuristic for Scalable MPE Inference
Brij Malhotra, Shivvrat Arya, Tahrima Rahman +1
We introduce learning to condition (L2C), a scalable, data-driven framework for accelerating Most Probable Explanation (MPE) inference in Probabilistic Graphical Models (PGMs), a f…