activity
20242026
collaborators

9 papers

cs.CV2026

A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding

Yue Zhang, Liqiang Jing, Jia Li +4

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approa…

cs.CV2026

Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos

Rohith Peddi, Saurabh, Shravan Shanmugam +4

Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason onl…

cs.AI2026

Learning to Guide Local Search for MPE Inference in Probabilistic Graphical Models

Brij Malhotra, Shivvrat Arya, Tahrima Rahman +1

Most Probable Explanation (MPE) inference in Probabilistic Graphical Models (PGMs) is a fundamental yet computationally challenging problem arising in domains such as diagnosis, pl…

cs.CV2025

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

Yue Zhang, Jilei Sun, Yunhui Guo +1

Video Large Multimodal Models (VLMMs) have made impressive strides in understanding video content, but they often struggle with abstract and adaptive reasoning-the ability to revis…

cs.AI2025

LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning

Navapat Nananukul, Yue Zhang, Ryan Lee +5

High-assurance reasoning, particularly in critical domains such as law and medicine, requires conclusions that are accurate, verifiable, and explicitly grounded in evidence. This r…

cs.LG2025

Learning to Condition: A Neural Heuristic for Scalable MPE Inference

Brij Malhotra, Shivvrat Arya, Tahrima Rahman +1

We introduce learning to condition (L2C), a scalable, data-driven framework for accelerating Most Probable Explanation (MPE) inference in Probabilistic Graphical Models (PGMs), a f…