4 papers
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
Pengyu Yan, Akhil Gorugantu, Mahesh Bhosale +3
Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-langu…
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi +3
Grounded multi-video question answering over real-world news events requires systems to surface query-relevant evidence across heterogeneous video archives while attributing every…
Artemis: Towards Referential Understanding in Complex Videos
Jihao Qiu, Yuan Zhang, Xi Tang +6
Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential un…
ChartReformer: Natural Language-Driven Chart Image Editing
Pengyu Yan, Mahesh Bhosale, Jay Lal +2
Chart visualizations are essential for data interpretation and communication; however, most charts are only accessible in image format and lack the corresponding data tables and su…