activity
20232026
most citedVideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs

1 citations · 1 across the 7 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Conformalized Large Language Models under Configuration Shift

Yuqicheng Zhu, Jialin Yu, Lin Li +7

Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets…

cs.CV2025

ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay

Gengyuan Zhang, Mingcong Ding, Jingpei Wu +2

Embodied exploration is a target-driven process that requires embodied agents to possess fine-grained perception and knowledge-enhanced decision making. While recent attempts lever…

cs.CV2025

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction

Gengyuan Zhang, Tanveer Hannan, Hermine Kleiner +6

An ideal vision-language agent serves as a bridge between the human users and their surrounding physical world in real-world applications like autonomous driving and embodied agent…

cs.CL2025

Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models

Jian Lan, Udo Schlegel, Tanveer Hannan +3

Although Vision-Language Models (VLMs) have achieved remarkable success, the knowledge mechanisms underlying their social biases remain a black box, where fairness- and ethics-rela…

cs.CV2025

Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs

Gengyuan Zhang, Mingcong Ding, Tong Liu +2

Multimodal large language models (MLLMs) have demonstrated strong performance in understanding videos holistically, yet their ability to process streaming videos-videos are treated…

cs.CV2024

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

Roberto Amoroso, Gengyuan Zhang, Rajat Koner +3

Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on cont…