works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

Chaoyu Li, Fei Tao, Pooyan Fazli

The paper proposes CASHEW, an inference-time method that stabilizes multi-step reasoning in vision-language models by aggregating multiple reasoning trajectories with visual verifi…

cs.CV2026

FrameOracle: Learning What to See and How Much to See in Videos

Chaoyu Li, Tianzhi Li, Fei Tao +6

Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of f…

cs.CV2026

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs

Chaoyu Li, Yogesh Kulkarni, Pooyan Fazli

The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing efficiency methods mainly target inference v…

cs.CV2025

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. Howeve…

cs.CV2025

VideoA11y: Method and Dataset for Accessible Video Description

Chaoyu Li, Sid Padmanabhuni, Maryam Cheema +2

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall…