14 citations · 23 across the 22 of their papers we have counts for
25 papers
Video-FLAIR: Not Whether to Reason, But How
Yogesh Kulkarni, Pooyan Fazli
Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others req…
VisionPulse: A Virtual Reality System Enabling Accessible Discovery and Navigation for Blind and Low Vision Users
Samuel Martin, Pooyan Fazli, Hasti Seifi
Free exploration is an important aspect of many engaging virtual reality (VR) experiences, yet remains largely inaccessible to blind and low vision (BLV) users due to its reliance…
ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
Maryam Cheema, Sina Elahimanesh, Pooyan Fazli +1
Advances in multimodal large language models enable automatic video narration and question answering (VQA), offering scalable alternatives to labor-intensive, human-authored audio…
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
Chaoyu Li, Fei Tao, Pooyan Fazli
Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sam…
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
Yogesh Kulkarni, Pooyan Fazli
Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) stru…
FrameOracle: Learning What to See and How Much to See in Videos
Chaoyu Li, Tianzhi Li, Fei Tao +6
Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of f…