3 citations · 3 across the 4 of their papers we have counts for
10 papers
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
Yolo Y. Tang, Jing Bi, Pinxin Liu +24
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Ali Vosoughi, Yongyi Zang, Qihui Yang +3
Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: t…
Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
Jing Bi, Guangyu Sun, Ali Vosoughi +2
Multimodal large language models (MLLMs) that integrate visual and textual reasoning leverage chain-of-thought (CoT) prompting to tackle complex visual tasks, yet continue to exhib…
Can Sound Replace Vision in LLaVA With Token Substitution?
Ali Vosoughi, Jing Bi, Pinxin Liu +2
What happens when we push audio-visual alignment to its absolute limits? To systematically investigate this question, we needed datasets with granular alignment quality annotations…
: Generating Instructional Illustrations via Text-Conditioned Diffusion
Jing Bi, Pinxin Liu, Ali Vosoughi +3
The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey comple…
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
Yolo Y. Tang, Pinxin Liu, Zhangyun Tan +11
Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains uncle…