11 papers
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
Ali Vosoughi, Ayoub Shahnazari, Yufeng Xi +4
We introduce OPENXRD, a comprehensive benchmarking framework for evaluating large language models (LLMs) and multimodal LLMs (MLLMs) in crystallography question answering. The fram…
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
Yolo Y. Tang, Jing Bi, Pinxin Liu +24
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
Yolo Y. Tang, Pinxin Liu, Zhangyun Tan +11
Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains uncle…
Video Understanding with Large Language Models: A Survey
Yolo Y. Tang, Jing Bi, Siting Xu +17
With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper
Integrating audio and visual data for training multimodal foundational models remains a challenge. The Audio-Video Vector Alignment (AVVA) framework addresses this by considering A…
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Ali Vosoughi, Yongyi Zang, Qihui Yang +3
Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: t…