1 citations · 1 across the 4 of their papers we have counts for
9 papers
Your Embedding Model is SMARTer Than You Think
Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly +3
Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discar…
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Christopher Clark, Jieyu Zhang, Zixian Ma +18
Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from…
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
Jianrui Zhang, Yue Yang, Rohun Tripathi +5
Token pruning is essential for enhancing the computational efficiency of vision-language models (VLMs), particularly for video-based tasks where temporal redundancy is prevalent. P…
Reasoning-Augmented Representations for Multimodal Retrieval
Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han +3
Universal Multimodal Retrieval (UMR) seeks any-to-any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolvi…
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
Shubham Bharti, Shiyun Cheng, Jihyun Rho +5
We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizati…
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Mu Cai, Reuben Tan, Jianrui Zhang +12
Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benc…