activity
20242026
most citedMolmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

1 citations · 1 across the 4 of their papers we have counts for

collaborators

9 papers

cs.IR2026

Your Embedding Model is SMARTer Than You Think

Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly +3

Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discar…

cs.CV20261 cited

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Christopher Clark, Jieyu Zhang, Zixian Ma +18

Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from…

cs.CV2026

Unified Spatio-Temporal Token Scoring for Efficient Video VLMs

Jianrui Zhang, Yue Yang, Rohun Tripathi +5

Token pruning is essential for enhancing the computational efficiency of vision-language models (VLMs), particularly for video-based tasks where temporal redundancy is prevalent. P…

cs.IR2026

Reasoning-Augmented Representations for Multimodal Retrieval

Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han +3

Universal Multimodal Retrieval (UMR) seeks any-to-any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolvi…

cs.AI2025

CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts

Shubham Bharti, Shiyun Cheng, Jihyun Rho +5

We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizati…

cs.CV2024

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Mu Cai, Reuben Tan, Jianrui Zhang +12

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benc…