collaborators

5 papers

cs.CV2025

VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery

Bojin Wu, Jing Chen

Monocular depth estimation can be broadly categorized into two directions: relative depth estimation, which predicts normalized or inverse depth without absolute scale, and metric…

cs.HC2025

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi +12

Chart-to-code reconstruction -- the task of recovering executable plotting scripts from chart images -- provides important insights into a model's ability to ground data visualizat…

cs.CV2025

R^3-VQA: "Read the Room" by Video Social Reasoning

Lixing Niu, Jiapeng Li, Xingping Yu +6

"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks an…

cs.CV2025

SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation

Hao Du, Bo Wu, Yan Lu +1

Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-…

cs.CV2025

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60

We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…