collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

Hunar Batra, Haoqin Tu, Hardy Chen +3

Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on…

cs.CV2026

VisualClaw: A Real-Time, Personalized Agent for the Physical World

Haoqin Tu, Jianwen Chen, Zijun Wang +14

Vision language models are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cos…

cs.CV2026

Omni-MMSI: Toward Identity-attributed Social Interaction Understanding

Xinpeng Li, Bolin Lai, Hardy Chen +5

We introduce Omni-MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity-attr…

cs.CV2026

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Jiawei Mao, Hardy Chen, Haoqin Tu +7

Large vision-language models (LVLMs) have become increasingly strong but remain prone to hallucinations in multimodal tasks, which significantly narrows their deployment. As traini…

cs.CV2025

Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales

Zhaofang Qian, Hardy Chen, Zeyu Wang +9

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real li…

cs.CV2025

ViLBench: A Suite for Vision-Language Process Reward Modeling

Haoqin Tu, Weitao Feng, Hardy Chen +3

Process-supervised reward models serve as a fine-grained function that provides detailed step-wise feedback to model responses, facilitating effective selection of reasoning trajec…