10 citations · 12 across the 10 of their papers we have counts for
7 papers · 1 filter
Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs
Yang Yang, Jiawei Chen, Tairan Chen +1
Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image…
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
Zichuan Lin, Yicheng Liu, Yang Yang +2
Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant compu…
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
Yang Xing, Jiong Wu, Savas Ozdemir +4
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (…
MeshRipple: Structured Autoregressive Generation of Artist-Meshes
Junkai Lin, Hang Long, Huipeng Guo +8
Meshes serve as a primary representation for 3D assets. Autoregressive mesh generators serialize faces into sequences and train on truncated segments with sliding-window inference…
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
Qirui Yang, Yang Yang, Ying Zeng +5
Text-guided diffusion models have greatly advanced image editing and generation. However, achieving physically consistent image retouching with precise parameter control (e.g., exp…
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
Chao Yuan, Yang Yang, Yehui Yang +1
Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localiza…