activity
20242026
most citedSora as a World Model? A Complete Survey on Text-to-Video Generation

10 citations · 12 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Yang Yang, Jiawei Chen, Tairan Chen +1

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image…

cs.CV2026

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

Zichuan Lin, Yicheng Liu, Yang Yang +2

Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant compu…

cs.CV2026

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

Yang Xing, Jiong Wu, Savas Ozdemir +4

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (…

cs.CV2025

MeshRipple: Structured Autoregressive Generation of Artist-Meshes

Junkai Lin, Hang Long, Huipeng Guo +8

Meshes serve as a primary representation for 3D assets. Autoregressive mesh generators serialize faces into sequences and train on truncated segments with sliding-window inference…

cs.CV2025

CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching

Qirui Yang, Yang Yang, Ying Zeng +5

Text-guided diffusion models have greatly advanced image editing and generation. However, achieving physically consistent image retouching with precise parameter control (e.g., exp…

cs.CV2025

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

Chao Yuan, Yang Yang, Yehui Yang +1

Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localiza…