14 papers
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs
Wenzhang Sun, Chunfeng Wang, Xiangchen Yin +3
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level.…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camer…
Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos
Yanan Liu, Qinya Li, Hao Zhang +5
Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and temporal discontinuity. To tac…
RePCM: Region-Specific and Phenotype-Adaptive Bi-Ventricular Cardiac Motion Synthesis
Xuan Yang, Xiaohan Yuan, Hao Li +3
Cardiac motion over a cardiac cycle is crucial for quantifying regional function and is strongly affected by cardiovascular diseases. Since temporally dense mesh sequences are diff…
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
Zhangchi Hu, Wenzhang Sun, Xiangchen Yin +5
Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded…
Lance: Unified Multimodal Modeling by Multi-Task Synergy
Fengyi Fu, Mengqi Huang, Shaojin Wu +10
We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity…