From the 1 of 11 linked papers with an AI index.
11 papers
Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
Junrui Zhang, Jiaqi Li, Yiran Wang +2
The paper introduces CapDepth, a framework that uses detailed long textual captions to guide monocular depth estimation, improving robustness on non‑Lambertian surfaces and in adve…
Prisma-World: Camera-Controllable Multi-Agent Video World Model
Huiqiang Sun, Zhan Peng, Size Wu +9
Video world models have made rapid progress in generating controllable visual experiences, but most of them still simulate the world from a single observer. Extending such models t…
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
Liao Shen, Wentao Jiang, Yiran Zhu +4
Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the appl…
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
Zihao Huang, Tianqi Liu, Zhaoxi Chen +7
Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage…
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
Tianqi Liu, Zhaoxi Chen, Zihao Huang +8
Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighti…
BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching
Yachuan Huang, Xianrui Luo, Qiwen Wang +6
Bokeh rendering simulates the shallow depth-of-field effect in photography, enhancing visual aesthetics and guiding viewer attention to regions of interest. Although recent approac…