From the 1 of 11 linked papers with an AI index.
11 papers
Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
Junrui Zhang, Jiaqi Li, Yiran Wang +2
The paper introduces CapDepth, a framework that uses detailed long textual captions to guide monocular depth estimation, improving robustness on non‑Lambertian surfaces and in adve…
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
Liao Shen, Wentao Jiang, Yiran Zhu +4
Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the appl…
BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching
Yachuan Huang, Xianrui Luo, Qiwen Wang +6
Bokeh rendering simulates the shallow depth-of-field effect in photography, enhancing visual aesthetics and guiding viewer attention to regions of interest. Although recent approac…
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
Huiqiang Sun, Liao Shen, Zhan Peng +9
Cinematic storytelling is profoundly shaped by the artful manipulation of photographic elements such as depth of field and exposure. These effects are crucial in conveying mood and…
MuGS: Multi-Baseline Generalizable Gaussian Splatting Reconstruction
Yaopeng Lou, Liao Shen, Tianqi Liu +4
We present Multi-Baseline Gaussian Splatting (MuGS), a generalized feed-forward approach for novel view synthesis that effectively handles diverse baseline settings, including spar…
Dynamic View Synthesis from Small Camera Motion Videos
Huiqiang Sun, Xingyi Li, Juewen Peng +4
Novel view synthesis for dynamic D scenes poses a significant challenge. Many notable efforts use NeRF-based approaches to address this task and yield impressive results. Howeve…