7 papers
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
Yu Guo, Zhiqiang Lao, Xiyun Song +2
Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited phys…
Inverse Rendering for High-Genus Surface Meshes from Multi-View Images
Xiang Gao, Xinmu Wang, Xiaolong Wu +8
We present a topology-informed inverse rendering approach for reconstructing high-genus surface meshes from multi-view images. Compared to 3D representations like voxels and point…
Neural Geometry Image-Based Representations with Optimal Transport (OT)
Xiang Gao, Yuanpeng Liu, Xinmu Wang +7
Neural representations for 3D meshes are emerging as an effective solution for compact storage and efficient processing. Existing methods often rely on neural overfitting, where a…
OT-Talk: Animating 3D Talking Head with Optimal Transportation
Xinmu Wang, Xiang Gao, Xiyun Song +4
Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech s…
ePBR: Extended PBR Materials in Image Synthesis
Yu Guo, Zhiqiang Lao, Xiyun Song +3
Realistic indoor or outdoor image synthesis is a core challenge in computer vision and graphics. The learning-based approach is easy to use but lacks physical consistency, while tr…
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Sakib Reza, Xiyun Song, Heather Yu +3
Integrating vision models into large language models (LLMs) has sparked significant interest in creating vision-language foundation models, especially for video understanding. Rece…