76 citations · 79 across the 6 of their papers we have counts for
8 papers · 1 filter
OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction
Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang +4
Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly…
Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models
Rwiddhi Chakraborty, Yinong, Wang +7
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The…
Racing in Volume with Flow Ensembles
Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle +4
Streaming 4D reconstruction has been demonstrated only indoors, on dense camera rigs surrounding subjects that move at human pace. Outdoor 4D reconstruction exists but relies eithe…
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produce…
Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3
3D Gaussian Splatting (3DGS) has transformed novel-view synthesis with its fast, interpretable, and high-fidelity rendering. However, its resource requirements limit its usability.…
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
Chen Wu, Fernando De la Torre
Text-to-image diffusion models have achieved remarkable performance in image synthesis, while the text interface does not always provide fine-grained control over certain image fac…