activity
20202026
most citedSelfPose: 3D Egocentric Pose Estimation from a Headset Mounted Camera

76 citations · 79 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang +4

Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly…

cs.CV2026

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

Rwiddhi Chakraborty, Yinong, Wang +7

Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The…

cs.CV2026

Racing in Volume with Flow Ensembles

Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle +4

Streaming 4D reconstruction has been demonstrated only indoors, on dense camera rigs surrounding subjects that move at human pace. Outdoor 4D reconstruction exists but relies eithe…

cs.CV2026

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produce…

cs.CV2024

Taming 3DGS: High-Quality Radiance Fields with Limited Resources

Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3

3D Gaussian Splatting (3DGS) has transformed novel-view synthesis with its fast, interpretable, and high-fidelity rendering. However, its resource requirements limit its usability.…

cs.CV2024★ 1 cited

Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models

Chen Wu, Fernando De la Torre

Text-to-image diffusion models have achieved remarkable performance in image synthesis, while the text interface does not always provide fine-grained control over certain image fac…