4 papers · 1 filter
Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models
Rwiddhi Chakraborty, Yinong, Wang +7
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The…
From Blurry to Believable: Enhancing Low-quality Talking Heads with 3D Generative Priors
Ding-Jiun Huang, Yuanhao Wang, Shao-Ji Yuan +4
Creating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield p…
Visual Data Diagnosis and Debiasing with Concept Graphs
Rwiddhi Chakraborty, Yinong Wang, Jialu Gao +3
The widespread success of deep learning models today is owed to the curation of extensive datasets significant in size and complexity. However, such models frequently pick up inher…
Generalizable Human Gaussians for Sparse View Synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu +9
Recent progress in neural rendering has brought forth pioneering methods, such as NeRF and Gaussian Splatting, which revolutionize view rendering across various domains like AR/VR,…