7 papers
Self-Ensembling Vision-Language Models for Chart Data Extraction
Thomas Berkane, Qianyi Wang, Maimuna S. Majumder
Charts effectively convey quantitative information, but the underlying data are often locked in image form, hindering reuse and analysis. Manually digitizing charts is time-consumi…
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
Zeyu Wang, Chang Liu, Eduardus Tjitrahardja +22
Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remain…
OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer
Dixuan Lin, Yuxiang Zhang, Mengcheng Li +5
In this paper, we introduce OmniHands, a universal approach to recovering interactive hand meshes and their relative movement from monocular or multi-view inputs. Our approach addr…
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Wenbin An, Feng Tian, Sicong Leng +6
Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsisten…
Unleashing the Potential of Model Bias for Generalized Category Discovery
Wenbin An, Haonan Lin, Jiahao Nie +5
Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another la…
Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
Haonan Lin, Mengmeng Wang, Jiahao Wang +7
Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires…