4 papers
AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
Omar Moured, Jiaming Zhang, M. Saquib Sarfraz +1
Chart summarization is a crucial task for blind and visually impaired individuals as it is their primary means of accessing and interpreting graphical data. Crafting high-quality d…
Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems
Qihao Yuan, Kailai Li, Jiaming Zhang
3D visual grounding (3DVG) aims to locate objects in a 3D scene with natural language descriptions. Supervised methods have achieved decent accuracy, but have a closed vocabulary a…
Deformable Mamba for Wide Field of View Segmentation
Jie Hu, Junwei Zheng, Jiale Wei +2
Recent advancements in the Mamba architecture, with its linear computational complexity, being a promising alternative to transformer architectures suffering from quadratic complex…
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
Xin Jiang, Junwei Zheng, Ruiping Liu +4
As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of pe…