8 papers
Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting
Amritansh Maurya, Navjot Singh, Mohammed Javed +1
Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, because Table Question-Answering…
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
Adriana Aida, Walid Amer, Katarina Bankovic +25
Industrial robotic manipulation demands reliable long-horizon execution across embodiments, tasks, and changing object distributions. While Vision-Language-Action models have demon…
AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
Omar Moured, Jiaming Zhang, M. Saquib Sarfraz +1
Chart summarization is a crucial task for blind and visually impaired individuals as it is their primary means of accessing and interpreting graphical data. Crafting high-quality d…
HybriDLA: Hybrid Generation for Document Layout Analysis
Yufan Chen, Omar Moured, Ruiping Liu +4
Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for ea…
TY-RIST: Tactical YOLO Tricks for Real-time Infrared Small Target Detection
Abdulkarim Atrash, Omar Moured, Yufan Chen +3
Infrared small target detection (IRSTD) is critical for defense and surveillance but remains challenging due to (1) target loss from minimal features, (2) false alarms in cluttered…
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
Alexander Vogel, Omar Moured, Yufan Chen +2
Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understandi…