activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA

Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri +4

Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer, rather than onl…

cs.CL2026

Structured Uncertainty guided Clarification for LLM Agents

Manan Suri, Puneet Mathur, Nedim Lipka +3

LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and task failures. Existing approaches ope…

cs.CL2025

Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents

Manan Suri, Puneet Mathur, Nedim Lipka +4

Flowcharts are a critical tool for visualizing decision-making processes. However, their non-linear structure and complex visual-textual relationships make it challenging to interp…

cs.CL2025

ChartLens: Fine-grained Visual Attribution in Charts

Manan Suri, Puneet Mathur, Nedim Lipka +3

The growing capabilities of multimodal large language models (MLLMs) have advanced tasks like chart understanding. However, these models often suffer from hallucinations, where gen…

cs.CL2025

VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation

Manan Suri, Puneet Mathur, Franck Dernoncourt +3

Understanding information from a collection of multiple documents, particularly those with visually rich elements, is important for document-grounded question answering. This paper…

cs.CL2024

DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding

Manan Suri, Puneet Mathur, Franck Dernoncourt +5

Document structure editing involves manipulating localized textual, visual, and layout components in document images based on the user's requests. Past works have shown that multim…