activity
20172026
most citedA Countrywide Traffic Accident Dataset

55 citations · 138 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2025

ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Dheeraj Kulshrestha +1

Document Visual Question Answering (VQA) requires models to not only extract accurate textual answers but also precisely localize them within document images, a capability critical…

cs.CV2025

MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Dheeraj Kulshrestha +1

Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spa…

cs.CV2025

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA

Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Ser-Nam Lim +1

Document visual question answering requires models not only to answer questions correctly, but also to precisely localize answers within complex document layouts. While large visio…

cs.CV2025

TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation

Amin Karimi Monsefi, Mridul Khurana, Rajiv Ramnath +3

We propose TaxaDiffusion, a taxonomy-informed training framework for diffusion models to generate fine-grained animal images with high morphological and identity accuracy. Unlike s…

cs.CV2024

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim +1

Document Visual Question Answering (VQA) demands robust integration of text detection, recognition, and spatial reasoning to interpret complex document layouts. In this work, we in…

cs.CV20241 cited

KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models

Pouyan Navard, Amin Karimi Monsefi, Mengxi Zhou +3

Recent advances in diffusion models have significantly improved text-to-image (T2I) generation, but they often struggle to balance fine-grained precision with high-level control. M…