2 papers
cs.AI2026
Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino +3
Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the…
cs.AI2026
Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino +2
Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or…