multimodal question answering 1policy optimization 1post-training 1reinforcement learning 1visual grounding 1
From the 1 of 3 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment
Harikrishnan P M, Goutham Vignesh, Ganesh Parab +4
The paper proposes Perception-RFT, a post‑training framework that uses Group Relative Policy Optimization to align visual features with grounding outputs for multimodal document qu…
cs.AI2026
Search-Based Risk Feature Discovery in Document Structure Spaces under a Constrained Budget
Saisubramaniam Gopalakrishnan, Harikrishnan P M, Dagnachew Birru
Enterprise-grade Intelligent Document Processing (IDP) systems support high-stakes workflows across finance, insurance, and healthcare. Early-phase system validation under limited…