2 papers
cs.AI2025
VaPR -- Vision-language Preference alignment for Reasoning
Rohan Wadhawan, Fabrice Y Harel-Canada, Zi-Yi Dou +3
Preference finetuning methods like Direct Preference Optimization (DPO) with AI-generated feedback have shown promise in aligning Large Vision-Language Models (LVLMs) with human pr…
cs.LG2025
The Curious Language Model: Strategic Test-Time Information Acquisition
Michael Cooper, Rohan Wadhawan, John Michael Giorgi +2
Decision-makers often possess insufficient information to render a confident decision. In these cases, the decision-maker can often undertake actions to acquire the necessary infor…