3 papers
cs.LO2026
interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification
Vishak K Bhat, Prateek Chanda, Vijval Ekbote +6
Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification important for ensuring correctness. Existing approaches either verify o…
cs.CL2025
Characterizing Deep Research: A Benchmark and Formal Definition
Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi +6
Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} --…
cs.CV2024
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
Sina Malakouti, Aysan Aghazadeh, Ashmit Khandelwal +1
Vision language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability…