3 papers
cs.LG2026
Multimodal Latent Reasoning via Predictive Embeddings
Ashutosh Adhikari, Mirella Lapata
Tool-augmented multimodal reasoning enables visual language models (VLMs) to improve perception by interacting with external tools (e.g., cropping, depth estimation). However, such…
cs.AI2025
Debating for Better Reasoning: An Unsupervised Multimodal Approach
Ashutosh Adhikari, Mirella Lapata
As Large Language Models (LLMs) gain expertise across diverse domains and modalities, scalable oversight becomes increasingly challenging, particularly when their capabilities may…
cs.CL2024
MCPDial: A Minecraft Persona-driven Dialogue Dataset
Seyed Hossein Alavi, Sudha Rao, Ashutosh Adhikari +7
We propose a novel approach that uses large language models (LLMs) to generate persona-driven conversations between Players and Non-Player Characters (NPC) in games. Showcasing the…