collaborators

5 papers

cs.IR2026

Multimodal Generative Recommendation for Fusing Semantic and Collaborative Signals

Moritz Vandenhirtz, Kaveh Hassani, Shervin Ghasemlou +5

Sequential recommender systems rank relevant items by modeling a user's interaction history and computing the inner product between the resulting user representation and stored ite…

cs.CV2025

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur +2

Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings in…

cs.CV2025

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

Ce Zhang, Yale Song, Ruta Desai +4

Visual Planning for Assistance (VPA) aims to predict a sequence of user actions required to achieve a specified goal based on a video showing the user's progress. Although recent a…

cs.CV2025

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

Yuxuan Li, Vijay Veerabadran, Michael L. Iuzzolino +3

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice…

cs.HC2025

Less or More: Towards Glanceable Explanations for LLM Recommendations Using Ultra-Small Devices

Xinru Wang, Mengjie Yu, Hannah Nguyen +10

Large Language Models (LLMs) have shown remarkable potential in recommending everyday actions as personal AI assistants, while Explainable AI (XAI) techniques are being increasingl…