10 citations · 52 across the 32 of their papers we have counts for
49 papers
PuzzleMate: Benchmarking MLLMs for Egocentric Puzzle Assistance
Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar +2
Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical activities. For these assista…
Flow Matching in Feature Space for Stochastic World Modeling
Francois Porcher, Nicolas Carion, Karteek Alahari +1
World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world models often struggle to satisfy both goa…
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
Yedidia Agnimo, Anna Korba, Annabelle Blangero +2
Large language models (LLMs) are prone to hallucinations, i.e., statements unsupported by the input or training data, hindering reliable deployment. In parallel, numerous uncertain…
Exploring High-Order Self-Similarity for Video Understanding
Manjin Kim, Heeseung Kwon, Karteek Alahari +1
Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this wo…
Flowception: Temporally Expansive Flow Matching for Video Generation
Tariq Berrada Ifriqi, John Nguyen, Karteek Alahari +2
We present Flowception, a novel non-autoregressive and variable-length video generation framework. Flowception learns a probability path that interleaves discrete frame insertions…
Online In-Context Distillation for Low-Resource Vision Language Models
Zhiqi Kang, Rahaf Aljundi, Vaggelis Dorovatas +1
As the field continues its push for ever more resources, this work turns the spotlight on a critical question: how can vision-language models (VLMs) be adapted to thrive in low-res…