1 citations · 1 across the 5 of their papers we have counts for
10 papers · 1 filter
Can Language Models Learn to Listen?
Evonne Ng, Sanjay Subramanian, Dan Klein +3
We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the s…
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1
The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…
Frozen Forecasting: A Unified Evaluation
Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera +7
Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "co…
Forecasting Motion in the Wild
Neerja Thakkar, Shiry Ginosar, Jacob Walker +3
Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectori…
Pose Priors from Language Models
Sanjay Subramanian, Evonne Ng, Lea Müller +3
Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large…
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Eunice Yiu, Maan Qraitem, Anisa Noor Majhi +5
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from on…