most citedCan Language Models Learn to Listen?

1 citations · 1 across the 2 of their papers we have counts for

collaborators

9 papers

cs.CV20261 cited

Can Language Models Learn to Listen?

Evonne Ng, Sanjay Subramanian, Dan Klein +3

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the s…

cs.CV2026

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1

The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…

cs.CV2026

Frozen Forecasting: A Unified Evaluation

Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera +7

Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "co…

cs.AI2026

Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions

Eunice Yiu, Kelsey Allen, Shiry Ginosar +1

Learning about the causal structure of the world is a fundamental problem for human cognition. Causal models and especially causal learning have proved to be difficult for large pr…

cs.CV2026

Forecasting Motion in the Wild

Neerja Thakkar, Shiry Ginosar, Jacob Walker +3

Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectori…

cs.CV2025

Pose Priors from Language Models

Sanjay Subramanian, Evonne Ng, Lea Müller +3

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large…