activity
20242026
most citedCan Language Models Learn to Listen?

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20261 cited

Can Language Models Learn to Listen?

Evonne Ng, Sanjay Subramanian, Dan Klein +3

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the s…

cs.CV2026

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1

The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…

cs.CV2026

Frozen Forecasting: A Unified Evaluation

Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera +7

Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "co…

cs.CV2026

Forecasting Motion in the Wild

Neerja Thakkar, Shiry Ginosar, Jacob Walker +3

Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectori…

cs.CV2025

Pose Priors from Language Models

Sanjay Subramanian, Evonne Ng, Lea Müller +3

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large…

cs.CV2025

KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models

Eunice Yiu, Maan Qraitem, Anisa Noor Majhi +5

This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from on…