collaborators

6 papers

cs.CV2026

Visual prompt engineering for video models

Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performan…

cs.AI2026

Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies

Sitao Cheng, Xunjian Yin, Ruiwen Zhou +5

Does Reinforcement Learning (RL) merely amplify existing skills, or synthesize novel skills? We investigate this question through the lens of Complementary Reasoning: the critical…

cs.CV2025

Accident Anticipation via Temporal Occurrence Prediction

Tianhao Zhao, Yiyang Zou, Zihao Mao +7

Accident anticipation aims to predict potential collisions in an online manner, enabling timely alerts to enhance road safety. Existing methods typically predict frame-level risk s…

cs.LG2025

Video models are zero-shot learners and reasoners

Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6

The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models.…

cs.AI2025

How well can LLMs provide planning feedback in grounded environments?

Yuxuan Li, Victor Zhong

Learning to plan in grounded environments typically requires carefully designed reward functions or high-quality annotated demonstrations. Recent works show that pretrained foundat…

cs.CV2025

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

Yuxuan Li, Vijay Veerabadran, Michael L. Iuzzolino +3

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice…