6 papers
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…
EvalCards: A Framework for Standardized Evaluation Reporting
Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11
Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…
What if Othello-Playing Language Models Could See?
Xinyi Chen, Yifei Yuan, Jiaang Li +3
Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…
Evaluation of Cultural Competence of Vision-Language Models
Srishti Yadav, Lauren Tilton, Maria Antoniak +10
Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in…
Leveraging GCN-based Action Recognition for Teleoperation in Daily Activity Assistance
Thomas M. Kwok, Jiaan Li, Yue Hu
Caregiving of older adults is an urgent global challenge, with many older adults preferring to age in place rather than enter residential care. However, providing adequate home-bas…
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
Lei Li, Sen Jia, Jianhao Wang +4
Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking…