collaborators

6 papers

cs.CL2026

Source-Modality Monitoring in Vision-Language Models

Etha Tianze Hua, Tian Yun, Ellie Pavlick

We define and investigate source-modality monitoring -- the ability of multimodal models to track and communicate the input source from which pieces of information originate. We co…

cs.CV2025

Video Finetuning Improves Reasoning Between Frames

Ruiqi Yang, Tian Yun, Zihan Wang +1

Multimodal large language models (LLMs) have made rapid progress in visual understanding, yet their extension from images to videos often reduces to a naive concatenation of frame…

cs.CL2025

$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources

Apoorv Khandelwal, Tian Yun, Nihal V. Nayak +4

Pre-training is notoriously compute-intensive and academic researchers are notoriously under-resourced. It is, therefore, commonly assumed that academics can't pre-train models. In…

cs.CL2025

What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models

Tian Yun, Chen Sun, Ellie Pavlick

Recent work has argued that large language models (LLMs) are not "abstract reasoners", citing their poor zero-shot performance on a variety of challenging tasks as evidence. We rev…

cs.CL2025

How Do Vision-Language Models Process Conflicting Information Across Modalities?

Tianze Hua, Tian Yun, Ellie Pavlick

AI models are increasingly required to be multimodal, integrating disparate input streams into a coherent state representation on which subsequent behaviors and actions can be base…

cs.CV2025

Pre-trained Vision-Language Models Learn Discoverable Visual Concepts

Yuan Zang, Tian Yun, Hao Tan +2

Do vision-language models (VLMs) pre-trained to caption an image of a "durian" learn visual concepts such as "brown" (color) and "spiky" (texture) at the same time? We aim to answe…