activity
20242026
collaborators

16 papers

cs.CV2026

Visual prompt engineering for video models

Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performan…

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.AI2026

PreScience: A Dataset and Benchmark for Scientific Forecasting

Anirudh Ajith, Amanpreet Singh, Jay DeYoung +7

Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built a…

cs.CL2026

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

Yuling Gu, Oyvind Tafjord, Hyunwoo Kim +4

Large language models (LLMs) are increasingly tested for a "Theory of Mind" (ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at expli…

cs.CL2025

Neologism Learning for Controllability and Self-Verbalization

John Hewitt, Oyvind Tafjord, Robert Geirhos +1

Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introdu…

cs.CL2025

2 OLMo 2 Furious

Team OLMo, Pete Walsh, Luca Soldaini +40

We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully rele…