13 papers
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
Suryaansh Jain, Rahasya Barkur, Vishal G +8
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss t…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in…
A Survey on LLM-based Conversational User Simulation
Bo Ni, Leyao Wang, Yu Wang +27
User simulation has long played a vital role in computer science due to its potential to support a wide range of applications. Language, as the primary medium of human communicatio…
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
Yifan Li, Seunghyun Yoon, Viet Dac Lai +4
Vision Transformers (ViTs) based vision foundation models (VFMs) have achieved remarkable performance across diverse vision tasks, but suffer from quadratic complexity that limits…
Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
Jiyeon Kim, Hyunji Lee, Dylan Zhou +6
LLMs operating in dynamic real-world contexts often encounter knowledge that evolves continuously or emerges incrementally. To remain accurate and effective, models must adapt to n…
Agentic Planning with Reasoning for Image Styling via Offline RL
Subhojyoti Mukherjee, Stefano Petrangeli, Branislav Kveton +3
Direct prompt-based editing often fails on complex transformations because vague and subjective prompts often require nuanced understanding of what should be changed in the image.…