3 papers
cs.CV2026
Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks
Jae Joong Lee
Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B paramete…
cs.CV2025
Language-Guided Invariance Probing of Vision-Language Models
Jae Joong Lee
Recent vision-language models (VLMs) such as CLIP, OpenCLIP, EVA02-CLIP and SigLIP achieve strong zero-shot performance, but it is unclear how reliably they respond to controlled l…
cs.MM2024
Pegasus-v1 Technical Report
Raehyuk Jung, Hyojun Go, Jaehyuk Yi +41
This technical report introduces Pegasus-1, a multimodal language model specialized in video content understanding and interaction through natural language. Pegasus-1 is designed t…