activity
20232026
most citedLarge Language Models are Temporal and Causal Reasoners for Video Question Answering

2 citations · 2 across the 14 of their papers we have counts for

collaborators

16 papers

cs.CL2026

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Ji Soo Lee, Xilun Chen, Pierce Chuang +5

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over…

cs.CV2026

Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG

Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6

Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, whil…

cs.CV2026

Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization

Byungoh Ko, Jinyoung Park, Jongha Kim +3

Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsisten…

cs.CV2026

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

Seohyun Lee, Seoung Choi, Dohwan Ko +2

As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video…

cs.CV2026

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

Jonggwon Park, Seongeun Lee, Junhyun Park +6

Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing rev…

cs.CV2026

Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

Minseok Joo, Dogyun Park, Taehoon Lee +2

Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving hist…