6 papers
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
Jingyang Lin, Jialian Wu, Jiang Liu +6
Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, result…
CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios
Jingyang Lin, Yingda Xia, Jianpeng Zhang +5
3D medical vision-language (VL) pretraining has shown potential in radiology by leveraging large-scale multimodal datasets with CT-report pairs. However, existing methods primarily…
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
Jingyang Lin, Jialian Wu, Ximeng Sun +8
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has…
Video Understanding with Large Language Models: A Survey
Yolo Y. Tang, Jing Bi, Siting Xu +17
With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…
Grounding Partially-Defined Events in Multimodal Data
Kate Sanders, Reno Kriz, David Etter +5
How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially…
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
Wenting Chen, Linlin Shen, Jingyang Lin +3
To address these issues, we propose a novel Adaptive patch-word Matching (AdaMatch) model to correlate chest X-ray (CXR) image regions with words in medical reports and apply it to…