5 papers
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
Songtao Jiang, Sibo Song, Chenyi Zhou +12
The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion…
UniVBench: Towards Unified Evaluation for Video Foundation Models
Jianhui Wei, Xiaotian Zhang, Yichen Li +6
Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-gen…
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
Zhaopeng Feng, Jiahan Ren, Jiayuan Su +3
Process reward models (PRMs) have shown success in complex reasoning tasks for large language models (LLMs). However, their application to machine translation (MT) remains underexp…
V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Yuan Wang, Jiaxiang Liu, Shujian Gao +5
Recent advances in multimodal techniques have led to significant progress in Medical Visual Question Answering (Med-VQA). However, most existing models focus on global image featur…
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
Songtao Jiang, Yan Zhang, Yeying Jin +5
Medical Vision-Language Models (Med-VLMs) have achieved success across various tasks, yet most existing methods overlook the modality misalignment issue that can lead to untrustwor…