8 papers
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
Jianshu Zhang, Chengxuan Qian, Haosen Sun +4
Estimating task progress requires reasoning over long-horizon dynamics rather than recognizing static visual content. While modern Vision-Language Models (VLMs) excel at describing…
What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View
Khouloud Saadi, Di Wang
Feature-based knowledge distillation aims to transfer intermediate representations from a teacher LLM model to a student. Existing approaches typically rely on direct feature match…
In-Context Operator Learning on the Space of Probability Measures
Frank Cole, Dixi Wang, Yineng Chen +2
We introduce \emph{in-context operator learning on probability measure spaces} for optimal transport (OT). The goal is to learn a single solution operator that maps a pair of distr…
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
Wenqian Ye, Di Wang, Guangtao Zheng +2
Large vision-language models, such as CLIP, have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models oft…
Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion
Jiabao Shi, Minfeng Qi, Lefeng Zhang +5
Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose…
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Yuying Ge, Yixiao Ge, Chen Li +15
Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…