activity
20242026
collaborators

8 papers

cs.CV2026

PROGRESSLM: Towards Progress Reasoning in Vision-Language Models

Jianshu Zhang, Chengxuan Qian, Haosen Sun +4

Estimating task progress requires reasoning over long-horizon dynamics rather than recognizing static visual content. While modern Vision-Language Models (VLMs) excel at describing…

cs.CL2026

What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View

Khouloud Saadi, Di Wang

Feature-based knowledge distillation aims to transfer intermediate representations from a teacher LLM model to a student. Existing approaches typically rely on direct feature match…

cs.LG2026

In-Context Operator Learning on the Space of Probability Measures

Frank Cole, Dixi Wang, Yineng Chen +2

We introduce \emph{in-context operator learning on probability measure spaces} for optimal transport (OT). The goal is to learn a single solution operator that maps a pair of distr…

cs.CV2025

SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

Wenqian Ye, Di Wang, Guangtao Zheng +2

Large vision-language models, such as CLIP, have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models oft…

cs.AI2025

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang +5

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose…

cs.CV2025

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li +15

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…