activity
20242026
most citedARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.LG2026

In-Context Operator Learning on the Space of Probability Measures

Frank Cole, Dixi Wang, Yineng Chen +2

We introduce \emph{in-context operator learning on probability measure spaces} for optimal transport (OT). The goal is to learn a single solution operator that maps a pair of distr…

cs.CV2025

SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

Wenqian Ye, Di Wang, Guangtao Zheng +2

Large vision-language models, such as CLIP, have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models oft…

cs.AI2025

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang +5

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose…

cs.CV20251 cited

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li +15

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…

cs.CL2025

Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions

Dingzriui Wang, Xuanliang Zhang, Keyan Xu +3

In-context learning (ICL) has emerged as an effective approach to enhance the performance of large language models (LLMs). However, its effectiveness varies significantly across mo…

cs.CV2024

CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information

Kaifan Zhang, Lihuo He, Xin Jiang +3

Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli.…