activity
20242026
most citedEAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration

15 citations · 17 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng +12

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models…

cs.CV2026

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

Ruofan Hu, Menghui Zhu, Jieming Zhu +8

Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typically combine dense visual e…

cs.CV2026

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

Jiayang Xu, Fan Zhuo, Majun Zhang +6

Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoup…

cs.CV2026

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

Sashuai Zhou, Qiang Zhou, Junpeng Ma +9

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most…

cs.CV2025

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision

Weicai Yan, Wang Lin, Zirun Guo +5

Prompt learning has demonstrated promising results in fine-tuning pre-trained multimodal models. However, the performance improvement is limited when applied to more complex and fi…

cs.CV2025

Astrea: A MOE-based Visual Understanding Model with Progressive Alignment

Xiaoda Yang, JunYu Lu, Hongshun Qiu +12

Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offering a powerful framework for inte…