19 citations · 45 across the 20 of their papers we have counts for
22 papers
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian, Yuting Wang, Hanshu Yao +5
While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and sta…
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
Zihao Zhu, Ruotong Wang, Siwei Lyu +2
The rapid advancement of text-to-video (T2V) models has revolutionized content creation, yet their commercial potential remains largely untapped. We introduce, for the first time,…
APAO: Bridging the Training-Inference Gap in Generative Recommendation via Adaptive Prefix-Aware Optimization
Yuanqing Yu, Yifan Wang, Weizhi Ma +2
Generative recommendation has recently emerged as a promising paradigm for sequential recommendation. It formulates the task as an autoregressive generation process, predicting tok…
GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder
Heng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo +2
Stickers have become a popular form of visual communication, yet understanding their semantic relationships remains challenging due to their highly diverse and symbolic content. In…
Integrating LLM and Diffusion-Based Agents for Social Simulation
Xinyi Li, Zhiqiang Guo, Qinglang Guo +3
Large language models (LLMs) offer strong semantic reasoning capabilities for user modeling, but applying LLM-based simulation to an entire social network is computationally expens…
Human vs. Agent in Task-Oriented Conversations
Zhefan Wang, Ning Geng, Zhiqiang Guo +2
Task-oriented conversational systems are essential for efficiently addressing diverse user needs, yet their development requires substantial amounts of high-quality conversational…