collaborators

6 papers

cs.IR2026

Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion

Huatuan Sun, Yunshan Ma, Changguang Wu +3

Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation due to their strong multimodal understanding. However, their integration lacks sy…

cs.CL2026

MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits

Yixin Xiang, Yunshan Ma, Xiaoyu Du +3

Document Question Answering (DQA) involves generating answers from a document based on a user's query, representing a key task in document understanding. This task requires interpr…

cs.CV2026

Stable Signer: Hierarchical Sign Language Generative Model

Sen Fang, Yalin Feng, Hongbin Zhong +2

Sign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages, an…

cs.IR2026

DMAP: Human-Aligned Structural Document Map for Multimodal Document Understanding

ShunLiang Fu, Yanxin Zhang, Yixin Xiang +2

Existing multimodal document question-answering (QA) systems predominantly rely on flat semantic retrieval, representing documents as a set of disconnected text chunks and largely…

cs.GR2025

SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation

Yongzhi Li, Saining Zhang, Yibing Chen +3

Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fid…

cs.AR2025

NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction

Shayne Wadle, Yanxin Zhang, Vikas Singh +1

The evaluation of new microprocessor designs is constrained by slow, cycle-accurate simulators that rely on unrepresentative benchmark traces. This paper introduces a novel deep le…