5 papers
Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
Denghao Ma, Qing Liu, Zulong Chen +5
Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings wit…
PARM: Pipeline-Adapted Reward Model
Xingyu Fan, Wei Shao, Jiacheng Liu +2
Reward models (RMs) are central to aligning large language models (LLMs) with human preferences, powering RLHF and advanced decoding strategies. While most prior work focuses on si…
Towards Privacy-Preserving Machine Translation at the Inference Stage: A New Task and Benchmark
Wei Shao, Lemao Liu, Yinqiao Li +3
Current online translation services require sending user text to cloud servers, posing a risk of privacy leakage when the text contains sensitive information. This risk hinders the…
Integrating Large Language Models into Recommendation via Mutual Augmentation and Adaptive Aggregation
Sichun Luo, Yuxuan Yao, Bowei He +9
Conventional recommendation methods have achieved notable advancements by harnessing collaborative or sequential information from user behavior. Recently, large language models (LL…
DiffETM: Diffusion Process Enhanced Embedded Topic Model
Wei Shao, Mingyang Liu, Linqi Song
The embedded topic model (ETM) is a widely used approach that assumes the sampled document-topic distribution conforms to the logistic normal distribution for easier optimization.…