11 papers
Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
Bo Wang, Ruixing Zhang, Yunqi Liu +4
User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profi…
SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay
Guilin Li, Jiaxing Zhang, Matthias Hwai Yong Tan +2
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful acti…
PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
Guilin Li, Yun Zhang, Xiuyuan Chen +6
Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world know…
Scaling Embeddings Outperforms Scaling Experts in Language Models
Hong Liu, Jiaqi Zhang, Chao Wang +13
While Mixture-of-Experts (MoE) architectures have become the standard for sparsity scaling in large language models, they increasingly face diminishing returns and system-level bot…
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
Hong Chen, Xiang Liu, Bo Wang +5
The linear growth of Key-Value (KV) cache remains a bottleneck for multi-turn LLM deployment. Existing KV cache compression methods often fail to account for the structural propert…
Urban In-Context Learning: Bridging Pretraining and Inference through Masked Diffusion for Urban Profiling
Ruixing Zhang, Bo Wang, Tongyu Zhu +2
Urban profiling aims to predict urban profiles in unknown regions and plays a critical role in economic and social censuses. Existing approaches typically follow a two-stage paradi…