7 papers
ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
Yuhang Zhan, Lisi Chen, Shuo Shang
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate cont…
From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs
Silin Zhou, Chenhao Wang, Yuntao Wen +3
Urban trajectories play a crucial role in modeling urban dynamics and supporting various smart city applications. However, privacy concerns restrict access to large-scale and high-…
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
Ruixiang Feng, Yuntao Wen, Silin Zhou +14
Language Reasoning Models (LRMs) achieve strong performance by scaling test-time computation but often suffer from ``overthinking'', producing excessively long reasoning traces tha…
V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like Chat
Qi Lin, Weikai Xu, Lisi Chen +1
With the continued proliferation of Large Language Model (LLM) based chatbots, there is a growing demand for generating responses that are not only linguistically fluent but also c…
Efficient Model-Agnostic Continual Learning for Next POI Recommendation
Chenhao Wang, Shanshan Feng, Lisi Chen +2
Next point-of-interest (POI) recommendation improves personalized location-based services by predicting users' next destinations based on their historical check-ins. However, most…
Region-Point Joint Representation for Effective Trajectory Similarity Learning
Hao Long, Silin Zhou, Lisi Chen +1
Recent learning-based methods have reduced the computational complexity of traditional trajectory similarity computation, but state-of-the-art (SOTA) methods still fail to leverage…