4 papers
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
Ziyan Wang, Enmao Diao, Qi Le +6
Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local…
MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches
Xin Wang, Chi Ma, Shaobin Chen +14
Generative recommendation (GR) offers superior modeling capabilities but suffers from prohibitive inference costs due to the repeated encoding of long user histories. While cross-r…
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
Ziyan Wang, Enmao Diao, Qi Le +5
Reasoning LLMs (RLMs) such as OpenAI o1, DeepSeek-R1, and Qwen3 deliver strong multi-step reasoning through chain-of-thought generation, but their large model sizes and lengthy dec…
Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques
Raju Challagundla, Mohsen Dorodchi, Pu Wang +1
As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially f…