2 papers
cs.NI2025
Toward an AI-Native Internet: Rethinking the Web Architecture for Semantic Retrieval
Muhammad Bilal, Zafar Qazi, Marco Canini
The rise of Generative AI Search is fundamentally transforming how users and intelligent systems interact with the Internet. LLMs increasingly act as intermediaries between humans…
cs.LG2025
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Keyu Wang, Tian Lyu, Guinan Su +4
Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing methods demonstrate strong performance reten…