3 papers
cs.AI2026
Inference-Time Budget Control for LLM Search Agents
Zhengru Fang, Senkang Forest Hu, Zhonghao Chang +6
LLM search agents increasingly rely on tools at inference time, but their trajectories are often constrained by hard limits on both tool calls and generated tokens. Under such dual…
cs.NI2026
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
Hongyao Liu, Liuqun Zhai, Junyi Wang +1
Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the ful…
cs.NI2026
An Efficient Wireless iBCI Headstage with Adaptive ADC Sample Rate
Hongyao Liu, Junyi Wang, Jinglong Chen +1
Implantable Brain-Computer Interfaces (iBCIs) are increasingly pivotal in clinical and daily applications. However, wireless iBCIs face severe constraints in power consumption and…