2 papers
cs.AI2026
Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents
Zedong Liu, Jiaan Wu, Xinyang Ma +5
Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading interfaces often expose entire objects even when only sparse evidence is needed.…
cs.AR2026
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Zedong Liu +17
The rapid scaling of Large Language Models presents significant challenges for their deployment and inference, particularly on resource-constrained specialized AI hardware accelera…