10 papers
Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning
Shanyong Wang, Shuhang Lin, Yining Zhao +2
Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major challenge remains in adapting L…
RAGRouter-Bench: A Dataset and Benchmark for Adaptive RAG Routing
Ziqi Wang, Xi Zhu, Shuhang Lin +3
Retrieval-augmented generation (RAG) has evolved into a family of paradigms with distinct performance profiles and resource demands, turning paradigm selection into a multi-criteri…
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
Kuan Lu, Shuhang Lin, Sai Wu +7
Large language models (LLMs) are increasingly applied in long-context scenarios such as multi-turn conversations. However, long contexts pose significant challenges for inference e…
OmniRouter: Budget and Performance Controllable Multi-LLM Routing
Kai Mei, Wujiang Xu, Minghao Guo +2
Large language models (LLMs) deliver superior performance but require substantial computational resources and operate with relatively low efficiency, while smaller models can effic…
Cache Mechanism for Agent RAG Systems
Shuhang Lin, Zhencan Peng, Lingyao Li +3
Recent advances in Large Language Model (LLM)-based agents have been propelled by Retrieval-Augmented Generation (RAG), which grants the models access to vast external knowledge ba…
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
Kai Mei, Xi Zhu, Hang Gao +2
We present AIOS 1.0, a novel platform designed to advance computer-use agent (CUA) capabilities through environmental contextualization. While existing approaches primarily focus o…