From the 1 of 5 linked papers with an AI index.
5 papers
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Bonan Zhang, Shiyu Dong, Quan Hung Tran +9
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and i…
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
Yuxin Chen, Liang Luo, Buyun Zhang +44
The paper introduces ROCS, a request-oriented compute sharing framework that restructures recommendation inference to evaluate shared request features once per request rather than…
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents
Peng Xu, Sijia Chen, Junzhuo Li +1
Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ti…
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
Sijia Chen, Yibo Wang, Yi-Feng Wu +5
Tool-augmented large language models (LLMs) leverage tools, often in the form of APIs, to improve their reasoning capabilities on complex tasks. This enables them to act as intelli…