6 papers
Orchestrating Specialized Agents for Trustworthy Enterprise RAG
Xincheng You, Qi Sun, Neha Bora +4
Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict tr…
Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation
Xiaoyan Liu, Kangrui Li, Yuehao Song +1
The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations a…
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
Zheng Xiong, Kang Li, Zilin Wang +3
Built upon language and vision foundation models with strong generalization ability and trained on large-scale robotic data, Vision-Language-Action (VLA) models have recently emerg…
Deep Research Agents: A Systematic Examination And Roadmap
Yuxuan Huang, Yihang Chen, Haozheng Zhang +10
The rapid progress of Large Language Models (LLMs) has given rise to a new category of autonomous AI systems, referred to as Deep Research (DR) agents. These agents are designed to…
Enterprise Large Language Model Evaluation Benchmark
Liya Wang, David Yi, Damien Jose +4
Large Language Models (LLMs) ) have demonstrated promise in boosting productivity across AI-powered tools, yet existing benchmarks like Massive Multitask Language Understanding (MM…
ViMo: A Generative Visual GUI World Model for App Agents
Dezhao Luo, Bohan Tang, Kang Li +6
App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle w…