5 papers
Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference
Jiayin Hu, Kai Yuan, Vanessa Hu +3
Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhead inherent in static inference, which pro…
Unifying Ranking and Generation in Query Auto-Completion via Retrieval-Augmented Generation and Multi-Objective Alignment
Kai Yuan, Anthony Zheng, Jia Hu +9
Query Auto-Completion (QAC) suggests query completions as users type, helping them articulate intent and reach results more efficiently. Existing approaches face fundamental challe…
Best-of-Q: Improving VLM agents with Q-function Action Ranking at Inference
Emilien Biré, MarÃa Santos, Kai Yuan
Vision-Language Models (VLMs) have become powerful backbones for agents to autonomously operate in digital environments like the web and operating systems. However, these models su…
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
Mathieu Andreux, Märt Bakler, Yanael Barbier +50
Building agents that generalize across web, desktop, and mobile environments remains an open challenge, as prior systems rely on environment-specific interfaces that limit cross-pl…