5 papers
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
Songshuo Lu, Hua Wang, Zhi Chen +1
Large-scale alignment pipelines typically pair a policy model with a separately trained reward model whose parameters remain frozen during reinforcement learning (RL). This separat…
DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents
Yibin Xu, Liang Yang, Hao Chen +3
The limitation of graphical user interface (GUI) data has been a significant barrier to the development of GUI agents today, especially for the desktop / computer use scenarios. To…
StableGS: A Floater-Free Framework for 3D Gaussian Splatting
Luchao Wang, Qian Ren, Kaimin Liao +3
3D Gaussian Splatting (3DGS) reconstructions are plagued by stubborn ``floater" artifacts that degrade their geometric and visual fidelity. We are the first to reveal the root caus…
SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models
Zicheng Cai, Yaohua Tang, Yutao Lai +3
We introduce SEKI, a novel large language model (LLM)-based neural architecture search (NAS) method. Inspired by the chain-of-thought (CoT) paradigm in modern LLMs, SEKI operates i…
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
Yaohua Tang, Zhicheng Hu, Kun Cheng +4
The increasing context window size in large language models (LLMs) has improved their ability to handle complex, long-text tasks. However, as the conversation rounds continue, it i…