5 papers
Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality
Vedant Badoni, Danqi Chen, Xinyi Wang
The performance of modern language models depends critically on pretraining data composition. Yet existing data selection methods rely on auxiliary classifiers for document scoring…
IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking
Ruihua Han, Shuai Wang, Chengyang Li +8
Simulation plays a key role in automated robotics research supported by large language models (LLMs). However, existing simulators often require custom code or complex interfaces,…
Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory
Yanming Liu, Xinyue Peng, Zixuan Yan +7
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, particularly when augmented with search mechanisms that enable systematic explora…
NC2C: Automated Convexification of Generic Non-Convex Optimization Problems
Xinyue Peng, Yanming Liu, Yihan Cang +4
Non-convex optimization problems are pervasive across mathematical programming, engineering design, and scientific computing, often posing intractable challenges for traditional so…
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
Yanming Liu, Xinyue Peng, Jiannan Cao +5
Large Language Models (LLMs) augmented with external tools have demonstrated remarkable capabilities in complex reasoning tasks. However, existing frameworks rely heavily on natura…