4 papers · 1 filter
Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality
Vedant Badoni, Danqi Chen, Xinyi Wang
The performance of modern language models depends critically on pretraining data composition. Yet existing data selection methods rely on auxiliary classifiers for document scoring…
Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory
Yanming Liu, Xinyue Peng, Zixuan Yan +7
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, particularly when augmented with search mechanisms that enable systematic explora…
NC2C: Automated Convexification of Generic Non-Convex Optimization Problems
Xinyue Peng, Yanming Liu, Yihan Cang +4
Non-convex optimization problems are pervasive across mathematical programming, engineering design, and scientific computing, often posing intractable challenges for traditional so…
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
Yanming Liu, Xinyue Peng, Jiannan Cao +5
Large Language Models (LLMs) augmented with external tools have demonstrated remarkable capabilities in complex reasoning tasks. However, existing frameworks rely heavily on natura…