4 papers
CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph
Chengtao Gan, Xiaoke Guo, Yushan Zhu +5
The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasingly tailored data requiremen…
Scaling LLM Knowledge Boundaries via Distribution-Optimized Synthesis
Songze Li, Yarong Lan, Zhongpu Bo +16
Knowledge injection via synthetic data is crucial for enhancing Large Language Models (LLMs). However, current synthesis methods simply stop at preset token counts or fixed data ra…
CRAFTQA: A Code-Driven Adaptive Framework for Complex Structured Data Reasoning
Chengtao Gan, Zhiqiang Liu, Long Jin +3
Real-world scenarios involve massive heterogeneous structured data (e.g., tables, knowledge graphs), making effective reasoning over such diverse data increasingly important. Unifi…
OntoTune: Ontology-Driven Self-training for Aligning Large Language Models
Zhiqiang Liu, Chengtao Gan, Junjie Wang +5
Existing domain-specific Large Language Models (LLMs) are typically developed by fine-tuning general-purposed LLMs with large-scale domain-specific corpora. However, training on la…