2 papers
cs.CL2026
KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
Yudong Li, Jiawei Cai, Linlin Shen
Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that humans naturally rely on to co…
cs.IR2026
FlexStructRAG: Flexible Structure-Aware Multi-Granular Relational Retrieval for RAG
Mengzhu Chen, Haodong Yang, Jia Cai +1
Retrieval-Augmented Generation (RAG) systems critically depend on how external knowledge is segmented, structured, and retrieved. Most existing approaches either retrieve fixed-len…