collaborators

10 papers

cs.CL2026

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

Chengtao Gan, Xiaoke Guo, Yushan Zhu +5

The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasingly tailored data requiremen…

cs.CL2026

Scaling LLM Knowledge Boundaries via Distribution-Optimized Synthesis

Songze Li, Yarong Lan, Zhongpu Bo +16

Knowledge injection via synthetic data is crucial for enhancing Large Language Models (LLMs). However, current synthesis methods simply stop at preset token counts or fixed data ra…

cs.CR2026

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents

Jin Xie, Songze Li

Large language model (LLM) agents increasingly act on a user's behalf -- reading personal files, calling tools, transacting with external services -- possibly leaking personally id…

cs.CV2026

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Ziang Yan, Sheng Xia, Jiashuo Yu +10

Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant se…

cs.AI2026

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

Songze Li, Xiaoke Guo, Tianqi Liu +5

Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct screen-based decision-making,…

cs.CL2026

Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

Songze Li, Zhiqiang Liu, Zhaoyan Gong +6

Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabling them to understand the logic…