collaborators

12 papers

cs.CL2026

TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation

Xinkai Ma, Zhiqi Bai, Dingling Zhang +21

Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmarks and systems remain predom…

cs.AI2026

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Wei Pang, Xiangru Jian, Hehan Li +10

Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they opera…

cs.CL2026

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

Haoran Xu, Hongyu Wang, Yifei Gao +3

On-policy distillation (OPD) trains a student on its own trajectories with dense per-token supervision from a stronger teacher, and often outperforms off-policy distillation and st…

cs.CV2026

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

Haoran Xu, Hongyu Wang, Yifei Gao +4

Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perceptual commitment and halluci…

cs.AI2026

Trajectory-Refined Distillation

Li Jiang, Haoran Xu, Yichuan Ding +1

On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollout…

cs.LG2026

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

Ningyuan Xi, Hao Xu, Hongsheng Xin +1

Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement learning with verifiable rewards…