works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.AI2026

Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution

Zhengbo Jiao, Hongyu Xian, Qinglong Wang +5

The paper introduces Policy of Thoughts (PoT), a test‑time training framework that continuously updates a lightweight LoRA adapter using online policy optimization to improve large…

cs.AI2026

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15

Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…

cs.SE2026

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

Chuan Xiao, Zhengbo Jiao, Shaobo Wang +5

LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-qualit…

cs.CV2026

Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models

Yuhang Han, Yuyang Wu, Zhengbo Jiao +6

Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its appli…

cs.AI2026

Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning

Zhengbo Jiao, Shaobo Wang, Zifan Zhang +4

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet how visual evidence is…

cs.CV2026

Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction

Zhengbo Jiao, Shaobo Wang, Zifan Zhang +4

Multimodal Large Language Models (MLLMs) have significantly advanced vision-language understanding. However, even state-of-the-art models struggle with geometric reasoning, reveali…