collaborators

8 papers

cs.SE2026

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

Zetong Xiong, Qiao Zhao, Jun Zhang +20

Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential…

cs.AI2026

SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

Dexu Yu, Youhua Li, Zhaoyang Guan +12

Agent skills have become a practical way to extend large language model agents, but the growing skill ecosystem still lacks a reliable way to judge whether a skill is worth deployi…

cs.SE2026

SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents

Qiao Zhao, JianYing Qu, Jun Zhang +3

Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generat…

cs.CL2026

Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts

Hanwen Du, Yuxin Dong, Xia Ning

Large Language Models (LLMs) excel at problem solving by generating chain of thoughts in natural language, but such verbal thinking is computationally costly and prone to overthink…

cs.CL2025

EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models

Xinyi Ling, Hanwen Du, Zhihui Zhu +1

E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enha…

cs.CL2025

Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data

Xinyi Ling, Hanwen Du, Bo Peng +2

Leveraging multimodal data to drive breakthroughs in e-commerce applications through Multimodal Foundation Models (MFMs) is gaining increasing attention from the research community…