collaborators

6 papers

cs.CV2025

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

Cheng Chen, Yunpeng Zhai, Yifan Zhao +3

In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, im…

cs.LG2025

Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model

Xue Wang, Tian Zhou, Jinyang Gao +2

We present a joint forecasting framework for time series prediction that contrasts with traditional direct or recursive methods. This framework achieves state-of-the-art performanc…

cs.SE2025

ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning

Lingzhe Zhang, Yunpeng Zhai, Tong Jia +6

As modern microservice systems grow increasingly popular and complex-often consisting of hundreds or even thousands of fine-grained, interdependent components-they are becoming mor…

cs.LG2025

RePO: Understanding Preference Learning Through ReLU-Based Optimization

Junkang Wu, Kexin Huang, Xue Wang +5

Aligning large language models (LLMs) with human preferences is critical for real-world deployment, yet existing methods like RLHF face computational and stability challenges. Whil…

cs.LG2025

Larger or Smaller Reward Margins to Select Preferences for Alignment?

Kexin Huang, Junkang Wu, Ziqian Chen +6

Preference learning is critical for aligning large language models (LLMs) with human values, with the quality of preference datasets playing a crucial role in this process. While e…

cs.CL2025

ToolCoder: A Systematic Code-Empowered Tool Learning Framework for Large Language Models

Hanxing Ding, Shuchang Tao, Liang Pang +5

Tool learning has emerged as a crucial capability for large language models (LLMs) to solve complex real-world tasks through interaction with external tools. Existing approaches fa…