2 citations · 2 across the 4 of their papers we have counts for
4 papers
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
Kaichen Zhang, Shenghao Gao, Yuzhong Hong +6
Current large language model post-training optimizes a risk-neutral objective that maximizes expected reward, yet evaluation relies heavily on risk-seeking metrics like Pass@k (at…
Streaming Multi-agent Pathfinding
Mingkai Tang, Lu Gan, Kaichen Zhang
The task of the multi-agent pathfinding (MAPF) problem is to navigate a team of agents from their start point to the goal points. However, this setup is unsuitable in the assembly…
Long Context Transfer from Language to Vision
Peiyuan Zhang, Kaichen Zhang, Bo Li +7
Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reduc…
Robust Reward Placement under Uncertainty
Petros Petsinis, Kaichen Zhang, Andreas Pavlogiannis +2
We consider a problem of placing generators of rewards to be collected by randomly moving agents in a network. In many settings, the precise mobility pattern may be one of several…