collaborators

7 papers

cs.CL2026

Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code

Victoria Basmov, Yoav Goldberg, Reut Tsarfaty

The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output and model behavior. In this…

cs.CL2026

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

Minzheng Wang, Run Luo, Yanbo Wang +6

While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a cr…

physics.chem-ph2026

AgentCAT: An LLM Agent for Extracting and Analyzing Catalytic Reaction Data from Chemical Engineering Literature

Wei Yang, Zihao Liu, Tao Tan +6

This paper presents a large language model (LLM) agent named AgentCAT, which extracts and analyzes catalytic reaction data from chemical engineering papers, %and supports natural l…

cs.AI2026

Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance

Wei Yang, Hong Xie, Tao Tan +3

While open sourced Vision-Language Models (VLMs) have proliferated, selecting the optimal pretrained model for a specific downstream task remains challenging. Exhaustive evaluation…

cs.LG2026

Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective

Hong Xie, Xiao Hu, Tao Tan +5

The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…

cs.LG2026

Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective

Xiao Hu, Hong Xie, Tao Tan +2

A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive…