activity
20242026
collaborators

6 papers

cs.CL2026

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

Huashan Sun, Shengyi Liao, Yansen Han +8

Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-world long-context information, primar…

cs.LG2026

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

Yuxuan Lu, Ziyi Wang, Jing Huang +10

Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on-policy training. Current web environm…

cs.CL2026

Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data

Yuxuan Lu, Jing Huang, Yan Han +9

Recent research shows that LLM Agents can generate ``believable'' human behaviors via prompt-only methods, and such agents have been increasingly adopted in downstream applications…

cs.CL2025

Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models

Yingqian Cui, Pengfei He, Jingying Zeng +11

Chain-of-Thought (CoT) reasoning, which breaks down complex tasks into intermediate reasoning steps, has significantly enhanced the performance of large language models (LLMs) on c…

cs.CL2025

Extracting and Understanding the Superficial Knowledge in Alignment

Runjin Chen, Gabriel Jacob Perin, Xuxi Chen +5

Alignment of large language models (LLMs) with human values and preferences, often achieved through fine-tuning based on human feedback, is essential for ensuring safe and responsi…

cs.AI2024

A Survey of Calibration Process for Black-Box LLMs

Liangru Xie, Hui Liu, Jingying Zeng +7

Large Language Models (LLMs) demonstrate remarkable performance in semantic understanding and generation, yet accurately assessing their output reliability remains a significant ch…