collaborators

5 papers

cs.AI2026

From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders

Yifan Luo, Yang Zhan, Jiedong Jiang +4

Sparse autoencoders (SAEs) have proven effective for extracting monosemantic features from large language models (LLMs), yet these features are typically identified in isolation. H…

cs.SE2025

BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution

Terry Yue Zhuo, Xiaolong Jin, Hange Liu +37

Crowdsourced model evaluation platforms, such as Chatbot Arena, enable real-time evaluation from human perspectives to assess the quality of model responses. In the coding domain,…

cs.SE2025

You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation

Yutong Bian, Xianhao Lin, Yupeng Xie +11

Large Language Models (LLMs) and code agents in software development are rapidly evolving from generating isolated code snippets to producing full-fledged software applications wit…

cs.LG2025

Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Zhoujun Cheng, Shibo Hao, Tianyang Liu +21

Reinforcement learning (RL) has emerged as a promising approach to improve large language model (LLM) reasoning, yet most open efforts focus narrowly on math and code, limiting our…

cs.CL2025

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

Yanbin Yin, Kun Zhou, Zhen Wang +11

The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard prac…