activity
20242026
most citedBeyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2026

ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Models

Jun Zhang, Jiasheng Zheng, Boxi Cao +5

The emergence of Large Reasoning Models has introduced exceptionally long Chain-of-Thought traces, creating a transparency burden where critical logic is often buried under massive…

cs.CL2026

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

Jiasheng Zheng, Boxi Cao, Boxi Yu +6

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However,…

cs.SE2026

ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models

Jiasheng Zheng, Xin Zheng, Boxi Cao +8

Code sandboxes have emerged as a critical infrastructure for advancing the coding capabilities of large language models, providing verifiable feedback for both RL training and eval…

cs.CV2025

Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models

Qiao Liang, Yanjiang Liu, Weixiang Zhou +7

Does the prior knowledge of the vision encoder constrain the capability boundary of Multi-modal Large Language Models (MLLMs)? While most existing research treats MLLMs as unified…

cs.CL2024

Multi-Facet Counterfactual Learning for Content Quality Evaluation

Jiasheng Zheng, Hongyu Lin, Boxi Cao +4

Evaluating the quality of documents is essential for filtering valuable content from the current massive amount of information. Conventional approaches typically rely on a single s…

cs.SE2024★ 1 cited

Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Jiasheng Zheng, Boxi Cao, Zhengzhao Ma +5

In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily…