activity
20242026
collaborators

5 papers

cs.CL2026

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

Chuanhao Yan, Fengdi Che, Xuhan Huang +12

Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant challenge: their verification proce…

cs.LG2026

Intrinsic Entropy of Context Length Scaling in LLMs

Jingzhe Shi, Qinwei Ma, Hongyi Liu +3

Long Context Language Models have drawn great attention in the past few years. There has been work discussing the impact of long context on Language Model performance: some find th…

cs.LG2025

PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning

Wanjia Zhao, Qinwei Ma, Jingzhe Shi +7

Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively explored. Most existing physics benchmarks eva…

cs.LG2025

Gradient Imbalance in Direct Preference Optimization

Qinwei Ma, Jingzhe Shi, Can Jin +3

Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). How…

cs.CV2024

Graph Canvas for Controllable 3D Scene Generation

Libin Liu, Shen Chen, Sen Jia +6

Spatial intelligence is foundational to AI systems that interact with the physical world, particularly in 3D scene generation and spatial comprehension. Current methodologies for 3…