6 papers
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
Hao Li, Ganlong Zhao, Yufei Liu +8
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show…
Edit-Based Refinement for Parallel Masked Diffusion Language Models
Houxing Ren, Mingjie Zhan, Zimu Lu +5
Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their performance degrades significant…
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
Houxing Ren, Mingjie Zhan, Zimu Lu +4
Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, existing large language model ba…
Alignment with Fill-In-the-Middle for Enhancing Code Generation
Houxing Ren, Zimu Lu, Weikang Shi +7
The code generation capabilities of Large Language Models (LLMs) have advanced applications like tool invocation and problem-solving. However, improving performance in code-related…
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
Zimu Lu, Yunqiao Yang, Houxing Ren +7
LLM-based agents have demonstrated great potential in generating and managing code within complex codebases. In this paper, we introduce WebGen-Bench, a novel benchmark designed to…
All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
Haowei Chen, Zhiwen Yang, Haotian Hou +4
All-in-one medical image restoration (MedIR) aims to address multiple MedIR tasks using a unified model, concurrently recovering various high-quality (HQ) medical images (e.g., MRI…