9 papers
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
Sunqi Fan, Lingshan Chen, Runqi Yin +4
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to th…
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
Xianjing Han, Bin Zhu, Shiqi Hu +4
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on pe…
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
Yuxuan Han, Meng-Hao Guo, Zhengning Liu +2
Optimizing GPU kernels manually is a challenging and time-consuming task. With the rapid development of LLMs, automated GPU kernel optimization is gradually becoming a tangible rea…
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
Zicong Cheng, Ruixuan Jia, Jia Li +3
Diffusion Large Language Models (DLLMs) are inherently ill-suited for variable-length generation, as their inference is defined on a fixed-length canvas and implicitly assumes a kn…
Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging
Jia-peng Zhang, Cheng-Feng Pu, Meng-Hao Guo +2
The rapid proliferation of generative 3D models has created a critical bottleneck in animation pipelines: rigging. Existing automated methods are fundamentally limited by their app…
DEER: Draft with Diffusion, Verify with Autoregressive Models
Zicong Cheng, Guo-Wei Yang, Jia Li +3
Efficiency, as a critical practical challenge for LLM-driven agentic and reasoning systems, is increasingly constrained by the inherent latency of autoregressive (AR) decoding. Spe…