1 citations · 1 across the 9 of their papers we have counts for
6 papers · 1 filter
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Shuangrui Ding, Xuanlang Dai, Long Xing +14
Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However, most agent benchmarks still…
DARE: Diffusion Large Language Models Alignment and Reinforcement Executor
Jingyi Yang, Yuxian Jiang, Xuhao Hu +3
Diffusion large language models (dLLMs) are emerging as a compelling alternative to dominant autoregressive models, replacing strictly sequential token generation with iterative de…
DeepSight: An All-in-One LM Safety Toolkit
Bo Zhang, Jiaxuan Guo, Lijun Li +17
As the development of Large Models (LMs) progresses rapidly, their safety is also a priority. In current Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) s…
-: Training-free Bidirectional Variable-Length Control for Masked Diffusion LLMs
Jingyi Yang, Yuxian Jiang, Jing Shao
Beyond parallel generation and global context modeling, current masked diffusion large language models (masked dLLMs, i.e., LLaDA) suffer from a fundamental limitation: they requir…
Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
Jingyi Yang, Guanxu Chen, Xuhao Hu +1
Masked diffusion language models (MDLMs) have recently emerged as a promising alternative to autoregressive (AR) language models, offering properties such as parallel decoding, fle…
Shall Your Data Strategy Work? Perform a Swift Study
Minlong Peng, Jingyi Yang, Zhongjun He +1
This work presents a swift method to assess the efficacy of particular types of instruction-tuning data, utilizing just a handful of probe examples and eliminating the need for mod…