6 papers
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
Jingyu Guo, Ziye Chen, Ziwen Li +7
Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descriptions with goal-centric evaluat…
Matryoshka Concept Bottleneck Models
Ziye Chen, Hongbin Lin, Jie Li +1
Concept Bottleneck Models (CBMs) have emerged as a prominent paradigm for interpretable deep learning, learning by grounding predictions in human-understandable concepts. However,…
AR1-ZO: Topology-Aware Rank-1 Zeroth-Order Queries for High-Rank LoRA Fine-Tuning
Ziye Chen, Hongbin Lin, Chenyu Zhang +3
Zeroth-order (ZO) optimization enables large-language-model fine-tuning without storing backpropagation activations, while LoRA supplies compact trainable adapters. Combining them…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Extended Self-similarity in Multimode Optical Fiber Speckles
Mengxin Wu, Ziye Chen, Guang Yang +1
Extended Self-Similarity (ESS) is a widely used tool for uncovering universal power-law scaling in systems dominated by nonlinear interactions. This work demonstrates that ESS scal…
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
Ziye Chen, Chengwei Qin, Yao Shu
As large language models (LLMs) reach high scores on established mathematical benchmarks, such as GSM8K and MATH, the research community has turned to International Mathematical Ol…