7 papers
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
Bingnan Liu, Chenhang Cui, Rui Huang +7
We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by…
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
Jiaping Lin, Fei Shen, Junzhe Li +4
Existing training-free approaches for GUI grounding often rely on multiple inference runs, such as iterative cropping or candidate aggregation, to identify target elements. Despite…
ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage
Wenbo Gao, Songbai Tan, Zhongan Wang +6
Smartphone scams are increasingly prevalent and typically manifest as multi-stage, cross-application processes with gradually emerging intent. Effective intervention thus requires…
RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting
Linrui Xu, Zhongan Wang, Fei Shen +4
Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically ad…
Chain-of-Trajectories: Unlocking the Intrinsic Generative Optimality of Diffusion Models via Graph-Theoretic Planning
Ping Chen, Xiang Liu, Xingpeng Zhang +7
Diffusion models operate in a reflexive System 1 mode, constrained by a fixed, content-agnostic sampling schedule. This rigidity arises from the curse of state dimensionality, wher…
Step-Audio 2 Technical Report
Boyong Wu, Chao Yan, Chen Hu +106
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…