activity
20242026
collaborators

8 papers

cs.CV2026

ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering

Wenhao Lu, Zhengqiu Zhu, Xiaofeng Wang +7

Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions. Existing outdoor EQA s…

cs.LG2026

One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability

Bhavith Chandra Challagundla, Sanskar Pandey, Param Thakkar +7

World models are now built on substantially different computational substrates. Latent recurrent state-space models such as PlaNet and the Dreamer family compress observations into…

cs.CL2026

AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions

Hengxing Cai, Yijie Rao, Ligang Huang +7

Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making…

cs.MA2026

GeomHerd: A Forward-looking Herding Quantification via Ricci Flow Geometry on Agent Interactive Simulations

Lake Yang, Junwei Su, Jingfeng Zeng +5

Herding -- where agents align their behaviors and act collectively -- is a central driver of market fragility and systemic risk. Existing approaches to quantify herding rely on pri…

cs.AI2026

Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback

Jiaye Lin, Mengdi Li, Xufeng Zhao +4

Reward models trained through Reinforcement Learning from AI Feedback (RLAIF) methods frequently suffer from limited generalizability, which hinders the alignment performance of po…

cs.CL2025

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +89

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…