papers

Publications (50)

cs.CV2025

From Physics to Foundation Models: A Review of AI-Driven Quantitative Remote Sensing Inversion

Zhenyu Yu, Mohd Yamani Idna Idris, Hua Wang +3

Quantitative remote sensing inversion aims to estimate continuous surface variables-such as biomass, vegetation indices, and evapotranspiration-from satellite observations, support…

cs.RO2025

VLM as Strategist: Adaptive Generation of Safety-critical Testing Scenarios via Guided Diffusion

Xinzheng Wu, Junyi Chen, Naiting Zhong +1

The safe deployment of autonomous driving systems (ADSs) relies on comprehensive testing and evaluation. However, safety-critical scenarios that can effectively expose system vulne…

cs.CV2026

: Permutation-Equivariant Visual Geometry Learning

Yifan Wang, Jianjun Zhou, Haoyi Zhu +7

We introduce , a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Pre…

cs.CV2024

EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE

Junyi Chen, Longteng Guo, Jia Sun +4

Building scalable vision-language models to learn from diverse, multimodal data remains an open challenge. In this paper, we introduce an Efficient Vision-languagE foundation model…

cs.CL2025

Pre: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation

Junyi Chen, Shihao Bai, Zaijun Wang +7

Extensive LLM applications demand efficient structured generations, particularly for LR(1) grammars, to produce outputs in specified formats (e.g., JSON). Existing methods primaril…

cs.CV2026

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Menglin Han, Yang Ding, Yulei Lu +6

Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained…