activity
20232026
collaborators

8 papers

cs.CV2026

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation

Ruizhe Zeng, Siyu Cao, Lu Zhang +1

Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Language Models (MLLMs) for vision…

cs.CV2026

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

Yiming Ding, Siyu Cao, Luyuan Jiao +4

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each…

cs.CV2026

Towards Temporal Compositional Reasoning in Long-Form Sports Videos

Siyu Cao, Lu Zhang, Ruizhe Zeng +1

Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in Multimodal Large Language M…

cs.CE2026

GA-Field: Geometry-Aware Vehicle Aerodynamic Field Prediction

Zhenhua Zheng, Lu Zhang, Junhong Zou +4

Accurate aerodynamic field prediction is crucial for vehicle drag evaluation, but the computational cost of high-fidelity CFD hinders its use in iterative design workflows. While l…

cs.LG2025

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

Yafei Lyu, Hao Zhou, Lu Zhang +2

Time series forecasting is central to data analysis and web technologies. The recent success of Large Language Models (LLMs) offers significant potential for this field, especially…

cs.CV2024

Boosting Open-Vocabulary Object Detection by Handling Background Samples

Ruizhe Zeng, Lu Zhang, Xu Yang +1

Open-vocabulary object detection is the task of accurately detecting objects from a candidate vocabulary list that includes both base and novel categories. Currently, numerous open…