most citedEnhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AI2025

FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI

Yuhang Peng, Yizhou Pan, Xinning He +6

As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex,…

cs.CV2025

RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

Xiyan Liu, Han Wang, Yuhu Wang +4

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous…

cs.RO2025

Embodied Cognition Augmented End2End Autonomous Driving

Ling Niu, Xiaoji Zheng, Han Wang +4

In recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networ…

cs.CV2025

Dynamic Double Space Tower

Weikai Sun, Shijie Song, Han Wang

The Visual Question Answering (VQA) task requires the simultaneous understanding of image content and question semantics. However, existing methods often have difficulty handling c…

cs.RO2025

Bench2FreeAD: A Benchmark for Vision-based End-to-end Navigation in Unstructured Robotic Environments

Yuhang Peng, Sidong Wang, Jihaoyu Yang +3

Most current end-to-end (E2E) autonomous driving algorithms are built on standard vehicles in structured transportation scenarios, lacking exploration of robot navigation for unstr…

cs.CV20241 cited

Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion

Peiyuan Chen, Zecheng Zhang, Yiping Dong +2

Visual Question Answering (VQA) is a challenging task that requires systems to provide accurate answers to questions based on image content. Current VQA models struggle with comple…