7 papers
Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object Detection
Linhai Zhuo, Junxi Cai, Tianwen Qian +2
Data augmentation, which simulates diverse visual variations to expand the source distribution and induce synthetic domain shifts, is a simple yet effective strategy for mitigating…
Omni-Supervised Motion Editing: Balancing Change and Invariance through Positive-Negative Learning
Zhenwu Shi, Jingyu Gong, Peiwei Wang +7
Text-based human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the consistency of the original motion. Existi…
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
Kailing Li, Tianwen Qian, Lijin Yang +4
Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions. Despite recent progress with vision-la…
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
Guowei Tang, Tianwen Qian, Huanran Zheng +2
Real-time, continuous understanding of visual signals is essential for real-world interactive AI applications, and poses a fundamental system-level challenge. Existing research on…
CLIP Based Region-Aware Feature Fusion for Automated BBPS Scoring in Colonoscopy Images
Yujia Fu, Zhiyu Dong, Tianwen Qian +3
Accurate assessment of bowel cleanliness is essential for effective colonoscopy procedures. The Boston Bowel Preparation Scale (BBPS) offers a standardized scoring system but suffe…
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
Yifei Wang, Zhenkai Li, Tianwen Qian +4
As embodied intelligence advances toward real-world deployment, the ability to continuously perceive and reason over streaming visual inputs becomes essential. In such settings, an…