5 papers
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
Chang Nie, Zhe Liu, Hesheng Wang
End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability is bounded by validated physical coverage.…
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
Chang Nie, Tianchen Deng, Guangming Wang +2
While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. T…
MID: A Self-supervised Multimodal Iterative Denoising Framework
Chang Nie, Tianchen Deng, Zhe Liu +1
Data denoising is a persistent challenge across scientific and engineering domains. Real-world data is frequently corrupted by complex, non-linear noise, rendering traditional rule…
MRASfM: Multi-Camera Reconstruction and Aggregation through Structure-from-Motion in Driving Scenes
Lingfeng Xuan, Chang Nie, Yiqing Xu +3
Structure from Motion (SfM) estimates camera poses and reconstructs point clouds, forming a foundation for various tasks. However, applying SfM to driving scenes captured by multi-…
MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep Thinking
Chang Nie, Yiqing Xu, Guangming Wang +3
Moving object segmentation plays a vital role in understanding dynamic visual environments. While existing methods rely on multi-frame image sequences to identify moving objects, s…