14 papers
Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning
Junyang Shu, Zhiwei Lin, Bingqing Wei +1
Reinforcement learning is a promising approach for improving the capabilities of vision-language-action (VLA) models while avoiding the heavy data requirements of imitation learnin…
CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search
Zhe Li, Zhiwei Lin, Yongtao Wang
The integration of Large Language Models (LLMs) with Neural Architecture Search (NAS) has introduced new possibilities for automating the design of neural architectures. However, m…
VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection
Chih-Chung Liu, Zhiwei Lin, Yongtao Wang
Open-world object detection aims to localize and recognize objects beyond a fixed closed-set label space. It is commonly divided into two categories, i.e., open-vocabulary detectio…
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
Xinhao Wang, Zhonyu Xia, Zhiwei Lin +2
Multimodal Large Language Models (MLLMs) have shown strong reasoning ability, but their high computational and memory costs hinder deployment in resource-constrained settings. Whil…
ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents
Bingqing Wei, Zhongyu Xia, Dingai Liu +3
Vision-language models (VLMs) have shown remarkable general capabilities, yet embodied agents built on them fail at complex tasks, often skipping critical steps, proposing invalid…
Multi-Representation Adapter with Neural Architecture Search for Efficient Range-Doppler Radar Object Detection
Zhiwei Lin, Weicheng Zheng, Yongtao Wang
Detecting objects efficiently from radar sensors has recently become a popular trend due to their robustness against adverse lighting and weather conditions compared with cameras.…