Publications (12)
UniDiffGrasp: A Unified Framework Integrating VLM Reasoning and VLM-Guided Part Diffusion for Open-Vocabulary Constrained Grasping with Dual Arms
Xueyang Guo, Hongwei Hu, Chengye Song +5
Open-vocabulary, task-oriented grasping of specific functional parts, particularly with dual arms, remains a key challenge, as current Vision-Language Models (VLMs), while enhancin…
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Donghui Feng, Fengxi Zhang, Changsheng Gao +6
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…
Adaptive Learned Image Compression with Graph Neural Networks
Yunuo Chen, Bing He, Zezheng Lyu +4
Efficient image compression relies on modeling both local and global redundancy. Most state-of-the-art (SOTA) learned image compression (LIC) methods are based on CNNs or Transform…
Linear Attention Modeling for Learned Image Compression
Donghui Feng, Zhengxue Cheng, Shen Wang +4
Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based tran…
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
Yuelin Hu, Zhengxue Cheng, Ronghua Wu +5
Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We p…
Restoring Network Evolution from Static Structure
Jiu Zhang, Zhanwei Du, Hongwei Hu +6
The dynamical evolution of complex networks underpins the structure-function relationships in natural and artificial systems. Yet, restoring a network's formation from a single sta…