4 citations · 12 across the 26 of their papers we have counts for
16 papers
Gripper-aware Vision Language Action Models
Hanyi Zhang, Zihong Luo, Tianyu Li +16
Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instru…
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10
Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning…
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6
Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
Nghia Vu, Tuong Do, Khang Nguyen +8
Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of o…
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
Bao Truong, Quang Nguyen, Baoru Huang +6
Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface flu…
Pessimistic Auxiliary Policy for Offline Reinforcement Learning
Fan Zhang, Baoru Huang, Xin Zhang
Offline reinforcement learning aims to learn an agent from pre-collected datasets, avoiding unsafe and inefficient real-time interaction. However, inevitable access to out-ofdistri…