activity
20232026
most citedLanguage-driven Scene Synthesis using Multi-conditional Diffusion Model

4 citations · 12 across the 26 of their papers we have counts for

collaborators

16 papers

cs.RO2026

Gripper-aware Vision Language Action Models

Hanyi Zhang, Zihong Luo, Tianyu Li +16

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instru…

cs.RO2026

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10

Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning…

cs.CV2026

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…

cs.CV2026

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

Nghia Vu, Tuong Do, Khang Nguyen +8

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of o…

cs.CV2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

Bao Truong, Quang Nguyen, Baoru Huang +6

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface flu…

cs.AI2026

Pessimistic Auxiliary Policy for Offline Reinforcement Learning

Fan Zhang, Baoru Huang, Xin Zhang

Offline reinforcement learning aims to learn an agent from pre-collected datasets, avoiding unsafe and inefficient real-time interaction. However, inevitable access to out-ofdistri…