3 papers
cs.RO2026
GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning
Kun Wang, Yiming Li, Mingcheng Qu +3
Implicit spatial relations and deep semantic structures encoded in object attributes are crucial for procedural planning in embodied AI systems. However, existing approaches often…
cs.RO2026
HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models
Kun Wang, Xiao Feng, Mingcheng Qu +1
Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-…
cs.CV2025
EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
Kun Wang, Donglin Di, Tonghua Su +1
Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models of…