most citedVoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion

8 citations · 14 across the 8 of their papers we have counts for

collaborators

5 papers

cs.CL20231 cited

InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition

Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu +5

The local and global features are both essential for automatic speech recognition (ASR). Many recent methods have verified that simply combining local and global features can furth…

cs.CL2023

Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding

Tian-Hao Zhang, Hai-Bo Qin, Zhi-Hao Lai +5

Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and seman…

cs.CV2023

SCTracker: Multi-object tracking with shape and confidence constraints

Huan Mao, Yulin Chen, Zongtan Li +2

Detection-based tracking is one of the main methods of multi-object tracking. It can obtain good tracking results when using excellent detectors but it may associate wrong targets…

cs.CV20238 cited

VoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion

Yiming Li, Zhiding Yu, Christopher Choy +5

Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in A…

cs.LG20231 cited

Efficient Communication via Self-supervised Information Aggregation for Online and Offline Multi-agent Reinforcement Learning

Cong Guan, Feng Chen, Lei Yuan +2

Utilizing messages from teammates can improve coordination in cooperative Multi-agent Reinforcement Learning (MARL). Previous works typically combine raw messages of teammates with…