12 citations · 24 across the 8 of their papers we have counts for
8 papers
Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng +4
Recent advancements in multimodal pre-training models have significantly advanced computational pathology. However, current approaches predominantly rely on visual-language models,…
Role Play: Learning Adaptive Role-Specific Strategies in Multi-Agent Interactions
Weifan Long, Wen Wen, Peng Zhai +1
Zero-shot coordination problem in multi-agent reinforcement learning (MARL), which requires agents to adapt to unseen agents, has attracted increasing attention. Traditional approa…
MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
Xiao Zhao, Xukun Zhang, Dingkang Yang +4
Accurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms…
HybridOcc: NeRF Enhanced Transformer-based Multi-Camera 3D Occupancy Prediction
Xiao Zhao, Bo Chen, Mingyang Sun +7
Vision-based 3D semantic scene completion (SSC) describes autonomous driving scenes through 3D volume representations. However, the occlusion of invisible voxels by scene surfaces…
Faster Diffusion Action Segmentation
Shuaibing Wang, Shunli Wang, Mingcheng Li +4
Temporal Action Segmentation (TAS) is an essential task in video analysis, aiming to segment and classify continuous frames into distinct action segments. However, the ambiguous bo…
Large Vision-Language Models as Emotion Recognizers in Context Awareness
Yuxuan Lei, Dingkang Yang, Zhaoyu Chen +3
Context-aware emotion recognition (CAER) is a complex and significant task that requires perceiving emotions from various contextual cues. Previous approaches primarily focus on de…