6 papers
Brain-Inspired Multimodal Spiking Neural Network for Image-Text Retrieval
Xintao Zong, Xian Zhong, Wenxuan Liu +3
Spiking neural networks (SNNs) have recently shown strong potential in unimodal visual and textual tasks, yet building a directly trained, low-energy, and high-performance SNN for…
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents
Tianmi Ma, Jiawei Du, Wenxin Huang +4
Large language models (LLMs) have demonstrated remarkable capabilities in natural language tasks, yet their performance in dynamic, real-world financial environments remains undere…
OccludeNet: A Causal Journey into Mixed-View Actor-Centric Video Action Recognition under Occlusions
Guanyu Zhou, Wenxuan Liu, Wenxin Huang +3
The lack of occlusion data in common action recognition video datasets limits model robustness and hinders consistent performance gains. We build OccludeNet, a large-scale occluded…
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification
Xiyu Han, Xian Zhong, Wenxin Huang +3
Cloth-changing person re-identification (CC-ReID) aims to match individuals across surveillance cameras despite variations in clothing. Existing methods typically mitigate the impa…
SpikeDerain: Unveiling Clear Videos from Rainy Sequences Using Color Spike Streams
Hanwen Liang, Xian Zhong, Wenxuan Liu +4
Restoring clear frames from rainy videos presents a significant challenge due to the rapid motion of rain streaks. Traditional frame-based visual sensors, which capture scene conte…
Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural Networks
Xian Zhong, Shengwang Hu, Wenxuan Liu +4
Spiking neural networks (SNNs) have garnered significant attention for their low power consumption and high biological interpretability. Their rich spatio-temporal information proc…