From the 1 of 23 linked papers with an AI index.
15 papers · 1 filter
Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching
Runwei Guan, Di Tian, Ningwei Ouyang +9
As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor e…
Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification
Honglei Jin, Wenshuo Chen, Shaofeng Liang +6
The paper presents Traffic-CBM, a multimodal framework that classifies encrypted network traffic by converting flow statistics, temporal features, and byte-level data into hierarch…
Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation
Runwei Guan, Shaofeng Liang, Jiacheng Weng +10
Male infertility contributes substantially to the global infertility burden, and sperm analysis remains central to diagnosis, treatment planning, and assisted reproductive technolo…
MemoGen: Can Past Experience Improve Future Text-to-Image Generation?
Wenshuo Chen, Kuimou Yu, Bowen Tian +10
Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational reasoning, or external knowled…
Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models
Haosen Li, Wenshuo Chen, Lei Wang +4
Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Free Guidance (CFG). However, des…
Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization
Haosen Li, Wenshuo Chen, Lei Wang +3
Text-to-image diffusion models have achieved remarkable generative capabilities, yet accurately aligning complex textual prompts with synthesized layouts remains an ongoing challen…