11 papers · 1 filter
Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification
Honglei Jin, Wenshuo Chen, Shaofeng Liang +6
Encrypted traffic classification has achieved strong performance, but its decision process remains difficult to interpret. Existing methods usually combine flow statistics, packet…
MemoGen: Can Past Experience Improve Future Text-to-Image Generation?
Wenshuo Chen, Kuimou Yu, Bowen Tian +10
Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational reasoning, or external knowled…
Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization
Haosen Li, Wenshuo Chen, Lei Wang +3
Text-to-image diffusion models have achieved remarkable generative capabilities, yet accurately aligning complex textual prompts with synthesized layouts remains an ongoing challen…
ECHO: Edge-Cloud Humanoid Orchestration for Language-to-Motion Control
Haozhe Jia, Jianfei Song, Yuan Zhang +5
We present ECHO, an edge--cloud framework for language-driven whole-body control of humanoid robots. A cloud-hosted diffusion-based text-to-motion generator synthesizes motion refe…
Guided Path Sampling: Steering Diffusion Models Back on Track with Principled Path Guidance
Haosen Li, Wenshuo Chen, Shaofeng Liang +3
Iterative refinement methods based on a denoising-inversion cycle are powerful tools for enhancing the quality and control of diffusion models. However, their effectiveness is crit…
POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models
Wenshuo Chen, Haosen Li, Shaofeng Liang +6
The Inversion-Denoising Paradigm, which is based on diffusion models, excels in diverse image editing and restoration tasks. We revisit its mechanism and reveal a critical, overloo…