Publications (30)
EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue
Jiawen Deng, Wei Li, Wentao Zhang +2
Intelligent dialogue systems are increasingly deployed in emotionally and ethically sensitive settings, where failures in either emotional attunement or ethical judgment can cause…
Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale
Yongsen Pan, Yuxin Chen, Zheng Hu +8
Modern recommendation systems involve massive catalogs of multimodal items, where scalable item identification must balance compactness, semantic fidelity, and downstream effective…
Exploring the Effects of Traditional Chinese Medicine Scents on Mitigating Driving Fatigue
Nengyue Su, Liang Luo, Yu Gu +1
The rise of autonomous driving technology has led to concerns about inactivity-induced fatigue. This paper explores Traditional Chinese Medicine (TCM) scents for mitigating. Two hu…
Graph Learning and Its Advancements on Large Language Models: A Holistic Survey
Shaopeng Wei, Jun Wang, Yu Zhao +6
Graph learning is a prevalent domain that endeavors to learn the intricate relationships among nodes and the topological structure of graphs. Over the years, graph learning has tra…
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support
Junchi Liao, Jiawen Deng, Fuji Ren
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the…
Auditing Evidence Use in Medical LLM Diagnosis
Junchi Liao, Jiawen Deng, Fuji Ren
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. W…
SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation
Peng Hu, Yu Gu, Liang Luo +1
Controllable video generation aims to synthesize video content that aligns precisely with user-provided conditions, such as text descriptions and initial images. However, a signifi…
WiFE: WiFi and Vision based Intelligent Facial-Gesture Emotion Recognition
Yu Gu, Xiang Zhang, Zhi Liu +1
Emotion is an essential part of Artificial Intelligence (AI) and human mental health. Current emotion recognition research mainly focuses on single modality (e.g., facial expressio…
Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts
Jiawen Deng, Wentao Zhang, Ziyun Jiao +1
Conversational AI is increasingly deployed in emotionally charged and ethically sensitive interactions. Previous research has primarily concentrated on emotional benchmarks or stat…
VISTA: Auditing Semantic Divergence in Vision-Language Models
Junchi Liao, Jiawen Deng, Fuji Ren
Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce…
BeSense: Leveraging WiFi Channel Data and Computational Intelligence for Behavior Analysis
Yu Gu, Xiang Zhang, Zhi Liu +1
The ever evolving informatics technology has gradually bounded human and computer in a compact way. Understanding user behavior becomes a key enabler in many fields such as sedenta…
Beyond Empathy: Integrating Diagnostic and Therapeutic Reasoning with Large Language Models for Mental Health Counseling
He Hu, Yucheng Zhou, Juzheng Si +6
Large language models (LLMs) hold significant potential for mental health support, capable of generating empathetic responses and simulating therapeutic conversations. However, exi…
SSP: A Simple and Safe automatic Prompt engineering method towards realistic image synthesis on LVM
Weijin Cheng, Jianzhi Liu, Jiawen Deng +1
Recently, text-to-image (T2I) synthesis has undergone significant advancements, particularly with the emergence of Large Language Models (LLM) and their enhancement in Large Vision…
Bridging the User-side Knowledge Gap in Knowledge-aware Recommendations with Large Language Models
Zheng Hu, Zhe Li, Ziyun Jiao +5
In recent years, knowledge graphs have been integrated into recommender systems as item-side auxiliary information, enhancing recommendation accuracy. However, constructing and int…
MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models
Zhongpu Chen, Yinfeng Liu, Long Shi +3
Large language models (LLMs) are expected to offer structured Markdown responses for the sake of readability in web chatbots (e.g., ChatGPT). Although there are a myriad of metrics…
Stop Treating Collisions Equally: Qualification-Aware Semantic ID Learning for Recommendation at Industrial Scale
Zheng Hu, Yuxin Chen, Yongsen Pan +13
Semantic IDs (SIDs) are compact discrete representations derived from multimodal item features, serving as a unified abstraction for ID-based and generative recommendation. However…
Two-level Attention with Two-stage Multi-task Learning for Facial Emotion Recognition
Xiaohua Wang, Muzi Peng, Lijuan Pan +3
Compared with facial emotion recognition on categorical model, the dimensional emotion recognition can describe numerous emotions of the real world more accurately. Most prior work…
WiFi-based Real-time Breathing and Heart Rate Monitoring during Sleep
Yu Gu, Xiang Zhang, Zhi Liu +1
Good quality sleep is essential for good health and sleep monitoring becomes a vital research topic. This paper provides a low cost, continuous and contactless WiFi-based vital sig…
Code Monitor Red Teaming for Public-Test-Passing Code
Junchi Liao, Jiawen Deng, Fuji Ren
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has p…
Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition
Yan Zhuang, Minhao Liu, Yanru Zhang +2
Multimodal Emotion Recognition (MER) has attracted growing attention with the rapid advancement of human-computer interaction. However, different modalities exhibit substantial dis…
MSSTNet: A Multi-Scale Spatio-Temporal CNN-Transformer Network for Dynamic Facial Expression Recognition
Linhuang Wang, Xin Kang, Fei Ding +2
Unlike typical video action recognition, Dynamic Facial Expression Recognition (DFER) does not involve distinct moving targets but relies on localized changes in facial muscles. Ad…
Disentangling Prosody Representations with Unsupervised Speech Reconstruction
Leyuan Qu, Taihao Li, Cornelius Weber +3
Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentanglin…
Generative Technology for Human Emotion Recognition: A Scope Review
Fei Ma, Yucheng Yuan, Yifan Xie +6
Affective computing stands at the forefront of artificial intelligence (AI), seeking to imbue machines with the ability to comprehend and respond to human emotions. Central to this…
A Review of Human Emotion Synthesis Based on Generative Technology
Fei Ma, Yukan Li, Yifan Xie +8
Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the…
Bert4XMR: Cross-Market Recommendation with Bidirectional Encoder Representations from Transformer
Zheng Hu, Satoshi Nakagawa, Shi-Min Cai +1
Real-world multinational e-commerce companies, such as Amazon and eBay, serve in multiple countries and regions. Some markets are data-scarce, while others are data-rich. In recent…
TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
Yan Zhuang, Minhao Liu, Yanru Zhang +2
Multimodal Sentiment Analysis (MSA) aims to infer human sentiment by integrating information from multiple modalities such as text, audio, and video. In real-world scenarios, howev…
Wital: A COTS WiFi Devices Based Vital Signs Monitoring System Using NLOS Sensing Model
Xiang Zhang, Yu Gu, Huan Yan +5
Vital sign (breathing and heartbeat) monitoring is essential for patient care and sleep disease prevention. Most current solutions are based on wearable sensors or cameras; however…
Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation
Wentao Zhang, Yan Zhuang, ZhuHang Zheng +3
Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are conspicuous and easy to detect…
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset
Yibing Weng, Yu Gu, Fuji Ren
Road rage, triggered by driving-related stimuli such as traffic congestion and aggressive driving, poses a significant threat to road safety. Previous research on road rage regulat…
EmoSense: Computational Intelligence Driven Emotion Sensing via Wireless Channel Data
Yu Gu, Yantong Wang, Tao Liu +6
Emotion is well-recognized as a distinguished symbol of human beings, and it plays a crucial role in our daily lives. Existing vision-based or sensor-based solutions are either obs…