6 papers
BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition
Bojing Hou, Ruohao Li, Yitong Zhu +3
To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this p…
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4
Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematica…
Harnessing LLM Agents with Skill Programs
Hongjun Liu, Yifei Ming, Shafiq Joty +1
Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizon tasks. However, such lesso…
Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
Chenruo Liu, Hongjun Liu, Zeyu Lai +3
To enhance group robustness to spurious correlations, prior work often relies on auxiliary group annotations and assumes identical sets of groups across training and test domains.…
SwarmSys: Decentralized Swarm-Inspired Agents for Scalable and Adaptive Reasoning
Ruohao Li, Hongjun Liu, Leyi Zhao +7
Large language model (LLM) agents have shown remarkable reasoning abilities. However, existing multi-agent frameworks often rely on fixed roles or centralized control, limiting sca…
MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
Hongjun Liu, Yinghao Zhu, Yuhui Wang +4
Recent progress in multimodal large language models (MLLMs) has demonstrated promising performance on medical benchmarks and in preliminary trials as clinical assistants. Yet, our…