5 papers
Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
Shuozhe Cheng, Kunlan Xiang, Mingxuan Li +3
Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead…
When Poison Fails After Retrieval: Revisiting Corpus Poisoning under Chunking and Reranking Pipelines
Xi Nie, Hongwei Li, Shenghao Wu +3
Retrieval-Augmented Generation (RAG) systems are vulnerable to corpus poisoning attacks that manipulate downstream model outputs through malicious knowledge injection. Existing stu…
Solvaformer: an SE(3)-equivariant graph transformer for small molecule solubility prediction
Jonathan Broadbent, Michael Bailey, Mingxuan Li +6
Accurate prediction of small molecule solubility using material-sparing approaches is critical for accelerating synthesis and process optimization, yet experimental measurement is…
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
Mingxuan Li, Hanchen Li, Chenhao Tan
Large language models (LLMs) have demonstrated great potential for automating the evaluation of natural language generation. Previous frameworks of LLM-as-a-judge fall short in two…
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
Haokun Liu, Yangqiaoyu Zhou, Mingxuan Li +2
AI holds promise for transforming scientific processes, including hypothesis generation. Prior work on hypothesis generation can be broadly categorized into theory-driven and data-…