6 papers
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
Zheyuan Liu, Dongwhi Kim, Yixin Wan +4
Multimodal large language models (MLLMs) are increasingly deployed as assistants that interact through text and images, making it crucial to evaluate contextual safety when risk de…
VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
Zhenghan Tai, Hanwei Wu, Qingchen Hu +24
Retrieval-Augmented Generation (RAG) is becoming increasingly essential for Question Answering (QA) in the financial sector, where accurate and contextually grounded insights from…
FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
Chuan Li, Qianyi Zhao, Fengran Mo +1
Efficiently enhancing the reasoning capabilities of large language models (LLMs) in federated learning environments remains challenging, particularly when balancing performance gai…
UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations
Fengran Mo, Yifan Gao, Chuan Meng +9
The rapid advancement of conversational search systems revolutionizes how information is accessed by enabling the multi-turn interaction between the user and the system. Existing c…
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
Yongan Yu, Qingchen Hu, Xianda Du +3
Climate change adaptation requires the understanding of disruptive weather impacts on society, where large language models (LLMs) might be applicable. However, their effectiveness…
Can Large Language Models Understand Preferences in Personalized Recommendation?
Zhaoxuan Tan, Zinan Zeng, Qingkai Zeng +4
Large Language Models (LLMs) excel in various tasks, including personalized recommendations. Existing evaluation methods often focus on rating prediction, relying on regression err…