6 papers
Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
Dota Tianai Dong, Yifan Luo, Po-Ya Angela Wang +2
Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly understood. To address this gap, we…
Probing the Lack of Stable Internal Beliefs in LLMs
Yifan Luo, Kangping Xu, Yanzhen Lu +2
Persona-driven large language models (LLMs) require consistent behavioral tendencies across interactions to simulate human-like personality traits, such as persistence or reliabili…
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
Yifan Zhang, Yifan Luo, Yang Yuan +1
We present Autonomous Data Selection (AutoDS), a method that leverages base language models themselves as zero-shot "generative classifiers" to automatically curate high-quality ma…
Towards Automated Formal Verification of Backend Systems with LLMs
Kangping Xu, Yifan Luo, Yang Yuan +1
Software testing plays a critical role in ensuring that systems behave as intended. However, existing automated testing approaches struggle to match the capabilities of human engin…
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
Kunlun Zhu, Yifan Luo, Dingling Xu +10
Retrieval-Augmented Generation (RAG) is a powerful approach that enables large language models (LLMs) to incorporate external knowledge. However, evaluating the effectiveness of RA…
Augmenting Math Word Problems via Iterative Question Composing
Haoxiong Liu, Yifan Zhang, Yifan Luo +1
Despite the advancements in large language models (LLMs) for mathematical reasoning, solving competition-level math problems remains a significant challenge, especially for open-so…