5 papers
SWE-Mutation: Can LLMs Generate Reliable Test Suites in Software Engineering?
Yuxuan Sun, Yuze Zhao, Yufeng Wang +6
Evaluating software engineering capabilities has become a core component of modern large language models (LLMs); however, the key bottleneck hindering further scaling lies not in t…
MED-COPILOT: A Medical Assistant Powered by GraphRAG and Similar Patient Case Retrieval
Shuheng Chen, Namratha Patil, Haonan Pan +4
Clinical decision-making requires synthesizing heterogeneous evidence, including patient histories, clinical guidelines, and trajectories of comparable cases. While large language…
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
Zhihao Li, Yao Du, Yang Liu +6
Multi-modal Large Language Models (MLLMs) have advanced greatly in general tasks. However, they still face challenges in geometric reasoning, a task that requires synergistic integ…
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
Ziyi Liu, Priyanka Dey, Jen-tse Huang +6
Cultural Intelligence (CQ) refers to the ability to understand unfamiliar cultural contexts, a crucial skill for large language models (LLMs) to effectively engage with globally di…
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
Yufang Liu, Yao Du, Tao Ji +6
Recent research has increasingly focused on multimodal mathematical reasoning, particularly emphasizing the creation of relevant datasets and benchmarks. Despite this, the role of…