8 papers
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
Yixuan Shen, Peng He, Honglu Liu +6
K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these i…
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
Xu Shen, Song Wang, Zhen Tan +5
Large language models (LLMs) increasingly rely on Chain-of-Thought (CoT) prompting to improve problem-solving and provide seemingly transparent explanations. However, growing evide…
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
Zhen Xu, Zhen Tan, Song Wang +2
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting large language models (LLMs) by decomposing token activations into combinations of human-understandable…
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
Pingzhi Li, Morris Yu-Chao Huang, Zhen Tan +6
Knowledge Distillation (KD) accelerates training of large language models (LLMs) but poses intellectual property protection and LLM diversity risks. Existing KD detection methods b…
Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs
Ruichen Zhang, Mufan Qiu, Zhen Tan +7
Web browsing agents powered by large language models (LLMs) have shown tremendous potential in automating complex web-based tasks. Existing approaches typically rely on large LLMs…
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
Shrey Pandit, Jiawei Xu, Junyuan Hong +4
Advancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliability. A critical challenge lies…