7 papers
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Qi Li, Yanzhe Zhao, Yongxin Zhou +4
Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various modalities for a given query. Ho…
Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models
Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou +1
Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decod…
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Yongxi Zhou, Lai Yun Choi, Jiaxi Wen +1
Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We investigate this accuracy--sta…
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
Yongxin Zhou, Philippe Mulhem, Didier Schwab
The evaluation of Retrieval-Augmented Generation (RAG) systems typically examines retrieval quality and generation parameters like temperature in isolation, overlooking their inter…
What Matters to an LLM? Behavioral and Computational Evidences from Summarization
Yongxin Zhou, Changshun Wu, Philippe Mulhem +2
Large Language Models (LLMs) are now state-of-the-art at summarization, yet the internal notion of importance that drives their information selections remains hidden. We propose to…