7 papers
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Shengwei Xu, Yuxuan Lu, Yifan Wu +2
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…
Truthful Calibration Errors for Multi-Class Prediction
Yuxuan Lu, Yifan Wu, Jason Hartline +1
Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate, compare, and tune pr…
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
Guolei Huang, Qinzhi Peng, Gan Xu +3
As Vision-Language Models (VLMs) move into interactive, multi-turn use, safety concerns intensify for multimodal multi-turn dialogue, which is characterized by concealment of malic…
ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization
Yang Zhan, Yunhao Li, Zhang Chao +2
Recent advancements in reinforcement fine-tuning have significantly improved the reasoning ability of large language models (LLMs). In particular, methods such as group relative po…
Jailbreaking LLMs via Calibration
Yuxuan Lu, Yongkang Guo, Yuqing Kong
Safety alignment in Large Language Models (LLMs) often creates a systematic discrepancy between a model's aligned output and the underlying pre-aligned data distribution. We propos…
Aligned Textual Scoring Rules
Yuxuan Lu, Yifan Wu, Jason Hartline +1
Scoring rules elicit probabilistic predictions from a strategic agent by scoring the prediction against a ground truth state. A scoring rule is proper if, from the agent's perspect…