collaborators

6 papers

cs.CY2026

A Scalable Approach to Evaluating Moral Sensitivity in LLMs

Daniel Kilov, Secil Yanik Guyot, Caroline Hendy +2

Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral compe…

cs.CL2026

Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents

Xushuo Tang, Junhe Zhang, Zihan Yang +4

Recent LLM role-playing systems build character agents from novels by extracting characters, scenes, and relations. Yet long-narrative role-playing suffers from two failures: Factu…

cs.CV2026

NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

Sichao Li, Sai Ma, Daniel Kilov +3

LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior. However, existing approaches eith…

cs.AI2026

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents

Sai Ma, Zhuang Li, Sichao Li +4

Earth Observation (EO) analysis is inherently interactive: resolving uncertainty often requires expanding the region of interest, retrieving historical observations, and switching…

cs.AI2026

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs

Junxian Li, Xinyue Xu, Sai Ma +2

Multimodal Large Language Models (MLLMs) frequently suffer from unfaithfulness, generating reasoning chains that drift from visual evidence or contradict final predictions. We prop…

cs.CL2025

Evaluating LLM Understanding via Structured Tabular Decision Simulations

Sichao Li, Xinyue Xu, Xiaomeng Li

Large language models (LLMs) often achieve impressive predictive accuracy, yet correctness alone does not imply genuine understanding. True LLM understanding, analogous to human ex…