7 papers
SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction
Yuting Ning, Zhehao Zhang, Yash Kumar Lal +8
Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surf…
Explaining GPTs' Schema of Depression: A Machine Behavior Analysis
Adithya V Ganesan, Vasudha Varadarajan, Yash Kumar Lal +9
Use of large language models such as ChatGPT (GPT-4/GPT-5) for mental health support has grown rapidly, emerging as a promising route to assess and help people with mood disorders…
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
Ananya Sadana, Yash Kumar Lal, Jiawei Zhou
Understanding causal relationships across modalities is a core challenge for multimodal models operating in real-world environments. We introduce ISO-Bench, a benchmark for evaluat…
MuSciClaims: Multimodal Scientific Claim Verification
Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan +3
Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large b…
: A Dataset for ynamic nformation nd ental modeling f umeric iscussions
Sayontan Ghosh, Mahnaz Koupaee, Yash Kumar Lal +4
Understanding multiparty conversations demands robust Theory of Mind (ToM) capabilities, including the ability to track dynamic information, manage knowledge asymmetries, and disti…
CaT-BENCH: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
Yash Kumar Lal, Vanya Cohen, Nathanael Chambers +2
Understanding the abilities of LLMs to reason about natural language plans, such as instructional text and recipes, is critical to reliably using them in decision-making systems. A…