1 citations · 1 across the 18 of their papers we have counts for
12 papers · 1 filter
MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
Jason Luo, Saibilila Abudukelimu, Judy Song +4
Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters…
Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure
Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw +3
Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it…
Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
Pratham Singla, Shivank Garg, Vihan Singh +1
Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors, the question's phrasing toge…
Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions
Shivank Garg, Sankalp Mittal, Manish Gupta
Communicating complex system designs or scientific processes through text alone is inefficient and prone to ambiguity. A system that automatically generates scientific architecture…
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
Zafir Shamsi, Nikhil Chekuru, Zachary Guzman +1
Large Language Models (LLMs) are increasingly integrated into high-stakes applications, making robust safety guarantees a central practical and commercial concern. Existing safety…
ViT Registers and Fractal ViT
Jason Chuan-Chih Chou, Abhinav Kumar, Shivank Garg
Drawing inspiration from recent findings including surprisingly decent performance of transformers without positional encoding (NoPE) in the domain of language models and how regis…