8 papers
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami +1
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional t…
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Mohammad Rostami
Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demon…
Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
Priyashree Roy, Sujitha Martin, Mohammad Rostami +6
Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Ex…
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3
Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full…
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
Zizhao Hu, Mohammad Rostami, Jesse Thomason
Persona prompting can steer LLM generation towards a domain-specific tone and pattern. This behavior enables use cases in multi-agent systems where diverse interactions are crucial…
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
Zizhao Hu, Mohammad Rostami, Jesse Thomason
Recent research has highlighted the risk of generative model collapse, where performance progressively degrades when continually trained on self-generated data. However, existing e…