collaborators

9 papers

cs.CV2026

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs

Jingfeng Chen, Jiawen Qian, Wendi Deng +5

Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token budgets. Existing methods often…

cs.IR2026

Document-as-Image Representations Fall Short for Scientific Retrieval

Ghazal Khalighinejad, Raghuveer Thirukovalluru, Alexander H. Oh +1

Many recent document embedding models are trained on document-as-image representations, embedding rendered pages as images rather than the underlying source. Meanwhile, existing be…

cs.CV2025

Text-Guided Semantic Image Encoder

Raghuveer Thirukovalluru, Xiaochuang Han, Bhuwan Dhingra +2

Image encoders, a fundamental component of vision-language models (VLMs), are typically pretrained independently before being aligned with a language model. This standard paradigm…

cs.CL2025

InData: Towards Secure Multi-Step, Tool-Based Data Analysis

Karthikeyan K, Raghuveer Thirukovalluru, Bhuwan Dhingra +1

Large language model agents for data analysis typically generate and execute code directly on databases. However, when applied to sensitive data, this approach poses significant se…

cs.CL2025

Additive Large Language Models for Semi-Structured Text

Karthikeyan K, Raghuveer Thirukovalluru, David Carlson

Large Language Models have advanced clinical text classification, but their opaque predictions remain a critical barrier to practical adoption in research and clinical settings whe…

cs.CL2025

ClinStructor: AI-Powered Structuring of Unstructured Clinical Texts

Karthikeyan K, Raghuveer Thirukovalluru, David Carlson

Clinical notes contain valuable, context-rich information, but their unstructured format introduces several challenges, including unintended biases (e.g., gender or racial bias), a…