collaborators

8 papers

cs.LG2026

Pretraining EHR Foundation Models with Patient-Aware Sampling

Joshua Placidi, Yuxuan Liu, Jinpei Han +2

Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into…

cs.LG2026

DyGnROLE: Asymmetric Pretraining for Edge Classification on Dynamic Graphs

Tyler Bonnet, Marek Rei

Edge classification on directed dynamic graphs requires modeling interactions between source and destination nodes exhibiting asymmetrical behavioral patterns and temporal dynamics…

cs.SE2026

Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations

Shahin Honarvar, Amber Gorzynski, James Lee-Jones +4

Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing pointwise benchmarks offer limited insi…

cs.CL2026

TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations

Jacob Si, Mike Qu, Michelle Lee +2

Incorporating external knowledge bases in traditional retrieval-augmented generation (RAG) relies on parsing the document, followed by querying a language model with the parsed inf…

cs.AI2025

Fine-tuning with RAG for Improving LLM Learning of New Skills

Humaid Ibrahim, Nikolai Rozanov, Marek Rei

Large language model (LLM) agents deployed for multi-step tasks frequently fail in predictable ways: attempting actions with unmet preconditions, issuing redundant commands, or mis…

cs.CL2025

DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising

Zhenhao Li, Huichi Zhou, Marek Rei +1

Pretrained language models have significantly advanced performance across various natural language processing tasks. However, adversarial attacks continue to pose a critical challe…