Publications (26)
Amortized Inference for Model Rocket Aerodynamics: Learning to Estimate Physical Parameters from Simulation
Rohit Pandey, Rohan Pandey
Accurate prediction of model rocket flight performance requires estimating aerodynamic parameters that are difficult to measure directly. Traditional approaches rely on computation…
Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
Rohan Pandey, Rulin Shao, Paul Pu Liang +2
Despite recent progress towards scaling up multimodal vision-language models, these models are still known to struggle on compositional generalization benchmarks such as Winoground…
Logical Undefinability of the Generalized Collatz Transition Relation in Büchi Arithmetic
Madhav Dhiman, Rohan Pandey
Let be an odd prime and let be an odd integer. We show that the arbitrary-step transition relation of the generalized Collatz map is not first-order definable in…
Non-Definability of Reachability in Büchi Arithmetic for a Family of Generalized Collatz Maps
Madhav Dhiman, Rohan Pandey
Let and be odd integers with a power of . We study the generalized Collatz map , a one-dimensional piecewise-affine map on the positive intege…
Athena 2.0: Discourse and User Modeling in Open Domain Dialogue
Omkar Patil, Lena Reed, Kevin K. Bowden +12
Conversational agents are consistently growing in popularity and many people interact with them every day. While many conversational agents act as personal assistants, they can hav…
Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners
Rohan Pandey, Eric Ye, Michael Li
As Large Language Models (LLMs) achieve increasingly sophisticated performance on complex reasoning tasks, current architectures serve as critical proxies for the internal heuristi…
CircuitBuilder: From Polynomials to Circuits via Reinforcement Learning
Weikun K. Zhang, Rohan Pandey, Bhaumik Mehta +5
Motivated by auto-proof generation and Valiant's VP vs. VNP conjecture, we study the problem of discovering efficient arithmetic circuits to compute polynomials, using addition and…
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
Paul Pu Liang, Chun Kai Ling, Yun Cheng +6
In many machine learning systems that jointly learn from multiple modalities, a core research question is to understand the nature of multimodal interactions: how modalities combin…
Predicting first-episode homelessness among US Veterans using longitudinal EHR data: time-varying models and social risk factors
Rohan Pandey, Haijuan Yan, Hong Yu +1
Homelessness among US veterans remains a critical public health challenge, yet risk prediction offers a pathway for proactive intervention. In this retrospective prognostic study,…
Parity-Dependent Real-Rootedness in Independence Polynomials of Generalized Petersen Graphs
Rohan Pandey
We investigate the distribution of zeros of the independence polynomial for the family of Generalized Petersen graphs in the complex plane. While t…
FactorLibrary: From Polynomials to Circuits via Recursive Subgoals
Rohan Pandey, Michael Ruofan Zeng, Weikun K. Zhang +5
Finding minimal arithmetic circuits for polynomials over finite fields is a combinatorially hard problem central to algebraic complexity theory. We formulate it as a reinforcement…
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
Rohan Pandey, Archit Bhujang
Large language models (LLMs) are increasingly used as analyst assistants in security operations centers (SOCs), where they ingest log and alert data to produce triage labels, incid…
Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix
Shree Murthy, Rohan Pandey
We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and…
A Cross-lingual Natural Language Processing Framework for Infodemic Management
Ridam Pal, Rohan Pandey, Vaibhav Gautam +2
The COVID-19 pandemic has put immense pressure on health systems which are further strained due to the misinformation surrounding it. Under such a situation, providing the right in…
gzip Predicts Data-dependent Scaling Laws
Rohan Pandey
Past work has established scaling laws that predict the performance of a neural language model (LM) as a function of its parameter count and the number of tokens it's trained on, e…
Syntax-guided Neural Module Distillation to Probe Compositionality in Sentence Embeddings
Rohan Pandey
Past work probing compositionality in sentence embedding models faces issues determining the causal impact of implicit syntax representations. Given a sentence, we construct a neur…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Semantic Composition in Visually Grounded Language Models
Rohan Pandey
What is sentence meaning and its ideal representation? Much of the expressive power of human language derives from semantic composition, the mind's ability to represent meaning hie…
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
Dhruv Saini, Rohan Pandey
Large language models often improve on difficult tasks by spending inference-time compute on a reasoning trace before producing the final answer. That extra computation can be usef…
Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP
Vedant Palit, Rohan Pandey, Aryaman Arora +1
Mechanistic interpretability seeks to understand the neural mechanisms that enable specific behaviors in Large Language Models (LLMs) by leveraging causality-based methods. While t…
Quantization Blindspots: How Model Compression Breaks Backdoor Defenses
Rohan Pandey, Eric Ye
Backdoor attacks embed input-dependent malicious behavior into neural networks while preserving high clean accuracy, making them a persistent threat for deployed ML systems. At the…
Analyzing the birth-death model of Oncostreams in Glioma, and the effects of Cytochalasin D treatment
Kai Poffenbarger, Rohan Pandey
This research project investigates the critical role of oncostreams in glioma aggressiveness, leveraging advanced ex-vivo 3D explants and in-vivo intravital imaging techniques to e…
Athena 2.0: Contextualized Dialogue Management for an Alexa Prize SocialBot
Juraj Juraska, Kevin K. Bowden, Lena Reed +12
Athena 2.0 is an Alexa Prize SocialBot that has been a finalist in the last two Alexa Prize Grand Challenges. One reason for Athena's success is its novel dialogue management strat…
A Machine Learning Application for Raising WASH Awareness in the Times of COVID-19 Pandemic
Rohan Pandey, Vaibhav Gautam, Ridam Pal +17
Background: The COVID-19 pandemic has uncovered the potential of digital misinformation in shaping the health of nations. The deluge of unverified information that spreads faster t…
The Möbius function of the poset of triangular numbers under divisibility
Rohan Pandey, Harry Richman
This paper analyzes the Möbius () function defined on the partially ordered set of triangular numbers () under the divisibility relation. We make conjectures…
(Un)Masked COVID-19 Trends from Social Media
Asmit Kumar Singh, Paras Mehan, Divyanshu Sharma +3
Wearing masks is a useful protection method against COVID-19, which has caused widespread economic and social impact worldwide. Across the globe, governments have put mandates for…