11 papers
Invariant Reasoning Directions in Latent Trajectories of Language Models
Arun Vignesh Malarkkan, Manan Roy Choudhury, Utkarsh Byahut +3
Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectories remains poorly understood. We show tha…
Causally-Guided Diffusion for Stable Feature Selection
Arun Vignesh Malarkkan, Xinyuan Wang, Kunpeng Liu +2
Feature selection is fundamental to robust data-centric AI, but most existing methods optimize predictive performance under a single data distribution. This often selects spurious…
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
Arun Vignesh Malarkkan, Manan Roy Choudhury, Guangwei Zhang +4
Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains…
Causally-Guided Automated Feature Engineering with Multi-Agent Reinforcement Learning
Arun Vignesh Malarkkan, Wangyang Ying, Yanjie Fu
Automated feature engineering (AFE) enables AI systems to autonomously construct high-utility representations from raw tabular data. However, existing AFE methods rely on statistic…
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
Xinyuan Wang, Kunpeng Liu, Arun Vignesh Malarkkan +1
Feature Transformation (FT) is a core data-centric AI task that improves feature space quality to advance downstream predictive performance. However, discovering effective transfor…
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
Arun Vignesh Malarkkan, Haoyue Bai, Anjali Kaushik +1
In real-world applications, domain data often contains identifiable or sensitive attributes, is subject to strict regulations (e.g., HIPAA, GDPR), and requires explicit data featur…