4 papers
Causally-Guided Diffusion for Stable Feature Selection
Arun Vignesh Malarkkan, Xinyuan Wang, Kunpeng Liu +2
Feature selection is fundamental to robust data-centric AI, but most existing methods optimize predictive performance under a single data distribution. This often selects spurious…
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
Arun Vignesh Malarkkan, Manan Roy Choudhury, Guangwei Zhang +4
Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains…
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
Guangwei Zhang, Qisheng Su, Jiateng Liu +4
Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but pose risks of inadvertently exposing copyrighted or proprietary data, especially when such da…
A Survey on Data-Centric AI: Tabular Learning from Reinforcement Learning and Generative AI Perspective
Wangyang Ying, Cong Wei, Nanxu Gong +7
Tabular data is one of the most widely used data formats across various domains such as bioinformatics, healthcare, and marketing. As artificial intelligence moves towards a data-c…