10 papers
De-attribute to Forget for LLM Unlearning
Xinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim +3
The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing interest in LLM unlearning. Many…
One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents
Yingchaojie Feng, Qiang Huang, Xiaoya Xie +4
Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain la…
MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition
Yifan Bao, Xinyu Xi, Xinyu Liu +8
Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation p…
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
Yiqun Sun, Qiang Huang, Anthony K. H. Tung +1
This position paper argues that text embedding research should move beyond surface meaning and embrace implicit semantics as a central modeling objective. Text embeddings are a fou…
Weight-Informed Self-Explaining Clustering for Mixed-Type Tabular Data
Lehao Li, Qiang Huang, Yihao Ang +3
Clustering mixed-type tabular data is fundamental for exploratory analysis, yet remains challenging due to misaligned numerical-categorical representations, uneven and context-depe…
RFOD: Random Forest-based Outlier Detection for Tabular Data
Yihao Ang, Peicheng Yao, Yifan Bao +4
Outlier detection in tabular data is crucial for safeguarding data integrity in high-stakes domains such as cybersecurity, financial fraud detection, and healthcare, where anomalie…