collaborators

6 papers

cs.LG2026

Weight-Informed Self-Explaining Clustering for Mixed-Type Tabular Data

Lehao Li, Qiang Huang, Yihao Ang +3

Clustering mixed-type tabular data is fundamental for exploratory analysis, yet remains challenging due to misaligned numerical-categorical representations, uneven and context-depe…

cs.LG2025

RFOD: Random Forest-based Outlier Detection for Tabular Data

Yihao Ang, Peicheng Yao, Yifan Bao +4

Outlier detection in tabular data is crucial for safeguarding data integrity in high-stakes domains such as cybersecurity, financial fraud detection, and healthcare, where anomalie…

cs.AI2025

Structured Agentic Workflows for Financial Time-Series Modeling with LLMs and Reflective Feedback

Yihao Ang, Yifan Bao, Lei Jiang +4

Time-series data is central to decision-making in financial markets, yet building high-performing, interpretable, and auditable models remains a major challenge. While Automated Ma…

q-fin.ST2025

CTBench: Cryptocurrency Time Series Generation Benchmark

Yihao Ang, Qiang Wang, Qiang Huang +5

Synthetic time series are essential tools for data augmentation, stress testing, and algorithmic prototyping in quantitative finance. However, in cryptocurrency markets, characteri…

cs.CL2025

Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation

Yingchaojie Feng, Yiqun Sun, Yandong Sun +4

In this work, we investigate an important task named instruction-following text embedding, which generates dynamic text embeddings that adapt to user instructions, highlighting spe…

cs.CL2025

PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder

Yiqun Sun, Qiang Huang, Anthony K. H. Tung +1

Semantic Text Embedding is a fundamental NLP task that encodes textual content into vector representations, where proximity in the embedding space reflects semantic similarity. Whi…