9 papers
Practical Anonymous Two-Party Gradient Boosting Decision Tree
Chenyu Huang, Fan Zhang, Minxin Du +6
Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High sp…
ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL
Zhaorui Yang, Huawei Zheng, Sen Yang +14
Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world databases often contain large…
EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL
Huawei Zheng, Sen Yang, Zhaorui Yang +12
Schema linking is a difficult and important step in large-scale Text-to-SQL, where systems must identify a compact yet sufficient schema context from large and ambiguous databases.…
Privacy-Preserving Screening for Record Linkage
Chenyu Huang, Fan Zhang, Huangxun Chen +4
In an era dominated by big data and machine learning, establishing valuable data collaboration has never been more critical. However, such collaborations must operate under regulat…
SiriusHelper: An LLM Agent-Based Operations Assistant for Big Data Platforms
Yu Shen, Shiyang Liu, Qihang He +14
Big data platforms are widely used in modern enterprises, and an in-production intelligent assistant is increasingly important to help users quickly find actionable guidance and re…
SQLGovernor: An LLM-powered SQL Toolkit for Real World Application
Jie Jiang, Siqi Shen, Haining Xie +8
SQL queries in real world analytical environments, whether written by humans or generated automatically often suffer from syntax errors, inefficiency, or semantic misalignment, esp…