collaborators

6 papers

cs.CL2025

DIDS: Domain Impact-aware Data Sampling for Large Language Model Training

Weijie Shi, Jipeng Zhang, Yaguang Wu +8

Large language models (LLMs) are commonly trained on multi-domain datasets, where domain sampling strategies significantly impact model performance due to varying domain importance…

cs.DB2025

E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency

Dongjie Xu, Yue Cui, Weijie Shi +9

SQL query rewriting aims to reformulate a query into a more efficient form while preserving equivalence. Most existing methods rely on predefined rewrite rules. However, such rule-…

cs.CL2025

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong

Sirui Han, Junqi Zhu, Ruiyuan Zhang +1

This paper presents the development of HKGAI-V1, a foundational sovereign large language model (LLM), developed as part of an initiative to establish value-aligned AI infrastructur…

cs.AI2025

LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning

Weijie Shi, Han Zhu, Jiaming Ji +7

Legal judgment prediction (LJP) aims to function as a judge by making final rulings based on case claims and facts, which plays a vital role in the judicial domain for supporting c…

cs.CL2025

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects

Jipeng Zhang, Haolin Yang, Kehao Miao +4

Recent text-to-SQL models have achieved strong performance, but their effectiveness remains largely confined to SQLite due to dataset limitations. However, real-world applications…

cs.CL2025

Benchmarking Multi-National Value Alignment for Large Language Models

Weijie Shi, Chengyi Ju, Chengzhong Liu +8

Do Large Language Models (LLMs) hold positions that conflict with your country's values? Occasionally they do! However, existing works primarily focus on ethical reviews, failing t…