6 papers
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
Kai Xiong, Yanwei Huang, Rongjunchen Zhang +3
High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language models (LLMs). While recent data…
KEditVis: A Visual Analytics System for Knowledge Editing of Large Language Models
Zhenning Chen, Hanbei Zhan, Yanwei Huang +4
Large Language Models (LLMs) demonstrate exceptional capabilities in factual question answering, yet they sometimes provide incorrect responses. To address this issue, knowledge ed…
Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring
Yunfan Zhou, Qiming Shi, Zhongsu Luo +5
LLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge,…
Vipera: Towards systematic auditing of generative text-to-image models at scale
Yanwei Huang, Wesley Hanwen Deng, Sijia Xiao +3
Generative text-to-image (T2I) models are known for their risks related such as bias, offense, and misinformation. Current AI auditing methods face challenges in scalability and th…
StructVizor: Interactive Profiling of Semi-Structured Textual Data
Yanwei Huang, Yan Miao, Di Weng +2
Data profiling plays a critical role in understanding the structure of complex datasets and supporting numerous downstream tasks, such as social media analytics and financial fraud…
Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts
Yunfan Zhou, Xiwen Cai, Qiming Shi +5
Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the…