Publications (21)
Eliminating Reasoning via Inferring with Planning: A New Framework to Guide LLMs' Non-linear Thinking
Yongqi Tong, Yifan Wang, Dawei Li +4
Chain-of-Thought(CoT) prompting and its variants explore equipping large language models (LLMs) with high-level reasoning abilities by emulating human-like linear cognition and log…
Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation
Ian D. Kivlichan, Zi Lin, Jeremiah Liu +1
Content moderation is often performed by a collaboration between humans and machine learning models. However, it is not well understood how to design the collaborative process so a…
Pruning Redundant Mappings in Transformer Models via Spectral-Normalized Identity Prior
Zi Lin, Jeremiah Zhe Liu, Zi Yang +2
Traditional (unstructured) pruning methods for a Transformer model focus on regularizing the individual weights by penalizing them toward zero. In this work, we explore spectral-no…
Semantic Role Labeling for Learner Chinese: the Importance of Syntactic Parsing and L2-L1 Parallel Data
Zi Lin, Yuguang Duan, Yuanyuan Zhao +2
This paper studies semantic parsing for interlanguage (L2), taking semantic role labeling (SRL) as a case task and learner Chinese as a case language. We first manually annotate th…
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Zi Lin, Zihan Wang, Yongqi Tong +4
Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays.…
Is Argument Structure of Learner Chinese Understandable: A Corpus-Based Analysis
Yuguang Duan, Zi Lin, Weiwei Sun
This paper presents a corpus-based analysis of argument structure errors in learner Chinese. The data for analysis includes sentences produced by language learners as well as their…
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness
Jeremiah Zhe Liu, Shreyas Padhy, Jie Ren +7
Accurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distrib…
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10
Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper,…
Fast Structured Decoding for Sequence Models
Zhiqing Sun, Zhuohan Li, Haoqing Wang +3
Autoregressive sequence models achieve state-of-the-art performance in domains like machine translation. However, due to the autoregressive factorization nature, these models suffe…
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10
Evaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences.…
Large-Scale Generative Data-Free Distillation
Liangchen Luo, Mark Sandler, Zi Lin +2
Knowledge distillation is one of the most popular and effective techniques for knowledge transfer, model compression and semi-supervised learning. Most existing distillation approa…
Neural-Symbolic Inference for Robust Autoregressive Graph Parsing via Compositional Uncertainty Quantification
Zi Lin, Jeremiah Liu, Jingbo Shang
Pre-trained seq2seq models excel at graph semantic parsing with rich annotated data, but generalize worse to out-of-distribution (OOD) and long-tail examples. In comparison, symbol…
Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
Zi Lin, Sheng Shen, Ilia Kulikov +3
Recent advances in large language models (LLMs) have improved their performance on coding benchmarks. However, improvement is plateauing due to the exhaustion of readily available…
Critique Ability of Large Language Models
Liangchen Luo, Zi Lin, Yinxiao Liu +4
Critical thinking is essential for rational decision-making and problem-solving. This skill hinges on the ability to provide precise and reasoned critiques and is a hallmark of hum…
Implanting Rational Knowledge into Distributed Representation at Morpheme Level
Zi Lin, Yang Liu
Previously, researchers paid no attention to the creation of unambiguous morpheme embeddings independent from the corpus, while such information plays an important role in expressi…
Hint-Based Training for Non-Autoregressive Machine Translation
Zhuohan Li, Zi Lin, Di He +4
Due to the unparallelizable nature of the autoregressive factorization, AutoRegressive Translation (ART) models have to generate tokens sequentially during decoding and thus suffer…
A Comparative Analysis of Knowledge-Intensive and Data-Intensive Semantic Parsers
Junjie Cao, Zi Lin, Weiwei Sun +1
We present a phenomenon-oriented comparative analysis of the two dominant approaches in task-independent semantic parsing: classic, knowledge-intensive and neural, data-intensive m…
Optimizing Language Model's Reasoning Abilities with Weak Supervision
Yongqi Tong, Sizhe Wang, Dawei Li +6
While Large Language Models (LLMs) have demonstrated proficiency in handling complex queries, much of the past work has depended on extensively annotated datasets by human experts.…
Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
Jeremiah Zhe Liu, Zi Lin, Shreyas Padhy +3
Bayesian neural networks (BNN) and deep ensembles are principled approaches to estimate the predictive uncertainty of a deep learning model. However their practicality in real-time…
Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots
Zongzheng Zhang, Zi Lin, Jiawen Yang +6
The paper introduces an automated pipeline that creates a manufacturable mechanical facial mechanism for animatronic robots from a single 2D portrait, and simultaneously generates…
Light-driven active phase separation and droplet division
Zi Lin, Thomas Beneyton, Suzanne Lafon +5
Phase separation organizes matter across scales, yet how it operates under sustained energy input remains poorly understood. Experimental approaches to driven phase separation have…