collaborators

6 papers

q-bio.GN2025

DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences

Xiaoxiao Zhou, Zihan Wang, Jingbo Shang +1

DNA language models have advanced genomics, but their downstream performance varies widely due to differences in tokenization, pretraining data, and architecture. We argue that a m…

cs.CL2025

Order Matters: Rethinking Prompt Construction in In-Context Learning

Warren Li, Yiqian Wang, Zihan Wang +1

In-context learning (ICL) enables large language models to perform new tasks by conditioning on a sequence of examples. Most prior work reasonably and intuitively assumes that whic…

cs.AI2024

Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Zihan Wang, Yunxuan Li, Yuexin Wu +4

Process supervision, using a trained verifier to evaluate the intermediate steps generated by a reasoner, has demonstrated significant improvements in multi-step problem solving. I…

cs.AI2024

Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving

Chenyang An, Zhibo Chen, Qihao Ye +6

Recent advances in Automated Theorem Proving have shown the effectiveness of leveraging a (large) language model that generates tactics (i.e. proof steps) to search through proof s…

cs.SE2024

Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step

Li Zhong, Zilong Wang, Jingbo Shang

Large language models (LLMs) are leading significant progress in code generation. Beyond one-pass code generation, recent works further integrate unit tests and program verifiers i…

cs.CL2024

MEMORYLLM: Towards Self-Updatable Large Language Models

Yu Wang, Yifan Gao, Xiusi Chen +9

Existing Large Language Models (LLMs) usually remain static after deployment, which might make it hard to inject new knowledge into the model. We aim to build models containing a c…