activity
20172026
most citedDeep Subdomain Adaptation Network for Image Classification

1.2k citations · 3.4k across the 148 of their papers we have counts for

collaborators
Showing 2024 · cs.CLShow all

24 papers · 2 filters

cs.CL2024★ 1 cited

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation

Zhuohao Yu, Weizheng Gu, Yidong Wang +5

Large Language Models excel at code generation yet struggle with complex programming tasks that demand sophisticated reasoning. To bridge this gap, traditional process supervision…

cs.CL2024

Learning from "Silly" Questions Improves Large Language Models, But Only Slightly

Tingyuan Zhu, Shudong Liu, Yidong Wang +4

Constructing high-quality Supervised Fine-Tuning (SFT) datasets is critical for the training of large language models (LLMs). Recent studies have shown that using data from a speci…

cs.CL2024★ 4 cited

LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output

Elise Karinshak, Amanda Hu, Kewen Kong +4

Immense effort has been dedicated to minimizing the presence of harmful or biased generative content and better aligning AI output to human intention; however, research investigati…

cs.CL2024★ 2 cited

CycleResearcher: Improving Automated Research via Automated Review

Yixuan Weng, Minjun Zhu, Guangsheng Bao +4

The automation of scientific discovery has been a long-standing goal within the research community, driven by the potential to accelerate knowledge creation. While significant prog…

cs.CL2024★ 1 cited

Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?

Yue Huang, Zhengqing Yuan, Yujun Zhou +8

Large Language Models (LLMs) are increasingly employed for simulations, enabling applications in role-playing agents and Computational Social Science (CSS). However, the reliabilit…

cs.CL2024★ 5 cited

On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Hao Chen, Abdul Waheed, Xiang Li +4

The rise of Large Language Models (LLMs) has accentuated the need for diverse, high-quality pre-training data. Synthetic data emerges as a viable solution to the challenges of data…