activity
20242026
collaborators

12 papers

cs.DB2026

Less Is More? When Dataset Context Hurts LLM-Generated Dataset Descriptions

Lisa-Yao Gan, Arunav Das, Johanna Walker +2

Dataset search and reuse are strongly constrained by the quality of metadata such as natural language descriptions, which are often sparse or inconsistent. Although large language…

cs.CL2026

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black +2

In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Cit…

cs.CL2026

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

Gerrit Quaremba, Elizabeth Black, Denny Vrandečić +1

Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such as Wikipedia. Existing detect…

cs.HC2026

Design Guidance Towards Addressing Over-Reliance on AI in Sensemaking

Yihang Zhao, Wenxin Zhang, Amy Rechkemmer +2

Sensemaking in collaborative work and learning is increasingly supported by GenAI systems, however, emerging evidence suggests that poorly designed GenAI systems tend to provide ex…

cs.HC2026

Exploring the Design of GenAI-Based Systems to Support Socially Shared Metacognition

Yihang Zhao, Wenxin Zhang, Amy Rechkemmer +2

Socially shared metacognition (SSM) refers to the collective monitoring and regulation of joint cognitive processes in collaborative problem-solving, and is essential for effective…

cs.AI2025

Schema Generation for Large Knowledge Graphs Using Large Language Models

Bohui Zhang, Yuan He, Lydia Pintscher +2

Schemas play a vital role in ensuring data quality and supporting usability in the Semantic Web and natural language processing. Traditionally, their creation demands substantial i…