12 papers
Less Is More? When Dataset Context Hurts LLM-Generated Dataset Descriptions
Lisa-Yao Gan, Arunav Das, Johanna Walker +2
Dataset search and reuse are strongly constrained by the quality of metadata such as natural language descriptions, which are often sparse or inconsistent. Although large language…
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black +2
In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Cit…
TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices
Gerrit Quaremba, Elizabeth Black, Denny VrandeÄiÄ +1
Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such as Wikipedia. Existing detect…
Design Guidance Towards Addressing Over-Reliance on AI in Sensemaking
Yihang Zhao, Wenxin Zhang, Amy Rechkemmer +2
Sensemaking in collaborative work and learning is increasingly supported by GenAI systems, however, emerging evidence suggests that poorly designed GenAI systems tend to provide ex…
Exploring the Design of GenAI-Based Systems to Support Socially Shared Metacognition
Yihang Zhao, Wenxin Zhang, Amy Rechkemmer +2
Socially shared metacognition (SSM) refers to the collective monitoring and regulation of joint cognitive processes in collaborative problem-solving, and is essential for effective…
Schema Generation for Large Knowledge Graphs Using Large Language Models
Bohui Zhang, Yuan He, Lydia Pintscher +2
Schemas play a vital role in ensuring data quality and supporting usability in the Semantic Web and natural language processing. Traditionally, their creation demands substantial i…