5 papers
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
Yi Zhang, Julia Rayz
We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its c…
A Fuzzy Evaluation of Sentence Encoders on Grooming Risk Classification
Geetanjali Bihani, Julia Rayz
With the advent of social media, children are becoming increasingly vulnerable to the risk of grooming in online settings. Detecting grooming instances in an online conversation po…
Evaluating Language Models on Grooming Risk Estimation Using Fuzzy Theory
Geetanjali Bihani, Tatiana Ringenberg, Julia Rayz
Encoding implicit language presents a challenge for language models, especially in high-risk domains where maintaining high precision is important. Automated detection of online ch…
Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes
Damin Zhang, Yi Zhang, Geetanjali Bihani +1
With the impressive performance in various downstream tasks, large language models (LLMs) have been widely integrated into production pipelines, like recruitment and recommendation…
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
Geetanjali Bihani, Julia Rayz
The advent of pre-trained language models (PLMs) has enabled significant performance gains in the field of natural language processing. However, recent studies have found PLMs to s…