activity
20242026
most citedDetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection

1 citations · 1 across the 9 of their papers we have counts for

collaborators

12 papers

cs.AI2026

Ostrakon-VL: Towards Domain-Expert MLLM for Food-Service and Retail Stores

Zhiyong Shen, Gongpeng Zhao, Jun Zhou +10

Multimodal Large Language Models (MLLMs) have recently achieved substantial progress in general-purpose perception and reasoning. Nevertheless, their deployment in Food-Service and…

cs.CV2026

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

Jing Tan, Zhaoyang Zhang, Yantao Shen +6

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects…

cs.CL2025

Ideology as a Problem: Lightweight Logit Steering for Annotator-Specific Alignment in Social Media Analysis

Wei Xia, Haowen Tang, Luozheng Li

LLMs internally organize political ideology along low-dimensional structures that are partially, but not fully aligned with human ideological space. This misalignment is systematic…

cs.CL2025

SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning

Wei Xia, Zhi-Hong Deng

With the rapid advancement of large language models (LLMs), their deployment in real-world applications has become increasingly widespread. LLMs are expected to deliver robust perf…

cs.AI2025

Experience-Guided Adaptation of Inference-Time Reasoning Strategies

Adam Stein, Matthew Trager, Benjamin Bowman +4

Enabling agentic AI systems to adapt their problem-solving approaches based on post-training interactions remains a fundamental challenge. While systems that update and maintain a…

cs.AI2025

e1: Learning Adaptive Control of Reasoning Effort

Michael Kleinman, Matthew Trager, Alessandro Achille +2

Increasing the thinking budget of AI models can significantly improve accuracy, but not all questions warrant the same amount of reasoning. Users may prefer to allocate different a…