papers

Publications (33)

cs.CV2021

MUSE: Textual Attributes Guided Portrait Painting Generation

Xiaodan Hu, Pengfei Yu, Kevin Knight +3

We propose a novel approach, MUSE, to illustrate textual attributes visually via portrait generation. MUSE takes a set of attributes written in text, in addition to facial features…

cs.LG2018

FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

Xu Han, Hao Zhu, Pengfei Yu +4

We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation…

cs.AI2026

OSExpert: Computer-Use Agents Learning Professional Skills via Exploration

Jiateng Liu, Zhenhailong Wang, Rushi Wang +6

General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less…

cs.AI2024

Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production

Kexuan Xin, Qingyun Wang, Junyu Chen +3

In the rapidly evolving field of metabolic engineering, the quest for efficient and precise gene target identification for metabolite production enhancement presents significant ch…

cs.AI2026

MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning

Jiawei Chen, Xintian Shen, Lihao Zheng +43

Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of auto…

cs.CL2025

RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines

Pengfei Yu, Dongming Shen, Silin Meng +8

We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…

cs.CV2026

StreamingClaw Technical Report

Jiawei Chen, Zhe Chen, Chaoqun Du +21

Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing st…

cs.CV2026

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

Chengjun Yu, Xuhan Zhu, Chaoqun Du +4

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…

cs.CL2025

Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models

Andrea Yaoyun Cui, Pengfei Yu

Language models are essentially probability distributions over token sequences. Auto-regressive models generate sentences by iteratively computing and sampling from the distributio…

cs.CL2025

The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

Yuji Zhang, Sha Li, Cheng Qian +8

Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which err…

cs.CL2023

Defining a New NLP Playground

Sha Li, Chi Han, Pengfei Yu +8

The recent explosion of performance of large language models (LLMs) has changed the field of Natural Language Processing (NLP) more abruptly and seismically than any other shift in…

cs.LG2019

AuxBlocks: Defense Adversarial Example via Auxiliary Blocks

Yueyao Yu, Pengfei Yu, Wenye Li

Deep learning models are vulnerable to adversarial examples, which poses an indisputable threat to their applications. However, recent studies observe gradient-masking defenses are…

eess.IV2025

Ordered-subsets Multi-diffusion Model for Sparse-view CT Reconstruction

Pengfei Yu, Bin Huang, Minghui Zhang +3

Score-based diffusion models have shown significant promise in the field of sparse-view CT reconstruction. However, the projection dataset is large and riddled with redundancy. Con…

cs.CL2024

EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries

Jiateng Liu, Pengfei Yu, Yuji Zhang +3

The dynamic nature of real-world information necessitates efficient knowledge editing (KE) in large language models (LLMs) for knowledge updating. However, current KE approaches, w…

gr-qc2012

Landau meets Newton: time translation symmetry breaking in classical mechanics

Liu Zhao, Wei Xu, Pengfei Yu

Every classical Newtonian mechanical system can be equipped with a nonstandard Hamiltonian structure, in which the Hamiltonian is the square of the canonical Hamiltonian up to a co…

cs.CV2026

LogiShot: Logically Coherent Cross-Shot Video Generation

Shuai Guo, Yuhang Yang, Zeyu Zhang +4

Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…

cs.CL2018

Tracking of enriched dialog states for flexible conversational information access

Yinpei Dai, Zhijian Ou, Dawei Ren +1

Dialog state tracking (DST) is a crucial component in a task-oriented dialog system for conversational information access. A common practice in current dialog systems is to define…

cs.CL2024

Information Association for Language Model Updating by Mitigating LM-Logical Discrepancy

Pengfei Yu, Heng Ji

Large Language Models~(LLMs) struggle with providing current information due to the outdated pre-training data. Existing methods for updating LLMs, such as knowledge editing and co…

cs.LG2025

SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents

Alejandro Cuadron, Pengfei Yu, Yang Liu +1

Despite rapid progress in LLM agents, performance on long-horizon, tool-using tasks remains fragile. To better understand this fragility, we ask a simple question: \emph{do all act…

cs.CV2025

MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm

Wei Chen, Chaoqun Du, Feng Gu +14

We present MindGPT-4ov, a multimodal large language model (MLLM) that introduces a general post-training paradigm spanning data production, model training, and efficient deployment…

cond-mat.mtrl-sci2026

Primary damage and mechanical degradation of WTaCrV refractory high-entropy alloy: effects of solid-solution and chemical ordering

Yihan Wu, Pengfei Yu, Yaohong Suo +1

As advanced nuclear reactors demand novel irradiation-tolerant materials, this study investigates the radiation damage and mechanical degradation of the promising WTaCrV refractory…

cs.CV2025

LDGen: Enhancing Text-to-Image Synthesis via Large Language Model-Driven Language Representation

Pengzhi Li, Pengfei Yu, Zide Liu +5

In this paper, we introduce LDGen, a novel method for integrating large language models (LLMs) into existing text-to-image diffusion models while minimizing computational demands.…

cs.CV2026

Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence

Yufei Zheng, Xuhan Zhu, Zide Liu +9

Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…

cs.LG2023

PET Tracer Conversion among Brain PET via Variable Augmented Invertible Network

Bohui Shen, Wei Zhang, Xubiao Liu +8

Positron emission tomography (PET) serves as an essential tool for diagnosis of encephalopathy and brain science research. However, it suffers from the limited choice of tracers. N…

cond-mat.mtrl-sci2025

Size-Dependent Tensile Behavior of Nanocrystalline HfNbTaTiZr High-Entropy Alloy: Roles of Solid-Solution and Short-Range Order

Yihan Wu, Gaosheng Yan, Pengfei Yu +3

This study investigates the size-dependent mechanical behavior of the HfNbTaTiZr refractory high-entropy alloy (RHEA) under uniaxial tension, with a focus on the effects of random…

cs.SD2022

SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure Bias

Zihao Wang, Qihao Liang, Kejun Zhang +8

Real-time music accompaniment generation has a wide range of applications in the music industry, such as music education and live performances. However, automatic real-time music a…

cs.SD2024

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

Zihao Wang, Shuyu Li, Tao Zhang +6

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music…

cs.CL2023

RCOT: Detecting and Rectifying Factual Inconsistency in Reasoning by Reversing Chain-of-Thought

Tianci Xue, Ziqi Wang, Zhenhailong Wang +3

Large language Models (LLMs) have achieved promising performance on arithmetic reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting. However, LLMs face ch…

cs.CL2026

When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning

Jiaqi Wei, Xuehang Guo, Pengfei Yu +5

In single-stream autoregressive interfaces, the same tokens both update the model state and constitute an irreversible public commitment. This coupling creates a silence tax: addit…

cs.CL2024

Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models

Yuji Zhang, Sha Li, Jiateng Liu +5

Hallucination is often regarded as a major impediment for using large language models (LLMs), especially for knowledge-intensive tasks. Even when the training corpus consists solel…

hep-th2012

Hamiltonian description of singular Lagrangian systems with spontaneously broken time translation symmetry

Liu Zhao, Pengfei Yu, Wei Xu

Shapere and Wilczek recently found some singular Lagrangian systems which spontaneously breaks time translation symmetry. The common feature of their models is that the energy func…

cs.CL2025

Why Does New Knowledge Create Messy Ripple Effects in LLMs?

Jiaxin Qin, Zixuan Zhang, Manling Li +2

Extensive previous research has focused on post-training knowledge editing (KE) for language models (LMs) to ensure that knowledge remains accurate and up-to-date. One desired prop…

cs.CV2026

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

Chenfeng Wang, Wei He, Xuhan Zhu +10

In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…