Publications (33)
MUSE: Textual Attributes Guided Portrait Painting Generation
Xiaodan Hu, Pengfei Yu, Kevin Knight +3
We propose a novel approach, MUSE, to illustrate textual attributes visually via portrait generation. MUSE takes a set of attributes written in text, in addition to facial features…
FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation
Xu Han, Hao Zhu, Pengfei Yu +4
We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation…
OSExpert: Computer-Use Agents Learning Professional Skills via Exploration
Jiateng Liu, Zhenhailong Wang, Rushi Wang +6
General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less…
Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production
Kexuan Xin, Qingyun Wang, Junyu Chen +3
In the rapidly evolving field of metabolic engineering, the quest for efficient and precise gene target identification for metabolite production enhancement presents significant ch…
MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
Jiawei Chen, Xintian Shen, Lihao Zheng +43
Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of auto…
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
Pengfei Yu, Dongming Shen, Silin Meng +8
We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…
StreamingClaw Technical Report
Jiawei Chen, Zhe Chen, Chaoqun Du +21
Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing st…
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Chengjun Yu, Xuhan Zhu, Chaoqun Du +4
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…
Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models
Andrea Yaoyun Cui, Pengfei Yu
Language models are essentially probability distributions over token sequences. Auto-regressive models generate sentences by iteratively computing and sampling from the distributio…
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
Yuji Zhang, Sha Li, Cheng Qian +8
Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which err…
Defining a New NLP Playground
Sha Li, Chi Han, Pengfei Yu +8
The recent explosion of performance of large language models (LLMs) has changed the field of Natural Language Processing (NLP) more abruptly and seismically than any other shift in…
AuxBlocks: Defense Adversarial Example via Auxiliary Blocks
Yueyao Yu, Pengfei Yu, Wenye Li
Deep learning models are vulnerable to adversarial examples, which poses an indisputable threat to their applications. However, recent studies observe gradient-masking defenses are…
Ordered-subsets Multi-diffusion Model for Sparse-view CT Reconstruction
Pengfei Yu, Bin Huang, Minghui Zhang +3
Score-based diffusion models have shown significant promise in the field of sparse-view CT reconstruction. However, the projection dataset is large and riddled with redundancy. Con…
EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries
Jiateng Liu, Pengfei Yu, Yuji Zhang +3
The dynamic nature of real-world information necessitates efficient knowledge editing (KE) in large language models (LLMs) for knowledge updating. However, current KE approaches, w…
Landau meets Newton: time translation symmetry breaking in classical mechanics
Liu Zhao, Wei Xu, Pengfei Yu
Every classical Newtonian mechanical system can be equipped with a nonstandard Hamiltonian structure, in which the Hamiltonian is the square of the canonical Hamiltonian up to a co…
LogiShot: Logically Coherent Cross-Shot Video Generation
Shuai Guo, Yuhang Yang, Zeyu Zhang +4
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…
Tracking of enriched dialog states for flexible conversational information access
Yinpei Dai, Zhijian Ou, Dawei Ren +1
Dialog state tracking (DST) is a crucial component in a task-oriented dialog system for conversational information access. A common practice in current dialog systems is to define…
Information Association for Language Model Updating by Mitigating LM-Logical Discrepancy
Pengfei Yu, Heng Ji
Large Language Models~(LLMs) struggle with providing current information due to the outdated pre-training data. Existing methods for updating LLMs, such as knowledge editing and co…
SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
Alejandro Cuadron, Pengfei Yu, Yang Liu +1
Despite rapid progress in LLM agents, performance on long-horizon, tool-using tasks remains fragile. To better understand this fragility, we ask a simple question: \emph{do all act…
MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
Wei Chen, Chaoqun Du, Feng Gu +14
We present MindGPT-4ov, a multimodal large language model (MLLM) that introduces a general post-training paradigm spanning data production, model training, and efficient deployment…
Primary damage and mechanical degradation of WTaCrV refractory high-entropy alloy: effects of solid-solution and chemical ordering
Yihan Wu, Pengfei Yu, Yaohong Suo +1
As advanced nuclear reactors demand novel irradiation-tolerant materials, this study investigates the radiation damage and mechanical degradation of the promising WTaCrV refractory…
LDGen: Enhancing Text-to-Image Synthesis via Large Language Model-Driven Language Representation
Pengzhi Li, Pengfei Yu, Zide Liu +5
In this paper, we introduce LDGen, a novel method for integrating large language models (LLMs) into existing text-to-image diffusion models while minimizing computational demands.…
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
Yufei Zheng, Xuhan Zhu, Zide Liu +9
Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…
PET Tracer Conversion among Brain PET via Variable Augmented Invertible Network
Bohui Shen, Wei Zhang, Xubiao Liu +8
Positron emission tomography (PET) serves as an essential tool for diagnosis of encephalopathy and brain science research. However, it suffers from the limited choice of tracers. N…
Size-Dependent Tensile Behavior of Nanocrystalline HfNbTaTiZr High-Entropy Alloy: Roles of Solid-Solution and Short-Range Order
Yihan Wu, Gaosheng Yan, Pengfei Yu +3
This study investigates the size-dependent mechanical behavior of the HfNbTaTiZr refractory high-entropy alloy (RHEA) under uniaxial tension, with a focus on the effects of random…
SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure Bias
Zihao Wang, Qihao Liang, Kejun Zhang +8
Real-time music accompaniment generation has a wide range of applications in the music industry, such as music education and live performances. However, automatic real-time music a…
MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
Zihao Wang, Shuyu Li, Tao Zhang +6
The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music…
RCOT: Detecting and Rectifying Factual Inconsistency in Reasoning by Reversing Chain-of-Thought
Tianci Xue, Ziqi Wang, Zhenhailong Wang +3
Large language Models (LLMs) have achieved promising performance on arithmetic reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting. However, LLMs face ch…
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
Jiaqi Wei, Xuehang Guo, Pengfei Yu +5
In single-stream autoregressive interfaces, the same tokens both update the model state and constitute an irreversible public commitment. This coupling creates a silence tax: addit…
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
Yuji Zhang, Sha Li, Jiateng Liu +5
Hallucination is often regarded as a major impediment for using large language models (LLMs), especially for knowledge-intensive tasks. Even when the training corpus consists solel…
Hamiltonian description of singular Lagrangian systems with spontaneously broken time translation symmetry
Liu Zhao, Pengfei Yu, Wei Xu
Shapere and Wilczek recently found some singular Lagrangian systems which spontaneously breaks time translation symmetry. The common feature of their models is that the energy func…
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
Jiaxin Qin, Zixuan Zhang, Manling Li +2
Extensive previous research has focused on post-training knowledge editing (KE) for language models (LMs) to ensure that knowledge remains accurate and up-to-date. One desired prop…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…