6 papers
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
Quanquan Li, Hongbo Zhang, Yihe Chi +9
Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods either use LLM-generated inter…
TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for Masked Diffusion Language Models
Zhengtao Yao, Liuyang Song, Hongbo Zhang +4
Masked diffusion language models (MDLMs) such as LLaDA now rival autoregressive (AR) LLMs, but every existing knowledge-editing and unlearning method (ROME, MEMIT, etc.) targets AR…
PersonaTree: Structured Lifecycle Memory for Person Understanding in LLM Agents
Yubo Hou, Jingwei Song, Hongbo Zhang +4
Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing agent memory methods emphasize…
LLM-based MOFs Synthesis Condition Extraction using Few-Shot Demonstrations
Lei Shi, Zhimeng Liu, Yi Yang +13
The extraction of Metal-Organic Frameworks (MOFs) synthesis route from literature has been crucial for the logical MOFs design with desirable functionality. The recent advent of la…
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
Hongbo Zhang, Han Cui, Guangsheng Bao +3
We introduce Direct Value Optimization (DVO), an innovative reinforcement learning framework for enhancing large language models in complex reasoning tasks. Unlike traditional meth…
How Likely Do LLMs with CoT Mimic Human Reasoning?
Guangsheng Bao, Hongbo Zhang, Cunxiang Wang +2
Chain-of-thought emerges as a promising technique for eliciting reasoning capabilities from Large Language Models (LLMs). However, it does not always improve task performance or ac…