Publications (99)
DynaRetarget: Dynamically-Feasible Retargeting using Sampling-Based Trajectory Optimization
Victor Dhedin, Ilyass Taouil, Shafeef Omar +4
In this paper, we introduce DynaRetarget, a complete pipeline for retargeting human motions to humanoid control policies. The core component of DynaRetarget is a novel Sampling-Bas…
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Zhiwei He, Tian Liang, Jiahao Xu +12
Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is suffic…
Document-Level Machine Translation with Large Language Models
Longyue Wang, Chenyang Lyu, Tianbo Ji +4
Large language models (LLMs) such as ChatGPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks. Taking document-level…
User independent Emotion Recognition with Residual Signal-Image Network
Guanghao Yin, Shouqian Sun, Hui Zhang +4
User independent emotion recognition with large scale physiological signals is a tough problem. There exist many advanced methods but they are conducted under relatively small data…
Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation
Dian Yu, Qingchuan Zhou, Bingkun Huang +2
The paper introduces Safe-Night VLA, a robot manipulation system that combines long-wave infrared thermal sensing with a vision‑language backbone and adds safety guarantees via con…
MIDAS: A Dialog Act Annotation Scheme for Open Domain Human Machine Spoken Conversations
Dian Yu, Zhou Yu
Dialog act prediction is an essential language comprehension task for both dialog system building and discourse analysis. Previous dialog act schemes, such as SWBD-DAMSL, are desig…
C-MORE: Pretraining to Answer Open-Domain Questions by Consulting Millions of References
Xiang Yue, Xiaoman Pan, Wenlin Yao +3
We consider the problem of pretraining a two-stage open-domain question answering (QA) system (retriever + reader) with strong transfer capabilities. The key challenge is how to co…
Gunrock: A Social Bot for Complex and Engaging Long Conversations
Dian Yu, Michelle Cohn, Yi Mang Yang +12
Gunrock is the winner of the 2018 Amazon Alexa Prize, as evaluated by coherence and engagement from both real users and Amazon-selected expert conversationalists. We focus on under…
Cross-Lingual Speaker Identification Using Distant Supervision
Ben Zhou, Dian Yu, Dong Yu +1
Speaker identification, determining which character said each utterance in literary text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-…
Price Interpretability of Prediction Markets: A Convergence Analysis
Dian Yu, Jianjun Gao, Weiping Wu +1
Prediction markets are long known for prediction accuracy. This study systematically explores the fundamental properties of prediction markets, addressing questions about their inf…
High-magnitude, spatially programmable, and sustained strain engineering of 2D semiconductors
Boran Kumral, Pedro Guerra Demingos, Peter Serles +13
Crystalline two-dimensional (2D) semiconductors often combine high elasticity and in-plane strength, making them ideal for strain-induced tuning of electronic characteristics, akin…
Language Embeddings for Typology and Cross-lingual Transfer Learning
Dian Yu, Taiqi He, Kenji Sagae
Cross-lingual language tasks typically require a substantial amount of annotated data or parallel translation data. We explore whether language representations that capture relatio…
DREAM: A Challenge Dataset and Models for Dialogue-Based Reading Comprehension
Kai Sun, Dian Yu, Jianshu Chen +3
We present DREAM, the first dialogue-based multiple-choice reading comprehension dataset. Collected from English-as-a-foreign-language examinations designed by human experts to eva…
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Xingyu Chen, Jiahao Xu, Tian Liang +11
The remarkable performance of models like the OpenAI o1 can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended c…
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu +4
While large language models (LLMs) have demonstrated impressive capabilities across tasks in language understanding and interactive decision making, their abilities for reasoning (…
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, Jeffrey Zhao +4
Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making proce…
Improving Pre-Trained Multilingual Models with Vocabulary Expansion
Hai Wang, Dian Yu, Kai Sun +2
Recently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks. However, in multilingual setting, it is extremely reso…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
MinT: Boosting Generalization in Mathematical Reasoning via Multi-View Fine-Tuning
Zhenwen Liang, Dian Yu, Xiaoman Pan +4
Reasoning in mathematical domains remains a significant challenge for relatively small language models (LMs). Many current methods focus on specializing LMs in mathematical reasoni…
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
Zhenwen Liang, Ruosen Li, Yujun Zhou +5
Assessing the quality of Large Language Model (LLM) outputs presents a critical challenge. Previous methods either rely on text-level information (e.g., reward models, majority vot…
Dialogue-Based Relation Extraction
Dian Yu, Kai Sun, Claire Cardie +1
We present the first human-annotated dialogue-based relation extraction (RE) dataset DialogRE, aiming to support the prediction of relation(s) between two arguments that appear in…
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
Xiyao Wang, Linfeng Song, Ye Tian +5
Monte Carlo Tree Search (MCTS) has recently emerged as a powerful technique for enhancing the reasoning capabilities of LLMs. Techniques such as SFT or DPO have enabled LLMs to dis…
Findings of the WMT 2023 Shared Task on Discourse-Level Literary Translation: A Fresh Orb in the Cosmos of LLMs
Longyue Wang, Zhaopeng Tu, Yan Gu +14
Translating literary works has perennially stood as an elusive dream in machine translation (MT), a journey steeped in intricate challenges. To foster progress in this domain, we h…
NarraSum: A Large-Scale Dataset for Abstractive Narrative Summarization
Chao Zhao, Faeze Brahman, Kaiqiang Song +3
Narrative summarization aims to produce a distilled version of a narrative to describe its most salient events and characters. Summarizing a narrative is challenging as it requires…
Retrieval Augmented End-to-End Spoken Dialog Models
Mingqiu Wang, Izhak Shafran, Hagen Soltau +4
We recently developed SLM, a joint speech and language model, which fuses a pretrained foundational speech model and a large language model (LLM), while preserving the in-context l…
Automatically Exposing Problems with Neural Dialog Models
Dian Yu, Kenji Sagae
Neural dialog models are known to suffer from problems such as generating unsafe and inconsistent responses. Even though these problems are crucial and prevalent, they are mostly m…
Attribute Alignment: Controlling Text Generation from Pre-trained Language Models
Dian Yu, Zhou Yu, Kenji Sagae
Large language models benefit from training with a large amount of unlabeled text, which gives them increasingly fluent and diverse generation capabilities. However, using these mo…
A Efficient Multimodal Framework for Large Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals
Guanghao Yin, Shouqian Sun, Dian Yu +2
Considerable attention has been paid for physiological signal-based emotion recognition in field of affective computing. For the reliability and user friendly acquisition, Electrod…
Improving Machine Reading Comprehension with General Reading Strategies
Kai Sun, Dian Yu, Dong Yu +1
Reading strategies have been shown to improve comprehension levels, especially for readers lacking adequate prior knowledge. Just as the process of knowledge accumulation is time-c…
Metasurface-Integrated Polarization-Insensitive LCoS for Projection Displays
Xiangnian Ou, Yueqiang Hu, Dian Yu +8
Liquid crystal on silicon (LCoS) panels, renowned for their high resolution and fill-factor, are integral to modern projection displays. However, their inherent polarization sensit…
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
Yujun Zhou, Zhenwen Liang, Haolin Liu +7
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…
Scaling Synthetic Data Creation with 1,000,000,000 Personas
Tao Ge, Xin Chan, Xiaoyang Wang +3
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully expl…
Guided Discovery of New Behaviors using Diffusion Policies
Dian Yu, Sebastian Sanokowski, Majid Khadiv
Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajectory distributions. However,…
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
Rui Liu, Dian Yu, Zhenwen Liang +6
Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exploiting language priors over…
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
Haolin Liu, Dian Yu, Sidi Lu +6
Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…
One Token to Fool LLM-as-a-Judge
Yulai Zhao, Haolin Liu, Dian Yu +4
Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…
ZeroKBC: A Comprehensive Benchmark for Zero-Shot Knowledge Base Completion
Pei Chen, Wenlin Yao, Hongming Zhang +4
Knowledge base completion (KBC) aims to predict the missing links in knowledge graphs. Previous KBC tasks and approaches mainly focus on the setting where all test entities and rel…
Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
Zhen Zhao, Zicheng Wang, Longyue Wang +3
Semi-supervised medical image segmentation studies have shown promise in training models with limited labeled data. However, current dominant teacher-student based approaches can s…
Few-shot Intent Classification and Slot Filling with Retrieved Examples
Dian Yu, Luheng He, Yuan Zhang +3
Few-shot learning arises in important practical scenarios, such as when a natural language understanding system needs to learn new semantic labels for an emerging, resource-scarce…
UniConFlow: A Unified Constrained Flow-Matching Framework for Certified Motion Planning
Zewen Yang, Xiaobing Dai, Dian Yu +4
Generative models have become increasingly powerful tools for robot motion generation, enabling flexible and multimodal trajectory generation across various tasks. Yet, most existi…
LiteSearch: Efficacious Tree Search for LLM
Ante Wang, Linfeng Song, Ye Tian +5
Recent research suggests that tree search algorithms (e.g. Monte Carlo Tree Search) can dramatically boost LLM performance on complex mathematical reasoning tasks. However, they of…
Unsupervised Slot Schema Induction for Task-oriented Dialog
Dian Yu, Mingqiu Wang, Yuan Cao +3
Carefully-designed schemas describing how to collect and annotate dialog corpora are a prerequisite towards building task-oriented dialog systems. In practical applications, manual…
Description-Driven Task-Oriented Dialog Modeling
Jeffrey Zhao, Raghav Gupta, Yuan Cao +6
Task-oriented dialogue (TOD) systems are required to identify key information from conversations for the completion of given tasks. Such information is conventionally specified in…
Filling Conversation Ellipsis for Better Social Dialog Understanding
Xiyuan Zhang, Chengxi Li, Dian Yu +2
The phenomenon of ellipsis is prevalent in social conversations. Ellipsis increases the difficulty of a series of downstream language understanding tasks, such as dialog act predic…
Reinforcing Multimodal Reasoning Against Visual Degradation
Rui Liu, Dian Yu, Haolin Liu +6
Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies remain brittle against real-wor…
Millimeter Wave Wireless Communications: New Results for Rural Connectivity
George R. MacCartney, Shu Sun, Theodore S. Rappaport +5
This paper shows the remarkable distances that can be achieved using millimeter wave communications, and presents a new rural macrocell (RMa) path loss model for millimeter wave fr…
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Yucheng Shi, Zhenwen Liang, Kishan Panaganti +3
We pursue a vision for self-improving language models in which the model does not merely generate problems or traces to imitate, but constructs the environments that train it. In z…
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…
Using Chatbots to Teach Languages
Yu Li, Chun-Yen Chen, Dian Yu +6
This paper reports on progress towards building an online language learning tool to provide learners with conversational experience by using dialog systems as conversation practice…
Self-Teaching Machines to Read and Comprehend with Large-Scale Multi-Subject Question-Answering Data
Dian Yu, Kai Sun, Dong Yu +1
In spite of much recent research in the area, it is still unclear whether subject-area question-answering data is useful for machine reading comprehension (MRC) tasks. In this pape…
Exploration Based Language Learning for Text-Based Games
Andrea Madotto, Mahdi Namazifar, Joost Huizinga +7
This work presents an exploration and imitation-learning-based agent capable of state-of-the-art performance in playing text-based computer games. Text-based computer games describ…
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
Yuheng Zhang, Dian Yu, Tao Ge +5
Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment…
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…
Gunrock 2.0: A User Adaptive Social Conversational System
Kaihui Liang, Austin Chau, Yu Li +8
Gunrock 2.0 is built on top of Gunrock with an emphasis on user adaptation. Gunrock 2.0 combines various neural natural language understanding modules, including named entity detec…
Improving Machine Reading Comprehension with Contextualized Commonsense Knowledge
Kai Sun, Dian Yu, Jianshu Chen +2
In this paper, we aim to extract commonsense knowledge to improve machine reading comprehension. We propose to represent relations implicitly by situating structured knowledge in a…
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
Dian Yu, Yulai Zhao, Kishan Panaganti +3
We propose Reinforcement Learning with Explicit Human Values (RLEV), a method that aligns Large Language Model (LLM) optimization directly with quantifiable human value signals. Wh…
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
Murong Yue, Wenlin Yao, Haitao Mi +3
Enhancing the capability of large language models (LLMs) in reasoning has gained significant attention in recent years. Previous studies have demonstrated the effectiveness of vari…
Towards Understanding What Code Language Models Learned
Toufique Ahmed, Dian Yu, Chengxuan Huang +3
Pre-trained language models are effective in a variety of natural language tasks, but it has been argued their capabilities fall short of fully learning meaning or understanding la…
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
Rui Liu, Dian Yu, Lei Ke +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalen…
Investigating Prior Knowledge for Challenging Chinese Machine Reading Comprehension
Kai Sun, Dian Yu, Dong Yu +1
Machine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chine…
GR2 Technical Report
Yufei Li, Zaiwei Zhang, Mingfu Liang +67
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step dispropo…
Teaching Pretrained Models with Commonsense Reasoning: A Preliminary KB-Based Approach
Shiyang Li, Jianshu Chen, Dian Yu
Recently, pretrained language models (e.g., BERT) have achieved great success on many downstream natural language understanding tasks and exhibit a certain level of commonsense rea…
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Runpeng Dai, Linfeng Song, Haolin Liu +8
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for enhancing the reasoning ability of Large Language Models (LLMs). Yet current RLVR methods often exp…
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
Ante Wang, Linfeng Song, Ye Tian +6
Recent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increas…
Evidence Sentence Extraction for Machine Reading Comprehension
Hai Wang, Dian Yu, Kai Sun +4
Remarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks. However, it is still difficult to interpret the predictions of…
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Yi Su, Dian Yu, Linfeng Song +5
Reinforcement learning with verifiable rewards (RLVR) has demonstrated significant success in enhancing mathematical reasoning and coding performance of large language models (LLMs…
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
Dian Yu, Baolin Peng, Ye Tian +3
There is a growing trend of teaching large language models (LLMs) to solve mathematical problems through coding. Existing studies primarily focus on prompting powerful, closed-sour…
Disco-Bench: A Discourse-Aware Evaluation Benchmark for Language Modelling
Longyue Wang, Zefeng Du, Donghuai Liu +7
Modeling discourse -- the linguistic phenomena that go beyond individual sentences, is a fundamental yet challenging aspect of natural language processing (NLP). However, existing…
Knowledge-grounded Dialog State Tracking
Dian Yu, Mingqiu Wang, Yuan Cao +3
Knowledge (including structured knowledge such as schema and ontology, and unstructured knowledge such as web corpus) is a critical part of dialog understanding, especially for uns…
Dependency Parsing for Spoken Dialog Systems
Sam Davidson, Dian Yu, Zhou Yu
Dependency parsing of conversational input can play an important role in language understanding for dialog systems by identifying the relationships between entities extracted from…
Skills-in-Context Prompting: Unlocking Compositionality in Large Language Models
Jiaao Chen, Xiaoman Pan, Dian Yu +4
We investigate how to elicit compositional generalization capabilities in large language models (LLMs). Compositional generalization empowers LLMs to solve complex problems by comb…
CLUE: A Chinese Language Understanding Evaluation Benchmark
Liang Xu, Hai Hu, Xuanwei Zhang +29
The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These com…
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Yue Wang, Qiuzhi Liu, Jiahao Xu +11
Large language models (LLMs) such as OpenAI's o1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep think…
Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations
Sihao Chen, Hongming Zhang, Tong Chen +7
We introduce sub-sentence encoder, a contrastively-learned contextual embedding model for fine-grained semantic representation of text. In contrast to the standard practice with se…
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
Zhenwen Liang, Dian Yu, Wenhao Yu +4
Large language models (LLMs) have demonstrated impressive capabilities in mathematical problem solving, particularly in single turn question answering formats. However, real world…
SafeFlow: Safe Robot Motion Planning with Flow Matching via Control Barrier Functions
Xiaobing Dai, Zewen Yang, Dian Yu +4
Recent advances in generative modeling have led to promising results in robot motion planning, particularly through diffusion and flow matching (FM)-based models that capture compl…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Neural Network-Assisted End-to-End Design for Dispersive Full-Parameter Control of Meta-Optics
Hanbin Chi, Yueqiang Hu, Xiangnian Ou +7
Flexible control light field across multiple parameters is the cornerstone of versatile and miniaturized optical devices. Metasurfaces, comprising subwavelength scatterers, offer a…
Improving Question Answering with External Knowledge
Xiaoman Pan, Kai Sun, Dian Yu +4
We focus on multiple-choice question answering (QA) tasks in subject areas such as science, where we require both broad background knowledge and the facts from the given subject-ar…
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Yuheng Zhang, Dian Yu, Baolin Peng +6
Reinforcement Learning with Human Feedback (RLHF) has achieved great success in aligning large language models (LLMs) with human preferences. Prevalent RLHF approaches are reward-b…
Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense Inventories
Wenlin Yao, Xiaoman Pan, Lifeng Jin +3
Word Sense Disambiguation (WSD) aims to automatically identify the exact meaning of one word according to its context. Existing supervised models struggle to make correct predictio…
Zemi: Learning Zero-Shot Semi-Parametric Language Models from Multiple Tasks
Zhenhailong Wang, Xiaoman Pan, Dian Yu +3
Although large language models have achieved impressive zero-shot ability, the huge model size generally incurs high cost. Recently, semi-parametric language models, which augment…
Conceptual and Unbiased Reasoning in Language Models
Ben Zhou, Hongming Zhang, Sihao Chen +5
Conceptual reasoning, the ability to reason in abstract and high-level perspectives, is key to generalization in human cognition. However, limited study has been done on large lang…
Evaluation of PID Performance at CEPC and Optimization with Combined dN/dx and Time-of-Flight Data
Dian Yu, Houqian Ding, Yongfeng Zhu +3
Charged-hadron identification (PID) is a critical requirement for the physics program of the Circular Electron-Positron Collider (CEPC). The baseline detector relies on ionization…
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
Xiaoyang Wang, Hongming Zhang, Tao Ge +3
Customizable role-playing in large language models (LLMs), also known as character generalization, is gaining increasing attention for its versatility and cost-efficiency in develo…
PAC-DP: PAC-Bayesian Diffusion Policy Learning
Mohammad Hasan Yeganegi, Dian Yu, Andrea Del Prete +2
Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which provides limited control over…
From Coated to Uncoated: Scanning Electron Microscopy Corrections to Estimate True Surface Pore Size in Nanoporous Membranes
Sima Zeinali Danalou, Dian Yu, Niher R. Sarker +4
Scanning electron microscopy (SEM) is the premier method for characterizing the nanoscale surface pores in ultrafiltration (UF) membranes and the support layers of reverse osmosis…
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
Jianhui Pang, Fanghua Ye, Longyue Wang +4
The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017), which have acted as benchmarks for progress in…
Teaching LLMs to Refine with Tools
Dian Yu, Yuheng Zhang, Jiahao Xu +5
Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods p…
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
Rui Liu, Dian Yu, Tong Zheng +8
Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
Zhihan Zhang, Tao Ge, Zhenwen Liang +5
Supervised fine-tuning enhances the problem-solving abilities of language models across various mathematical reasoning tasks. To maximize such benefits, existing research focuses o…
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Ye Tian, Baolin Peng, Linfeng Song +4
Despite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Recent work p…
SLM: Bridge the thin gap between speech and text foundation models
Mingqiu Wang, Wei Han, Izhak Shafran +15
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM…
Self-Rewarding Vision-Language Model via Reasoning Decomposition
Zongxia Li, Wenhao Yu, Chengsong Huang +8
Vision-Language Models (VLMs) often suffer from visual hallucinations: generating things that are not consistent with visual inputs and language shortcuts, where they skip the visu…
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
Xiaoman Pan, Wenlin Yao, Hongming Zhang +3
Fully-parametric language models generally require a huge number of model parameters to store the necessary knowledge for solving multiple natural language tasks in zero/few-shot s…
Recurrent Chunking Mechanisms for Long-Text Machine Reading Comprehension
Hongyu Gong, Yelong Shen, Dian Yu +2
In this paper, we study machine reading comprehension (MRC) on long texts, where a model takes as inputs a lengthy document and a question and then extracts a text span from the do…
DarkSHINE Baseline Design Report: Physics Prospects and Detector Technologies
Jing Chen, Ji-Yuan Chen, Jun-Feng Chen +39
DarkSHINE is a newly proposed fixed-target experiment initiative to search for the invisible decay of Dark Photon via missing energy/momentum signatures, based on the high repetiti…
Learning-by-Narrating: Narrative Pre-Training for Zero-Shot Dialogue Comprehension
Chao Zhao, Wenlin Yao, Dian Yu +3
Comprehending a dialogue requires a model to capture diverse kinds of key information in the utterances, which are either scattered around or implicitly implied in different turns…
Speech-to-Text Adapter and Speech-to-Entity Retriever Augmented LLMs for Speech Understanding
Mingqiu Wang, Izhak Shafran, Hagen Soltau +4
Large Language Models (LLMs) have been applied in the speech domain, often incurring a performance drop due to misaligned between speech and language representations. To bridge thi…