NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (67)

cs.CL2026

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7

The paper introduces DeepSearch-World, a deterministic, verifiable web environment, and DeepSearch-Evolve, a self‑distillation framework that lets web search agents improve from th…

#self-distillation#web search agents#verifiable environment#multi-hop QA
cs.CL2026

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

Dadi Guo, Yuejin Xie, Qingyu Liu +11

cs.CV2025

CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions

Yuchen Huang, Zhiyuan Fan, Zhitao He +3

cs.CL2025

The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

Yuji Zhang, Sha Li, Cheng Qian +8

cs.CL2025

DocCHA: Towards LLM-Augmented Interactive Online diagnosis System

Xinyi Liu, Dachun Sun, Yi R. Fung +2

cs.CL2026

Diversity-Enhanced Reasoning for Subjective Questions

Yumeng Wang, Zhiyuan Fan, Jiayu Liu +2

cs.CL2024

CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets

Lifan Yuan, Yangyi Chen, Xingyao Wang +3

cs.AI2025

EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce

Rui Min, Zile Qiao, Ze Xu +18

cs.AI2025

Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability

Ruida Wang, Yuxin Li, Yi R. Fung +1

cs.CL2023

Enhanced Chart Understanding in Vision and Language Task via Cross-modal Pre-training on Plot Table Pairs

Mingyang Zhou, Yi R. Fung, Long Chen +3

cs.CV2026

EMCompress: Video-LLMs with Endomorphic Multimodal Compression

Zheyu Fan, Jiateng Liu, Yuji Zhang +4

cs.CL2025

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?

Zhitao He, Zongwei Lyu, Dazhong Chen +2

cs.CL2025

Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization

Shujin Wu, Cheng Qian, Yi R. Fung +2

cs.AI2026

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

Jiayu Liu, Cheng Qian, Zhaochen Su +4

cs.CL2025

Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data Refinement

Chenkai Sun, Ke Yang, Revanth Gangi Reddy +5

cs.AI2025

UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios

Haotian Luo, Huaisong Zhang, Xuelin Zhang +15

cs.CL2021

COVID-19 Literature Knowledge Graph Construction and Drug Repurposing Report Generation

Qingyun Wang, Manling Li, Xuan Wang +24

cs.CL2024

Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models

Yuji Zhang, Sha Li, Jiateng Liu +5

cs.LG2026

Empowering Reliable Visual-Centric Instruction Following in MLLMs

Weilei He, Feng Ju, Zhiyuan Fan +3

cs.CV2025

MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing

Minghao Liu, Zhitao He, Zhiyuan Fan +2

cs.AI2026

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

Yuchen Huang, Xiang Li, Zhenqing Ling +5

cs.CL2024

MACAROON: Training Vision-Language Models To Be Your Engaged Partners

Shujin Wu, Yi R. Fung, Sha Li +3

cs.CL2024

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning

Kung-Hsiang Huang, Mingyang Zhou, Hou Pong Chan +5

cs.CL2024

LEMMA: Towards LVLM-Enhanced Multimodal Misinformation Detection with External Knowledge Augmentation

Keyang Xuan, Li Yi, Fan Yang +3

cs.CL2025

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Jianshu Zhang, Dongyu Yao, Renjie Pi +2

cs.CV2025

Video Signature: Implicit Watermarking for Video Diffusion Models

Yu Huang, Junhao Chen, Shuliang Liu +6

cs.CL2024

From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models

Kung-Hsiang Huang, Hou Pong Chan, Yi R. Fung +5

cs.CL2026

Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority

Zhanming Shen, Zeyu Qin, Jiaqi Hu +7

cs.CL2025

Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization

Zhitao He, Zijun Liu, Peng Li +5

cs.CL2026

On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility

Zhitao He, Haolin Yang, Rui Min +2

cs.CL2024

CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models

Cheng Qian, Chi Han, Yi R. Fung +3

cs.CV2025

TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models

Jaewoo Lee, Keyang Xuan, Chanakya Ekbote +3

cs.CL2022

NewsClaims: A New Benchmark for Claim Detection from News with Attribute Knowledge

Revanth Gangi Reddy, Sai Chetan, Zhenhailong Wang +8

cs.AI2026

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Guanyu Jiang, Zhaochen Su, Xiaoye Qu +1

cs.CL2026

Scalable Token-Level Hallucination Detection in Large Language Models

Rui Min, Tianyu Pang, Chao Du +2

cs.CV2025

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

Zhiyuan Fan, Yumeng Wang, Sandeep Polisetty +1

cs.CL2026

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces

Xinyu Geng, Yanjing Xiao, Yuyang Zhang +5

cs.CL2025

Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning

Qi He, Cheng Qian, Xiusi Chen +3

cs.CL2023

Defining a New NLP Playground

Sha Li, Chi Han, Pengfei Yu +8

cs.AI2025

Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

Dadi Guo, Jiayu Liu, Zhiyuan Fan +5

cs.CL2024

R-Tuning: Instructing Large Language Models to Say `I Don't Know'

Hanning Zhang, Shizhe Diao, Yong Lin +6

cs.CL2026

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

Haolin Yang, Jipeng Zhang, Zhitao He +2

cs.CL2026

Reinforcement Learning from Denoising Feedback

Qi He, Huan Chen, Ya Guo +3

cs.CV2025

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Zhaochen Su, Peng Xia, Hangyu Guo +12

cs.CL2023

Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response Forecasting

Chenkai Sun, Jinning Li, Yi R. Fung +4

cs.CV2026

AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios

Zhaochen Su, Jincheng Gao, Hangyu Guo +10

cs.AI2025

Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4

Yuxin Li, Minghao Liu, Ruida Wang +6

cs.CR2026

RedAct: Redacting Agent Capability Traces for Procedural Skill Protection

Shuwen Xu, Zhitao He, Yi R. Fung

cs.CL2025

Scaling Laws of Synthetic Data for Language Models

Zeyu Qin, Qingxiu Dong, Xingxing Zhang +10

cs.DL2026

CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation

Yee Man Choi, Xuehang Guo, Yi R. Fung +1

cs.CL2025

MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models

Hengzhi Li, Megan Tjandrasuwita, Yi R. Fung +2

cs.CL2024

NormSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-Fly

Yi R. Fung, Tuhin Chakraborty, Hao Guo +3

cs.CV2026

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

Yiyang Fang, Pei Fu, Jinjie Li +7

cs.AI2026

Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm

Dadi Guo, Tianyi Zhou, Dongrui Liu +8

cs.CL2024

If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Ke Yang, Jiateng Liu, John Wu +9

cs.LG2026

On the Geometry of On-Policy Distillation

Zhennan Shen, Yanshu Li, Qingyu Yin +6

cs.SE2025

SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation

Yixiang Chen, Tianshi Zheng, Shijue Huang +2

cs.AI2025

Are Your Agents Upward Deceivers?

Dadi Guo, Qingyu Liu, Dongrui Liu +13

cs.CL2026

RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection

Yiming Huang, Junyan Zhang, Zihao Wang +5

cs.AI2025

AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting

Shijue Huang, Hongru Wang, Wanjun Zhong +4

cs.CL2026

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Jiayu Liu, Cheng Qian, Zhenhailong Wang +10

cs.SI2022

A Weibo Dataset for the 2022 Russo-Ukrainian Crisis

Yi R. Fung, Heng Ji

cs.CL2024

SmartBook: AI-Assisted Situation Report Generation for Intelligence Analysts

Revanth Gangi Reddy, Daniel Lee, Yi R. Fung +6

cs.CL2026

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Shijue Huang, Hangyu Guo, Guanting Dong +8

cs.CL2025

MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration

Zhitao He, Sandeep Polisetty, Zhiyuan Fan +3

cs.CL2026

Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking

Feng Ju, Zeyu Qin, Rui Min +3

cs.LG2025

Environment Scaling for Interactive Agentic Experience Collection: A Survey

Yuchen Huang, Sijia Li, Minghao Liu +5