papers

Publications (37)

cs.AI2025

SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration

Keyan Ding, Jing Yu, Junjie Huang +3

Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools demands substantial domain expertise. While Large Language Models…

cs.AI2026

SkillNet: Create, Evaluate, and Connect AI Skills

Yuan Liang, Ruobin Zhong, Haoming Xu +46

Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Wi…

cs.CL2026

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10

The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…

cs.LG2026

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

Tong Xu, Xinzhe Cao, Zhihui Zhu +2

Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of…

cs.LG2026

KEPLA: A Knowledge-Enhanced Deep Learning Framework for Accurate Protein-Ligand Binding Affinity Prediction

Han Liu, Keyan Ding, Peilin Chen +4

Accurate prediction of protein-ligand binding affinity is critical for drug discovery. While recent deep learning approaches have demonstrated promising results, they often rely so…

cs.CL2026

Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

Yuqi Tang, Kehua Feng, Yunfeng Wang +6

Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, w…

cs.CV2018

A Simple Method to improve Initialization Robustness for Active Contours driven by Local Region Fitting Energy

Keyan Ding, Linfang Xiao

Active contour models based on local region fitting energy can segment images with intensity inhomogeneity effectively, but their segmentation results are easy to error if the init…

cs.CL2025

Advancing biomolecular understanding and design following human instructions

Xiang Zhuang, Keyan Ding, Tianwen Lyu +9

Understanding and designing biomolecules, such as proteins and small molecules, is central to advancing drug discovery, synthetic biology and enzyme engineering. Recent breakthroug…

cs.CL2025

SAFER: Advancing Safety Alignment via Efficient Ex-Ante Reasoning

Kehua Feng, Keyan Ding, Yuhao Wang +5

Recent advancements in large language models (LLMs) have accelerated progress toward artificial general intelligence, yet their potential to generate harmful content poses critical…

cs.CL2025

SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models

Jing Yu, Yuqi Tang, Kehua Feng +8

Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains r…

cs.CV2024

Deep Shape-Texture Statistics for Completely Blind Image Quality Evaluation

Yixuan Li, Peilin Chen, Hanwei Zhu +3

Opinion-Unaware Blind Image Quality Assessment (OU-BIQA) models aim to predict image quality without training on reference images and subjective quality scores. Thereinto, image st…

cs.MM2019

Intrinsic Image Popularity Assessment

Keyan Ding, Kede Ma, Shiqi Wang

The goal of research in automatic image popularity assessment (IPA) is to develop computational models that can accurately predict the potential of a social image to go viral on th…

cs.CV2024

Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics

Zhangkai Ni, Yue Liu, Keyan Ding +3

Deep learning-based methods have significantly influenced the blind image quality assessment (BIQA) field, however, these methods often require training using large amounts of huma…

cs.AI2026

Embodied Science: Closing the Discovery Loop with Agentic Embodied AI

Xiang Zhuang, Chenyi Zhou, Kehua Feng +10

Artificial intelligence has demonstrated remarkable capability in predicting scientific properties, yet scientific discovery remains an inherently physical, long-horizon pursuit go…

cs.CL2025

CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning

Kehua Feng, Keyan Ding, Zhihui Zhu +3

While chain-of-thought (CoT) distillation from advanced large language models (LLMs) has proven effective in general reasoning tasks, it struggles in scientific domains where even…

cs.AI2026

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

Shuofei Qiao, Yunxiang Wei, Jiazheng Fan +8

The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowled…

cs.AI2026

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

Mingyang Rao, Kehua Feng, Zhihui Zhu +4

While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical reaction diagrams. We identi…

cs.CL2024

SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks

Tianhao Li, Jingyu Lu, Chuangxin Chu +12

Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring…

q-bio.BM2023

InstructProtein: Aligning Human and Protein Language via Knowledge Instruction

Zeyuan Wang, Qiang Zhang, Keyan Ding +4

Large Language Models (LLMs) have revolutionized the field of natural language processing, but they fall short in comprehending biological sequences such as proteins. To address th…

cs.CL2025

SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

Kehua Feng, Xinyi Shen, Weijie Wang +4

Large language models (LLMs) are playing an increasingly important role in scientific research, yet there remains a lack of comprehensive benchmarks to evaluate the breadth and dep…

cs.LG2025

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

Kehua Feng, Keyan Ding, Hongzhi Tan +8

Reliable evaluation of large language models (LLMs) is impeded by two key challenges: objective metrics often fail to reflect human perception of natural language, and exhaustive h…

cs.LG2025

Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra

Yiwen Zhang, Keyan Ding, Yihang Wu +4

Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library m…

cs.AI2026

TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

Yuhang Zhang, Keyan Ding, Peilin Chen +5

Enzyme-reaction retrieval is a fundamental problem in computational biology, underpinning enzyme characterization, reaction mechanism elucidation, and the rational design of metabo…

cs.LG2025

MolMark: Safeguarding Molecular Structures through Learnable Atom-Level Watermarking

Runwen Hu, Peilin Chen, Keyan Ding +1

AI-driven molecular generation is reshaping drug discovery and materials design, yet the lack of protection mechanisms leaves AI-generated molecules vulnerable to unauthorized reus…

cs.AI2025

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization

Yuhao Wang, Keyan Ding, Kehua Feng +5

Protein language models have emerged as powerful tools for sequence generation, offering substantial advantages in functional optimization and denovo design. However, these models…

eess.IV2020

Comparison of Image Quality Models for Optimization of Image Processing Systems

Keyan Ding, Kede Ma, Shiqi Wang +1

The performance of objective image quality assessment (IQA) models has been evaluated primarily by comparing model predictions to human quality judgments. Perceptual datasets gathe…

cs.CV2020

Image Quality Assessment: Unifying Structure and Texture Similarity

Keyan Ding, Kede Ma, Shiqi Wang +1

Objective measures of image quality generally operate by comparing pixels of a "degraded" image to those of the original. Relative to human observers, these measures are overly sen…

cs.LG2023

Graph Sampling-based Meta-Learning for Molecular Property Prediction

Xiang Zhuang, Qiang Zhang, Bin Wu +3

Molecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been…

eess.IV2021

Locally Adaptive Structure and Texture Similarity for Image Quality Assessment

Keyan Ding, Yi Liu, Xueyi Zou +2

The latest advances in full-reference image quality assessment (IQA) involve unifying structure and texture similarity based on deep representations. The resulting Deep Image Struc…

cs.CL2026

Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions

Shunyang Luo, Peibei Cao, Zhihui Zhu +3

Reward models (RMs) are central to aligning large language models, yet their practical effectiveness hinges on generalization to unseen prompts and shifting distributions. Most exi…

cs.LG2023

Learning Invariant Molecular Representation in Latent Discrete Space

Xiang Zhuang, Qiang Zhang, Keyan Ding +5

Molecular representation learning lays the foundation for drug discovery. However, existing methods suffer from poor out-of-distribution (OOD) generalization, particularly when dat…

cs.CL2025

ClinDEF: A Dynamic Evaluation Framework for Large Language Models in Clinical Reasoning

Yuqi Tang, Jing Yu, Zichang Su +7

Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through p…

cs.CV2025

AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment

Hanwei Zhu, Yu Tian, Keyan Ding +4

Image quality assessment (IQA) is inherently complex, as it reflects both the quantification and interpretation of perceptual quality rooted in the human visual system. Conventiona…

cs.CL2025

OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases

Yongrui Chen, Zhiqiang Liu, Jing Yu +21

Large Language Models (LLMs) have demonstrated substantial progress on reasoning tasks involving unstructured text, yet their capabilities significantly deteriorate when reasoning…

cs.CL2025

Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning

Xiang Zhuang, Bin Wu, Jiyu Cui +6

Molecular structure elucidation involves deducing a molecule's structure from various types of spectral data, which is crucial in chemical experimental analysis. While large langua…

cs.AI2026

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Yuqi Tang, Chenyi Zhou, Libin Wang +3

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on pred…

cs.AI2025

Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning

Tianwen Lyu, Xiang Zhuang, Keyan Ding +5

Understanding complex biomolecular mechanisms requires multi-step reasoning across molecular interactions, signaling cascades, and metabolic pathways. While large language models(L…