24 papers
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns
Yijie Wang, Zhen-Yu Yin, Zhenheng Tang +1
Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of wor…
SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction
Guofan Yu, Sitian Chen, Zhenheng Tang +2
Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-ass…
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang +4
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluatio…
CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection
Qihua Pan, Zhenheng Tang, Peijie Dong +4
Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing…
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
Xiaoyan Su, Peijie Dong, Zhenheng Tang +8
Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for profess…
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
Xiang Liu, Shimiao Yuan, Zhenheng Tang +5
LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…