8 papers
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
Huyen Nguyen, Haoxuan Zhang, Yang Zhang +2
Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document lengths. We conduct a compre…
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
Huyen Nguyen, Haoxuan Zhang, Yang Zhang +2
Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores wit…
Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator
Haoxuan Zhang, Ruochi Li, Yang Zhang +4
Large language models (LLMs) are increasingly used in scientific research and discovery, supporting tasks ranging from literature retrieval and synthesis to hypothesis generation,…
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
Jinze Li, Yang Zhang, Xin Yang +5
Autonomous LLM agents increasingly operate in long-horizon, interactive settings where success depends on reusing experience accumulated over extended histories. However, existing…
MetaGAI: A Large-Scale and High-Quality Benchmark for Generative AI Model and Data Card Generation
Haoxuan Zhang, Ruochi Li, Yang Zhang +4
The rapid proliferation of Generative AI necessitates rigorous documentation standards for transparency and governance. However, manual creation of Model and Data Cards is not scal…
AdaQE-CG: Adaptive Query Expansion for Web-Scale Generative AI Model and Data Card Generation
Haoxuan Zhang, Ruochi Li, Zhenni Liang +7
Transparent and standardized documentation is essential for building trustworthy generative AI (GAI) systems. However, existing automated methods for generating model and data card…