12 papers
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Runming He, Zhen Hao Wong, Hao Liang +4
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as per…
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
Qijie You, Hao Liang, Mingrui Chen +4
As video becomes increasingly central to information dissemination and multimodal large language models (MLLMs) continue to advance, evaluating video retrieval has become increasin…
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Hao Liang, Qihan Lin, Zhaoyang Han +5
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is str…
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
Zhen Hao Wong, Jingwen Deng, Yuzhao Wang +6
Textbooks are among the richest repositories of human-verified reasoning knowledge, yet their complex layouts contain multi-column typesetting, cross-page question answer separatio…
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
Hao Liang, Zhengyang Zhao, Meiyi Qiang +22
Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, a…