5 papers
Beyond Catalogue Counts: the Dataset Visibility Asymmetry in Low-Resource Multilingual NLP
Zhiyin Tan, Changxu Duan
Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record onl…
Semantically Orthogonal Framework for Citation Classification: Disentangling Intent and Content
Changxu Duan, Zhiyin Tan
Understanding the role of citations is essential for research assessment and citation-aware digital libraries. However, existing citation classification frameworks often conflate c…
Multi-Disciplinary Dataset Discovery from Citation-Verified Literature Contexts
Zhiyin Tan, Changxu Duan
Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often…
Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
Changxu Duan
Converting data from machine-unreadable formats like PDFs into Markdown has the potential to enhance the accessibility of scientific research. Existing end-to-end decoder transform…
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
Changxu Duan
Academic documents stored in PDF format can be transformed into plain text structured markup languages to enhance accessibility and enable scalable digital library workflows. Marku…