activity
20122025
most citedCulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

19 citations · 109 across the 39 of their papers we have counts for

collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2025

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi +8

Since the SciCap datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the fir…

cs.CL2025

On Mechanistic Circuits for Extractive Question-Answering

Samyadeep Basu, Vlad Morariu, Zichao Wang +4

Large language models are increasingly used to process documents and facilitate question-answering on them. In our paper, we extract mechanistic circuits for this real-world langua…

cs.CL2025

ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution

Kanika Goswami, Puneet Mathur, Ryan Rossi +1

Large Language Models (LLMs) can perform chart question-answering tasks but often generate unverified hallucinated responses. Existing answer attribution methods struggle to ground…

cs.CL20251 cited

PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback

Kanika Goswami, Puneet Mathur, Ryan Rossi +1

Scientific data visualization is pivotal for transforming raw data into comprehensible visual representations, enabling pattern recognition, forecasting, and the presentation of da…

cs.CL2024

A Framework for Fine-Tuning LLMs using Heterogeneous Feedback

Ryan Aponte, Ryan A. Rossi, Shunan Guo +5

Large language models (LLMs) have been applied to a wide range of tasks, including text summarization, web navigation, and chatbots. They have benefitted from supervised fine-tunin…

cs.CL2024

Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models

Minh Nguyen, Franck Dernoncourt, Seunghyun Yoon +6

We introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives. Despi…