2 papers
cs.CL2026
ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
Zhongyi Zhou, Kohei Uehara, Haoyu Zhang +5
Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS). This leads to inevitable anno…
cs.CV2025
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando +3
This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integr…