From the 1 of 5 linked papers with an AI index.
5 papers
SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning
Shiyu Yuan, Sourav Sanjukta Bhabesh, Zhe Wang +3
The paper introduces SD-MAR, a synthetic-data framework and reinforcement‑learning fine‑tuning method to improve vision‑language models' ability to reason analytically across multi…
Reinforcing the Generation Order of Multimodal Masked Diffusion Models
Yidong Ouyang, Zhe Wang, Sourav Bhabesh +1
Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation order…
MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
Haris Riaz, Sourav Bhabesh, Vinayak Arannil +2
Recent smaller language models such Phi-3.5 and Phi-4 rely on synthetic data generated using larger Language models. Questions remain about leveraging synthetic data for other use…
DoPAMine: Domain-specific Pre-training Adaptation from seed-guided data Mining
Vinayak Arannil, Neha Narwal, Sourav Sanjukta Bhabesh +5
Large Language Models (LLMs) have shown remarkable ability to generalize effectively across numerous industry domains while executing a range of tasks. Many of these competencies a…
Towards Building a Robust Toxicity Predictor
Dmitriy Bespalov, Sourav Bhabesh, Yi Xiang +2
Recent NLP literature pays little attention to the robustness of toxicity language predictors, while these systems are most likely to be used in adversarial contexts. This paper pr…