7 papers
StarVector: Generating Scalable Vector Graphics Code from Images and Text
Juan A. Rodriguez, Abhay Puri, Shubham Agarwal +6
Scalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectoriza…
LitLLMs, LLMs for Literature Review: Are we there yet?
Shubham Agarwal, Gaurav Sahu, Abhay Puri +5
Literature reviews are an essential component of scientific research, but they remain time-intensive and challenging to write, especially due to the recent influx of research paper…
LitLLM: A Toolkit for Scientific Literature Review
Shubham Agarwal, Gaurav Sahu, Abhay Puri +5
Conducting literature reviews for scientific papers is essential for understanding research, its limitations, and building on existing work. It is a tedious task which makes an aut…
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
Juan Rodriguez, Xiangru Jian, Siba Smarak Panigrahi +40
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and sum…
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
Shaharukh Khan, Ayush Tarun, Ali Faraz +7
In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We in…
Krutrim LLM: Multilingual Foundational Model for over a Billion People
Aditya Kallappa, Palash Kamble, Abhinav Ravi +10
India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundatio…