papers

Publications (120)

cs.CL2024

Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets

Vatsal Gupta, Pranshu Pandya, Tushar Kataria +2

Language models, characterized by their black-box nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust. To enhance trust, it is im…

cs.CL2021

Incorporating External Knowledge to Enhance Tabular Reasoning

J. Neeraja, Vivek Gupta, Vivek Srikumar

Reasoning about tabular information presents unique challenges to modern NLP approaches which largely rely on pre-trained contextualized embeddings of text. In this paper, we study…

cs.CL2025

PRAISE: Enhancing Product Descriptions with LLM-Driven Structured Insights

Adnan Qidwai, Srija Mukhopadhyay, Prerana Khatiwada +2

Accurate and complete product descriptions are crucial for e-commerce, yet seller-provided information often falls short. Customer reviews offer valuable details but are laborious…

cs.CL2026

DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity

Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi +3

Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbat…

astro-ph.HE2019

Detection of a Glitch in PSR J09084913 by UTMOST

Marcus E. Lower, Matthew Bailes, Ryan M. Shannon +20

We report the first detection of a glitch in the radio pulsar PSR J09084913 (PSR B090649) during regular timing observations by the Molonglo Observatory Synthesis Telescope (…

cs.IR2025

REaR: Retrieve, Expand and Refine for Effective Multitable Retrieval

Rishita Agarwal, Himanshu Singhal, Peter Baile Chen +3

Answering natural language queries over relational data often requires retrieving and reasoning over multiple tables, yet most retrievers optimize only for query-table relevance an…

cs.DB2022

Share the Tensor Tea: How Databases can Leverage the Machine Learning Ecosystem

Yuki Asada, Victor Fu, Apurva Gandhi +8

We demonstrate Tensor Query Processor (TQP): a query processor that automatically compiles relational operators into tensor programs. By leveraging tensor runtimes such as PyTorch,…

astro-ph.CO2020

The Cosmic Dispersion Measure in the EAGLE Simulations

Adam J. Batten, Alan R. Duffy, Nastasha Wijers +4

The dispersion measure (DM) of fast radio bursts (FRBs) provides a unique way to probe ionised baryons in the intergalactic medium (IGM). Cosmological models with different paramet…

cs.AI2025

Weaver: Interweaving SQL and LLM for Table Reasoning

Rohit Khoja, Devanshu Gupta, Yanjie Fu +2

Querying tables with unstructured data is challenging due to the presence of text (or image), either embedded in the table or in external paragraphs, which traditional SQL struggle…

cs.CY2025

REDDIX-NET: A Novel Dataset and Benchmark for Moderating Online Explicit Services

MSVPJ Sathvik, Manan Roy Choudhury, Rishita Agarwal +2

The rise of online platforms has enabled covert illicit activities, including online prostitution, to pose challenges for detection and regulation. In this study, we introduce REDD…

cs.CL2026

CORE-T: COherent REtrieval of Tables for Text-to-SQL

Hassan Soliman, Vivek Gupta, Dan Roth +1

Realistic text-to-SQL workflows often require joining multiple tables. As a result, accurately retrieving the relevant set of tables becomes a key bottleneck for end-to-end perform…

cs.CL2024

Enhancing Question Answering on Charts Through Effective Pre-training Tasks

Ashim Gupta, Vivek Gupta, Shuo Zhang +3

To completely understand a document, the use of textual information is not enough. Understanding visual cues, such as layouts and charts, is also required. While the current state-…

cs.AI2026

OSCAR: Orchestrated Self-verification and Cross-path Refinement

Yash Shah, Abhijit Chakraborty, Naresh Kumar Devulapally +2

Diffusion language models (DLMs) expose their denoising trajectories, offering a natural handle for inference-time control; accordingly, an ideal hallucination mitigation framework…

cs.CL2026

Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective

Krishna Singh Rajput, Tejas Anvekar, Chitta Baral +1

Recent advances in multimodal question answering have primarily focused on combining heterogeneous modalities or fine-tuning multimodal large language models. While these approache…

cs.CV2025

NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

Pranshu Pandya, Vatsal Gupta, Agney S Talwarr +3

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spa…

cs.CL2026

TabRank: Chain-of-Thought Distillation for Table Re-Rankers

Adarsh Singh, Kushal Raj Bhandari, Jianxi Gao +2

The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely heavily on rerankers to refin…

astro-ph.GA2025

Mapping the Spatial Distribution of Fast Radio Bursts within their Host Galaxies

Alexa C. Gordon, Wen-fai Fong, Adam T. Deller +23

We present deep optical and near-infrared observations of the host galaxies of 34 fast radio bursts (FRBs) detected by the Commensal Real-time ASKAP Fast Transient (CRAFT) survey o…

astro-ph.HE2018

Detection of a glitch in the pulsar J1709-4429

Marcus E. Lower, Chris Flynn, Matthew Bailes +24

We report the detection of a glitch event in the pulsar J17094429 (also known as B170644) during regular monitoring observations with the Molonglo Observatory Synthesis Teles…

astro-ph.HE2024

The impact of the FREDDA dedispersion algorithm on estimations with FRBs

Jordan Hoffmann, Clancy W. James, Hao Qiu +12

Fast radio bursts (FRBs) are transient radio signals of extragalactic origins that are subjected to propagation effects such as dispersion and scattering. It follows then that thes…

cs.CL2025

SPORTSQL: An Interactive System for Real-Time Sports Reasoning and Visualization

Sebastian Martinez, Naman Ahuja, Fenil Bardoliya +2

We present a modular, interactive system, SPORTSQL, for natural language querying and visualization of dynamic sports data, with a focus on the English Premier League (EPL). The sy…

cs.CL2022

Realistic Data Augmentation Framework for Enhancing Tabular Reasoning

Dibyakanti Kumar, Vivek Gupta, Soumya Sharma +1

Existing approaches to constructing training data for Natural Language Inference (NLI) tasks, such as for semi-structured table reasoning, are either via crowdsourcing or fully aut…

cs.LG2026

Invariant Reasoning Directions in Latent Trajectories of Language Models

Arun Vignesh Malarkkan, Manan Roy Choudhury, Utkarsh Byahut +3

Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectories remains poorly understood. We show tha…

cs.AI2026

JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI

Mayank Vyas, Abhijit Chakraborty, Vivek Gupta

Recruiters and job seekers rely on search systems to navigate labor markets, making candidate matching engines critical for hiring outcomes. Most systems act as keyword filters, fa…

cs.CL2025

Map&Make: Schema Guided Text to Table Generation

Naman Ahuja, Fenil Bardoliya, Chitta Baral +1

Transforming dense, detailed, unstructured text into an interpretable and summarised table, also colloquially known as Text-to-Table generation, is an essential task for informatio…

cs.CL2025

LLM-Symbolic Integration for Robust Temporal Tabular Reasoning

Atharv Kulkarni, Kushagra Dixit, Vivek Srikumar +2

Temporal tabular question answering presents a significant challenge for Large Language Models (LLMs), requiring robust reasoning over structured data, which is a task where tradit…

cs.LG2017

Efficient Estimation of Generalization Error and Bias-Variance Components of Ensembles

Dhruv Mahajan, Vivek Gupta, S Sathiya Keerthi +3

For many applications, an ensemble of base classifiers is an effective solution. The tuning of its parameters(number of classes, amount of data on which each classifier is to be tr…

cs.LG2017

Leveraging Distributional Semantics for Multi-Label Learning

Rahul Wadbude, Vivek Gupta, Piyush Rai +3

We present a novel and scalable label embedding framework for large-scale multi-label learning a.k.a ExMLDS (Extreme Multi-Label Learning using Distributional Semantics). Our appro…

cs.LG2019

Equalizing Recourse across Groups

Vivek Gupta, Pegah Nokhiz, Chitradeep Dutta Roy +1

The rise in machine learning-assisted decision-making has led to concerns about the fairness of the decisions and techniques to mitigate problems of discrimination. If a negative d…

cs.AI2026

Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities

Manan Roy Choudhury, Adithya Chandramouli, Mannan Anand +1

The rapid integration of large language models (LLMs) into high-stakes legal work has exposed a critical gap: no benchmark exists to systematically stress-test their reliability ag…

cs.CL2017

SCDV : Sparse Composite Document Vectors using soft clustering over distributional representations

Dheeraj Mekala, Vivek Gupta, Bhargavi Paranjape +1

We present a feature vector formation technique for documents - Sparse Composite Document Vector (SCDV) - which overcomes several shortcomings of the current distributional paragra…

cs.CL2021

RETRONLU: Retrieval Augmented Task-Oriented Semantic Parsing

Vivek Gupta, Akshat Shrivastava, Adithya Sagar +2

While large pre-trained language models accumulate a lot of knowledge in their parameters, it has been demonstrated that augmenting it with non-parametric retrieval-based memory ha…

cs.CL2026

ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation

Jesus-German Ortiz-Barajas, Jonathan Tonglet, Vivek Gupta +1

Multimodal large language models (MLLMs) are increasingly used to automate chart generation from data tables, improving analysis and reporting efficiency while introducing new misu…

cs.CL2025

UNJOIN: Enhancing Multi-Table Text-to-SQL Generation via Schema Simplification

Poojah Ganesan, Rajat Aayush Jha, Dan Roth +1

Recent advances in large language models (LLMs) have greatly improved Text-to-SQL performance for single-table queries. But, it remains challenging in multi-table databases due to…

cs.LG2018

Bayes-optimal Hierarchical Classification over Asymmetric Tree-Distance Loss

Dheeraj Mekala, Vivek Gupta, Purushottam Kar +1

Hierarchical classification is supervised multi-class classification problem over the set of class labels organized according to a hierarchy. In this report, we study the work by R…

cs.HC2025

RELATE-Sim: Leveraging Turning Point Theory and LLM Agents to Predict and Understand Long-Term Relationship Dynamics through Interactive Narrative Simulations

Matthew Yue, Zhikun Xu, Vivek Gupta +3

Most dating technologies optimize for getting together, not staying together. We present RELATE-Sim, a theory-grounded simulator that models how couples behave at consequential tur…

cs.LG2025

AI-driven Inverse Design of Band-Tunable Mechanical Metastructures for Tailored Vibration Mitigation

Tanuj Gupta, Arun Kumar Sharma, Ankur Dwivedi +5

On-demand vibration mitigation in a mechanical system needs the suitable design of multiscale metastructures, involving complex unit cells. In this study, immersing in the world of…

cs.CL2025

RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables

Nikhil Abhyankar, Purvi Chaurasia, Sanchit Kabra +3

Existing tabular reasoning benchmarks mostly test models on small, uniform tables, underrepresenting the complexity of real-world data and giving an incomplete view of Large Langua…

cs.CL2026

Moneyball with LLMs: Analyzing Tabular Summarization in Sports Narratives

Ritam Upadhyay, Naman Ahuja, Rishabh Baral +2

Large language model (LLM) approaches to tabular summarization rely on extensive prompt engineering, decomposition pipelines, or entity-level intermediate representations to achiev…

cs.CL2025

MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering

Varun Srivastava, Fan Lei, Srija Mukhopadhyay +2

Recent advancements in multimodal large language models (MLLMs) have driven researchers to explore how well these models read data visualizations, e.g., bar charts, scatter plots.…

cs.CL2025

Evidence-Guided Schema Normalization for Temporal Tabular Reasoning

Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta +1

Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose a SQL-based approach that involves (1) generating a 3NF schema from Wiki…

astro-ph.HE2020

The UTMOST pulsar timing programme II: Timing noise across the pulsar population

Marcus E. Lower, Matthew Bailes, Ryan M. Shannon +15

While pulsars possess exceptional rotational stability, large scale timing studies have revealed at least two distinct types of irregularities in their rotation: red timing noise a…

cs.CL2021

TabPert: An Effective Platform for Tabular Perturbation

Nupur Jain, Vivek Gupta, Anshul Rai +1

To truly grasp reasoning ability, a Natural Language Inference model should be evaluated on counterfactual data. TabPert facilitates this by assisting in the generation of such cou…

cs.DB2025

H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables

Nikhil Abhyankar, Vivek Gupta, Dan Roth +1

Tabular reasoning involves interpreting natural language queries about tabular data, which presents a unique challenge of combining language understanding with structured data anal…

cs.CL2026

CRAFT: Training-Free Cascaded Retrieval for Tabular QA

Adarsh Singh, Kushal Raj Bhandari, Jianxi Gao +2

Open-Domain Table Question Answering (TQA) involves retrieving relevant tables from a large corpus to answer natural language queries. Traditional dense retrieval models such as DT…

cs.CL2022

Leveraging Data Recasting to Enhance Tabular Reasoning

Aashna Jena, Vivek Gupta, Manish Shrivastava +1

Creating challenging tabular inference data is essential for learning complex reasoning. Prior work has mostly relied on two data generation strategies. The first is human annotati…

cs.CL2024

Knowledge-Aware Reasoning over Multimodal Semi-structured Tables

Suyash Vardhan Mathur, Jainit Sushil Bafna, Kunal Kartik +5

Existing datasets for tabular question answering typically focus exclusively on text within cells. However, real-world data is inherently multimodal, often blending images such as…

cs.CV2025

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

Shikhhar Siingh, Abhinav Rawat, Chitta Baral +1

Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevan…

cs.CL2026

Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"

Dhruv Madhwal, Lyuxin David Zhang, Dan Roth +2

Large language models often struggle to recognize their knowledge limits in closed-book question answering, leading to confident hallucinations. While decomposed prompting is typic…

cs.AI2026

Synapse: Federated Tool Routing via Typed Compendium Artifacts

Abhijit Chakraborty, Yash Shah, Vivek Gupta

The unit of collaboration in federated learning determines what guarantees are even expressible. Flat units like weights, prompts, raw examples, carry no type signature on which pr…

cs.LG2020

DeepSumm -- Deep Code Summaries using Neural Transformer Architecture

Vivek Gupta

Source code summarizing is a task of writing short, natural language descriptions of source code behavior during run time. Such summaries are extremely useful for software developm…

astro-ph.HE2025

Ultra-Wideband Polarimetry of the April 2021 Profile Change Event in PSR J1713+0747

Rami F. Mandow, Andrew Zic, J. R. Dawson +15

The millisecond pulsar PSR J1713+0747 is a high-priority target for pulsar timing array experiments due to its long-term timing stability, and bright, narrow pulse profile. In Apri…

cs.CL2023

Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular Data

Mubashara Akhtar, Abhilash Shankarampeta, Vivek Gupta +3

Numbers are crucial for various real-world domains such as finance, economics, and science. Thus, understanding and reasoning with numbers are essential skills for language models…

cs.CL2026

FD-NL2SQL: Feedback-Driven Clinical NL2SQL that Improves with Use

Suparno Roy Chowdhury, Tejas Anvekar, Manan Roy Choudhury +5

Clinicians exploring oncology trial repositories often need ad-hoc, multi-constraint queries over biomarkers, endpoints, interventions, and time, yet writing SQL requires schema ex…

cs.CV2026

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding

Prasham Titiya, Jainil Trivedi, Chitta Baral +1

Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remain difficult for Multimodal Large Langu…

cs.CV2025

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

Tejas Anvekar, Fenil Bardoliya, Pavan K. Turaga +2

Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, mos…

cs.CV2025

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization

Muhammad Kashif Ali, Eun Woo Im, Dongjin Kim +4

Video stabilization remains a fundamental problem in computer vision, particularly pixel-level synthesis solutions for video stabilization, which synthesize full-frame outputs, add…

cs.SI2018

Characterizing the spread of exaggerated news content over social media

Jasabanta Patro, Sabyasachee Baruah, Vivek Gupta +3

In this paper, we consider a dataset comprising press releases about health research from different universities in the UK along with a corresponding set of news articles. First, w…

astro-ph.IM2020

Estimating fast transient detection pipeline efficiencies at UTMOST via real-time injection of mock FRBs

Vivek Gupta, Chris Flynn, Wael Farah +7

Dedicated surveys using different detection pipelines are being carried out at multiple observatories to find more Fast Radio Bursts (FRBs). Understanding the efficiency of detecti…

cs.CL2019

Improving Document Classification with Multi-Sense Embeddings

Vivek Gupta, Ankit Saw, Pegah Nokhiz +2

Efficient representation of text documents is an important building block in many NLP tasks. Research on long text categorization has shown that simple weighted averaging of word v…

cs.CL2025

No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning

Abhishek Rajgaria, Kushagra Dixit, Mayank Vyas +3

Temporal Table Reasoning is a critical challenge for Large Language Models (LLMs), requiring effective reasoning to extract relevant insights. Despite existence of multiple prompti…

physics.app-ph2025

Harnessing Piezoelectric Shear Actuators for Vibration Control in Sandwich Beams

Mark Baken, Vivek Gupta, Bas Jansen +1

Our study found that integrating shear piezo-transducers inside the beam offers a compact and efficient solution that enables localized damping control without compromising structu…

cs.CY2016

Assisting humans to achieve optimal sleep by changing ambient temperature

Vivek Gupta, Siddhant Mittal, Sandip Bhaumik +1

Environment plays a vital role in the sleep mechanism of a human. It has been shown from many studies that sleeping and waking environment, waking time and hours of sleep is of ver…

cs.AI2026

FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles

Arun Vignesh Malarkkan, Manan Roy Choudhury, Guangwei Zhang +4

Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains…

cs.LG2026

TIMEGATE: Sustainable Time-Boxed Promotion Gates for Continual ML Adaptation Under Resource Constraints

Abhijit Chakraborty, Suddhasvatta Das, Yash Shah +2

As machine learning(ML) systems evolve to continual adaptation, each re-training cycle uses compute, annotation, and energy. We introduce TIMEGATE, a policy layer managing adaptati…

cs.CV2026

MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps

Sharat Bhat, Harshita Khandelwal, Tushar Kataria +1

Maps are powerful carriers of structured and contextual knowledge, encompassing geography, demographics, infrastructure, and environmental patterns. Reasoning over such knowledge r…

cs.CL2022

IndicXNLI: Evaluating Multilingual Inference for Indian Languages

Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end,…

cs.CL2024

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

Shubhankar Singh, Purvi Chaurasia, Yerram Varun +4

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchm…

astro-ph.HE2022

The ultra narrow FRB20191107B, and the origins of FRB scattering

Vivek Gupta, Chris Flynn, Wael Farah +4

We report the detection of FRB20191107B with the UTMOST radio telescope at a dispersion measure (DM) of 714.9 . The burst consists of three components, the bright…

cs.CL2026

Improving Robustness of Tabular Retrieval via Representational Stability

Kushal Raj Bhandari, Adarsh Singh, Jianxi Gao +2

Transformer-based table retrieval systems flatten structured tables into token sequences, making retrieval sensitive to the choice of serialization even when table semantics remain…

cs.CV2026

Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models

Eun Woo Im, Muhammad Kashif Ali, Vivek Gupta

Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal capabilities, but they inherit the tendency to hallucinate from their underlying language models. While…

cs.CL2020

P-SIF: Document Embeddings Using Partition Averaging

Vivek Gupta, Ankit Saw, Pegah Nokhiz +3

Simple weighted averaging of word vectors often yields effective representations for sentences which outperform sophisticated seq2seq neural models in many tasks. While it is desir…

cs.CL2026

TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation

Vihang Pancholi, Jainit Bafna, Tejas Anvekar +2

Evaluating tables qualitatively and quantitatively poses a significant challenge, as standard metrics often overlook subtle structural and content-level discrepancies. To address t…

cs.CL2023

TempTabQA: Temporal Question Answering for Semi-Structured Tables

Vivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora +4

Semi-structured data, such as Infobox tables, often include temporal information about entities, either implicitly or explicitly. Can current NLP systems reason about such informat…

astro-ph.HE2025

The discovery of a 41s radio pulsar PSR J0311+1402 with ASKAP

Yuanming Wang, Pavan Uttarkar, Ryan Shannon +16

The emerging population of long-period radio transients (LPTs) show both similarities and differences with normal pulsars. A key difference is that their radio emission is too brig…

astro-ph.IM2024

Efficient Summation of Arbitrary Masks -- ESAM

Vivek Gupta, Keith Bannister, Chris Flynn +1

Searches for impulsive, astrophysical transients are often highly computationally demanding. A notable example is the dedispersion process required for performing blind searches fo…

cs.CV2026

DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams

Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Gaurav Najpande +4

Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Rec…

cs.CL2026

Integrity Shield A System for Ethical AI Use & Authorship Transparency in Assessments

Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi +3

Large Language Models (LLMs) can now solve entire exams directly from uploaded PDF assessments, raising urgent concerns about academic integrity and the reliability of grades and c…

cs.CL2026

TabReX : Tabular Referenceless eXplainable Evaluation

Tejas Anvekar, Junha Park, Aparna Garimella +1

Evaluating the quality of tables generated by large language models (LLMs) remains an open challenge: existing metrics either flatten tables into text, ignoring structure, or rely…

cs.CL2026

SCOPE:Planning for Hybrid Querying over Clinical Trial Data

Suparno Roy Chowdhury, Manan Roy Choudhury, Tejas Anvekar +5

We study clinical trial table reasoning, where answers are not directly stored in visible cells but must be reasoned from semantic understanding through normalization, classificati…

cs.CL2021

Unsupervised Contextualized Document Representation

Ankur Gupta, Vivek Gupta

Several NLP tasks need the effective representation of text documents. Arora et. al., 2017 demonstrate that simple weighted averaging of word vectors frequently outperforms neural…

cs.CL2026

TraceBack: Multi-Agent Decomposition for Fine-Grained Table Attribution

Tejas Anvekar, Junha Park, Rajat Jha +4

Question answering (QA) over structured tables requires not only accurate answers but also transparency about which cells support them. Existing table QA systems rarely provide fin…

cs.AI2016

Product Classification in E-Commerce using Distributional Semantics

Vivek Gupta, Harish Karnick, Ashendra Bansal +1

Product classification is the task of automatically predicting a taxonomy path for a product in a predefined taxonomy hierarchy given a textual product description or title. For ef…

cs.AI2019

A Logic-Driven Framework for Consistency of Neural Models

Tao Li, Vivek Gupta, Maitrey Mehta +1

While neural models show remarkable accuracy on individual predictions, their internal beliefs can be inconsistent across examples. In this paper, we formalize such inconsistency a…

cs.CL2026

ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution

Yahia Alqurnawi, Preetom Biswas, Anmol Rao +3

Multimodal Large Language Models (mLLMs) are often used to answer questions in structured data such as tables in Markdown, JSON, and images. While these models can often give corre…

cs.CL2026

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA

Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri +4

Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer, rather than onl…

cs.CL2017

Text Summarization using Abstract Meaning Representation

Shibhansh Dohare, Harish Karnick, Vivek Gupta

With an ever increasing size of text present on the Internet, automatic summary generation remains an important problem for natural language understanding. In this work we explore…

cs.LG2025

Distribution Shift Aware Neural Tabular Learning

Wangyang Ying, Nanxu Gong, Dongjie Wang +5

Tabular learning transforms raw features into optimized spaces for downstream tasks, but its effectiveness deteriorates under distribution shifts between training and testing data.…

cs.CL2025

Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering

Pragya Srivastava, Manuj Malik, Vivek Gupta +2

Large Language Models (LLMs), excel in natural language understanding, but their capability for complex mathematical reasoning with an amalgamation of structured tables and unstruc…

cs.CL2024

Enhancing Temporal Understanding in LLMs for Semi-structured Tables

Irwin Deng, Kushagra Dixit, Vivek Gupta +1

Temporal reasoning over tabular data presents substantial challenges for large language models (LLMs), as evidenced by recent research. In this study, we conduct a comprehensive an…

cs.CL2022

Enhancing Tabular Reasoning with Pattern Exploiting Training

Abhilash Reddy Shankarampeta, Vivek Gupta, Shuo Zhang

Recent methods based on pre-trained language models have exhibited superior performance over tabular tasks (e.g., tabular NLI), despite showing inherent problems such as not using…

cs.CL2026

DeALOG: Decentralized Multi-Agents Log-Mediated Reasoning Framework

Abhijit Chakraborty, Ashish Raj Shekhar, Shiven Agarwal +1

Complex question answering across text, tables and images requires integrating diverse information sources. A framework supporting specialized processing with coordination and inte…

cs.CL2025

Is Architectural Complexity Overrated? Competitive and Interpretable Knowledge Graph Completion with RelatE

Abhijit Chakraborty, Chahana Dahal, Ashutosh Balasubramaniam +2

We revisit the efficacy of simple, real-valued embedding models for knowledge graph completion and introduce RelatE, an interpretable and modular method that efficiently integrates…

astro-ph.GA2021

Fast radio bursts as probes of feedback from active galactic nuclei

Adam J. Batten, Alan R. Duffy, Chris Flynn +3

Fast Radio Bursts (FRBs) are a promising tool for studying the low-density universe as their dispersion measures (DM) are extremely sensitive probes of electron column density. Act…

cs.RO2025

Hybrid Robotic Meta-gripper for Tomato Harvesting: Analysis of Auxetic Structures with Lattice Orientation Variations

Shahid Ansari, Vivek Gupta, Bishakh Bhattacharya

The agricultural sector is rapidly evolving to meet growing global food demands, yet tasks like fruit and vegetable handling remain labor-intensive, causing inefficiencies and post…

cs.CL2025

Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents

Manan Suri, Puneet Mathur, Nedim Lipka +4

Flowcharts are a critical tool for visualizing decision-making processes. However, their non-linear structure and complex visual-textual relationships make it challenging to interp…

cs.CL2022

Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning

Vivek Gupta, Riyaz A. Bhat, Atreya Ghosal +3

Neural models command state-of-the-art performance across NLP tasks, including ones involving "reasoning". Models claiming to reason about the evidence presented to them should att…

cs.CL2022

SLATE: A Sequence Labeling Approach for Task Extraction from Free-form Inked Content

Apurva Gandhi, Ryan Serrao, Biyi Fang +8

We present SLATE, a sequence labeling approach for extracting tasks from free-form content such as digitally handwritten (or "inked") notes on a virtual whiteboard. Our approach al…

cs.CL2025

Leveraging LLM For Synchronizing Information Across Multilingual Tables

Siddharth Khincha, Tushar Kataria, Ankita Anand +2

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedi…

physics.app-ph2026

Digitally Controlled Mechatronic Metamaterials for Actively Induced Targeted Bandgaps

Vivek Gupta, Aditya Natu, S. Hassan HosseinNia

This paper presents an experimental framework for inducing and tuning vibration bandgaps in digitally controlled mechatronic metamaterials. A slender-beam structure instrumented wi…

cs.CL2024

ChartCheck: Explainable Fact-Checking over Real-World Chart Images

Mubashara Akhtar, Nikesh Subedi, Vivek Gupta +3

Whilst fact verification has attracted substantial interest in the natural language processing community, verifying misinforming statements against data visualizations such as char…