6 papers
Position: Evaluation Scores Are Perishable Knowledge Claims
Sankalp Gilda, Shlok Gilda
Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results…
tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series
Sankalp Gilda
Finance, sensing, and demand streams violate the exchangeability that IID conformal prediction and the IID bootstrap assume, and existing libraries implement either a general resam…
Structured Abductive-Deductive-Inductive Reasoning for LLMs via Algebraic Invariants
Sankalp Gilda, Shlok Gilda
Large language models exhibit systematic limitations in structured logical reasoning: they conflate hypothesis generation with verification, cannot distinguish conjecture from vali…
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
Sankalp Gilda, Shlok Gilda
This position paper argues that AI-assisted software engineering requires explicit mechanisms for tracking the epistemic status and temporal validity of architectural decisions. LL…
deep-REMAP: Probabilistic Parameterization of Stellar Spectra Using Regularized Multi-Task Learning
Sankalp Gilda
In the era of exploding survey volumes, traditional methods of spectroscopic analysis are being pushed to their limits. In response, we develop deep-REMAP, a novel deep learning fr…
tsbootstrap: Enhancing Time Series Analysis with Advanced Bootstrapping Techniques
Sankalp Gilda, Benedikt Heidrich, Franz Kiraly
In time series analysis, traditional bootstrapping methods often fall short due to their assumption of data independence, a condition rarely met in time-dependent data. This paper…