works on

From the 1 of 25 linked papers with an AI index.

collaborators

25 papers

cs.AI2026

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

Swati Rajwal, Sanjay Das, Tirthankar Ghosal

Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled sci…

cs.LG2026

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

Xuehang Guo, Pengyuan Li, Tom Hope +3

As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fun…

cs.CL2026

Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

Shashwat Sourav, Subhadeep Pal, Markus J. Buehler +4

AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically meaningful mechanism. We present a graph-to-a…

cs.DC2026

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage +25

The paper updates a community roadmap for autonomous scientific laboratories, emphasizing trust, verification, reproducibility, safety, security, and governance as central challeng…

cs.AI2026

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Subhadeep Pal, Shashwat Sourav, Tirthankar Ghosal +1

Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models…

cs.AI2026

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

Sanjay Das, Ran Elgedawy, Ethan Seefried +2

Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identification. While large language m…