What Are Best Practices for AI Use in Early-Stage Scientific Research?

October 5, 2026

Before a researcher runs an experiment or collects data, they face decisions that can shape months or years of work: What do we already know? Where are the gaps? Which questions are worth pursuing? AI tools promise to make that groundwork faster, but the quality of their answers can influence everything that follows.

When Sanjana Gautam was a Bullard Research Fellow at the iSchool from 2024-2025, working with the NSF-Simons AI Institute for Cosmic Origins – better known as CosmicAI – campus was abuzz with conversations about AI’s impact on scientific research. "There were a lot of changes in how we navigated research projects," Gautam says. "I often found myself in rooms where researchers were discussing how they were changing their pipelines, what they thought AI could do well and what they thought it could not." 

Sanjana Gautam, 2024-2025 iSchool Bullard Research Fellow
Sanjana Gautam, 2024-2025 Bullard Research Fellow

For example, when a scientist becomes curious about a new question, an LLM can help by rapidly summarizing existing knowledge on the topic. But unlike scientists, who value precise and limited claims, LLMs sometimes overstate findings or even get facts wrong. LLM errors can send researchers down unproductive paths or lead them to repeat existing work; risks that graduate students and postdocs still learning a field may struggle to spot.

Guatam’s thoughtful response to this conundrum is a recent paper, “How Researchers Navigate Accountability, Transparency, and Trust when Using AI Tools in Early-Stage Research: A Think-Aloud Study,” co-authored with CosmicAI co-director and iSchool professor Matthew Lease, alongside iSchool graduate students Houjiang Liu and Yujin Choi. The findings offer practical lessons for researchers weighing where AI can save them time, which claims need checking and when to rely on their own judgment before committing to a research path.

Watching Scientists Think in Real Time

Rather than survey researchers about their AI habits after the fact, Gautam’s team asked participants to narrate their thinking as they worked — a "think-aloud" approach borrowed from usability research and clinical decision-making studies. Gautam says the method was chosen to avoid distortions that come with hindsight. "The thing with survey design is it's very post hoc and surface-level," Gautam says. "A think-aloud study allows for less observer effect.”

Participants worked with two commercially available tools, Research Rabbit and Elicit AI, chosen to represent two different ways AI might assist in early research: one for exploring citation networks, the other for generating AI-written summaries and syntheses of the literature. Because early-stage research looks broadly similar across disciplines, Gautam’s 15 participants hailed from various branches of academia and included one researcher working in industry.

The Confidence Problem

The single biggest source of complaints, Gautam says, was the tendency of LLMs to overstate claims. The paper frames this as a problem of accountability: a tool that can't reliably signal its own uncertainty makes it hard for researchers to know which outputs should be flagged for further investigation. "People were put off by how assertive AI was and how confidently wrong it could be," she said. "The more confident the tool sounded, the more researchers were suspicious of its response."

Closely related is what participants described as a "black box" quality to how these tools retrieve and construct information. More than once, experts saw AI tools inexplicably recommend less relevant papers that didn't match their own sense of the field, which made them doubt they could trust the tools for questions outside their expertise. "Opacity formed the basis of mistrust,” Gautam says. 

How Researchers Compensate

Faced with tools they couldn't fully trust, participants attempted to build workarounds that could allow them to take advantage of AI’s positive contributions. Many leaned on what the paper calls "social credibility heuristics," applying their own experience as an additional filter to AI literature reviews. If a recommended paper came from an author they recognized and respected, they took the suggestion; if not, they met it with healthy skepticism.

Participants also modulated how much they leaned on AI depending on the stakes involved, deferring to LLMs to brainstorm and summarize lists of papers to review, but eventually reading the key papers themselves. "In a higher-stakes scenario, they wanted to retain more agency," Gautam says, while for lower-stakes tasks, "they were okay to defer that work to AI."

An Unsettled Question for the Next Generation

Gautam identified other concerns that may prove harder to navigate. Senior researchers in the study worried about younger scholars who may never learn to evaluate relevant papers the old-fashioned way, and who haven’t been around long enough to filter AI-curated reading lists by name recognition.

"Anybody starting now, if they use AI, might not be able to leverage similar workarounds" built from years of experience in a field, she says. That, in turn, raises the odds of a young researcher building on a hallucinated, irrelevant or less trustworthy source. The question of what universities might do to address this challenge fell outside the scope of the study, but Gautam calls it "a problem worth investing in.”

Designing Tools Researchers Can Actually Trust

The paper also points toward design fixes for AI tools. First, Gautam suggests a "phased transition": rather than asking researchers to abandon indexed tools like Google Scholar and trust in a black box intelligence, LLM outputs should preserve cues that researchers can use to reverse-engineer findings, such as search terms that explain why a result surfaced. "Not everything needs to change at once," Gautam said. 

She also suggests a readout that tells the user how sure an LLM is about a given claim or correlation. This is a fix that likely must be addressed at the model level, Gautam notes. An LLM model that informs users of its varying confidence levels in its own outputs could help build trust in AI far beyond the world of scientific research.

A Universe of Possibilities

Matthew Lease, iSchool professor and co-director of CosmicAI
Matthew Lease, iSchool professor and co-director of CosmicAI

Since completing her postdoc at UT, Gautam has taken a role at Microsoft evaluating AI agents for trust and accountability, work she describes as a natural extension of the questions raised in this study. Meanwhile, her paper, accepted to the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT) in Montreal this June, lands at a moment when nearly every research lab is renegotiating its relationship with AI tools, often without guidance on best practices.

For example, at CosmicAI, co-director Lease also co-leads the "Explorable Universe" research pillar, focused on building trustworthy generative AI for astronomy. He describes the current moment in AI as "exciting and anxiety-provoking.” He sees Gautam’s insights as crucial for scientists hoping to harness AI’s vast potential to assist in major breakthroughs. 

"When the confident tone of AI outputs misrepresents epistemic uncertainty, this makes it more difficult for researchers to identify which outputs require the greatest scrutiny," he says. “Trust in AI is fragile, context-dependent, and easily eroded."

 

Learn more about researchers at CosmicAI using cutting-edge technologies to advance frontiers of human knowledge.

 

 

Share this content