James Howison Wins Grant to Make Science Software Visible in the Scholarly Record

August 27, 2026

Every modern scientific discovery relies on good code. Yet if you look at the official record of science — the papers and citation-count databases that funders and universities use to measure impact — that software is almost invisible. 

James Howison, professor at the iSchool, has been trying to fix that oversight and make research software easier to track in the scholarly graph. Doing so could help ensure that the people who build and maintain essential scientific tools receive the recognition, funding and career opportunities their work deserves. 

James Howison
Dr. James Howison

That project takes a major step forward with the announcement of a two-year grant from Schmidt Sciences, which brings over half a million dollars to UT in the first year alone.

Unsung Heroes of Science

Howison's interest in the problem goes back more than a decade, to a research project that traced the software behind published papers back to the people who wrote it. Speaking to these coders, he found that the invisibility of their work had limited their professional advancement. 

“We recognize the authors of research papers, but not the people who write the code that makes much of that research possible,” Howison says. In turn, this lack of recognition harms the scientific software ecosystem, making coders less likely to improve and maintain software.

The problem is more urgent in the age of AI. AI coding tools draw on the same foundation of under-recognized code, which becomes harder than ever to attribute. "AI can do what it can do because it's building on software packages," Howison says. "Without the packages, it would have to write a lot more code, work deep down into the stack, and that is a much more error-prone situation." He adds that AI coding also threatens recognition of the value of coordinating and curating packages.

Iterations at the iSchool 

Beginning seven years ago, SoftCite, Howison’s previous software attribution project at the iSchool, trained a model to recognize when software was mentioned in scientific text, eventually publishing extractions from over 26 million PDFs. Howison accomplished this with support from the Sloan Foundation, iSchool Ph.D. Johanna Cohoon (now a UX researcher at Berkeley Lab) and roughly 30 students who read and annotated 5,000 scientific papers. SoftCite has been a first step in giving coders their due, but it had room for improvement in areas like disambiguation. For example, he says, "There are about five different pieces of software called STAR."

The Schmidt Sciences-funded project takes a more comprehensive approach. Working with Andrew Nesbitt at ecosyste.ms, Howison will begin by identifying which of the hundreds of millions of repositories on GitHub are likely scientific software. From there, the model will match these candidates against mentions in academic literature. By surveying both software packages and mentions, it becomes easier to map connections between the two data sets.  

A Database Open for All

The grant pairs Howison and his team with OpenAlex, an open-access answer to Google Scholar, which tracks and indexes academic work. Another significant aspect of the project involves expanding OpenAlex's data model to represent software properly — accounting for the frequent versioning and complex dependency chains that make software different from papers. 

"The results of what we do will become part of the OpenAlex graph, which is available to all, both via their APIs and through download dumps," Howison says.

What Success Looks Like

Asked about the impact he hopes the grant will have, Howison points to two outcomes. "It’ll be easier for the people who are shepherding software to make their case for impact," he says. He sees this as overdue justice for research software engineers whose careers have long been penalized by a system that couldn't see their work. 

The second outcome has a broader benefit. As AI systems increasingly assist scientists, better data about software will lead to more helpful context. "People within a field of science will get better recommendations about how to do their analyses, because agents will have more structured data to work with," Howison said. "This means better AI-driven science."

To learn more about Howison’s work on software development and collaboration, visit his research profile.

 

Share this content