Monday, August 3, 2026 Login
Breaking
Trent Rockets withstand Annabel Sutherland assault to secure tight win Webb telescope finds signs of ancient disaster for Neptune's moons – Reuters Liver Health: Which Fruits Should You Eat to Cleanse Your Liver? Find Out – Lokshahi English News Congress keeps Datia seat in Madhya Pradesh, defeats BJP rival by 6,016 votes Two tankers with Saudi oil exit Red Sea over weekend, data shows – Reuters
Science

For AI to Drive Science Discoveries, Highly Reproducible Data Is Key

BYLINE: Catherine Zandonella

Read this story in the SLAC News Center

Key takeaways

  • AI modeling with experimental data could predict catalyst performance in converting carbon dioxide into fuels, researchers at SLAC propose.
  • Variability in results from four independent labs during catalyst testing highlighted the need for extra steps to standardize methods to achieve reproducible outcomes.
  • The study provides essential guidance for scientists to ensure reliable data for AI models by carefully considering experimental methods and conditions.

Newswise — To turn abundant carbon dioxide into valuable fuel, we need a fast and efficient way of determining which catalysts work best over the longest time. AI models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data you put into them.

By convening four laboratories from across the nation to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, researchers at SLAC National Accelerator Laboratory have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results in Nature Catalysis.

“Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes,” said Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara, and first author on the study.

Speeding up catalyst development – with AI

With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.

In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time periods (days), but catalyst deactivation occurs over time (months to years) due to buildup of impurities and repeated exposure to high temperatures.

AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst.

Squaring the data from round-robin experiments

To the researchers’ surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem. 

Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.

“It was a bit of an eye-opener,” said SLAC staff scientist Adam Hoffman, senior author of the study. “This experience shines light on the practical challenges of including real-world data into machine learning models.”

Painstakingly the teams evaluated their methods. Through rigorous testing they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred. 

With further standardization across the four labs – which in addition to SLAC included groups at Pennsylvania State University, Stanford University and University of California, Santa Barbara – the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols and experimental conditions.

Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.

“We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models,” Hoffman said.

This work was supported in part by the U.S. Department of Energy (DOE) Office of Science. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. The SLAC portion of the research took place at the Stanford Synchrotron Radiation Lightsource (SSRL), a DOE Office of Science user facility.

About SLAC

SLAC National Accelerator Laboratory explores how the universe works at the biggest, smallest and fastest scales and invents powerful tools used by researchers around the globe. As world leaders in ultrafast science and bold explorers of the physics of the universe, we forge new ground in understanding our origins and building a healthier and more sustainable future. Our discovery and innovation help develop new materials and chemical processes and open unprecedented views of the cosmos and life’s most delicate machinery. Building on more than 60 years of visionary research, we help shape the future by advancing areas such as quantum technology, scientific computing and the development of next-generation accelerators.

SLAC is operated by Stanford University for the U.S. Department of Energy’s Office of Science. The Office of Science is the single largest supporter of basic research in the physical sciences in the United States and is working to address some of the most pressing challenges of our time.



Source link

Related Stories

Leave a Comment

Your email address will not be published. Required fields are marked *