The human physique and its intestine microbiome produce hundreds of small molecules that form how the physique capabilities—influencing immunity, metabolism, and extra. Figuring out what these molecules truly are has been one of many biomedical sciences’ most persistent bottlenecks. Greater than 80% of compounds detected in a typical organic pattern can’t be matched to any identified construction utilizing present strategies.
Researchers on the Boyce Thompson Institute (BTI) and Cornell College have developed a instrument that begins to alter that. AIMe, brief for AI Molecule Explorer, makes use of a type of synthetic intelligence known as neuro-symbolic AI to foretell, arrange, and search the mass spectra of the whole lot of identified small natural molecules, greater than 100 million compounds—successfully constructing an unlimited searchable map of chemical house that may speed up speculation era and compound identification.
The work is a collaboration between Frank Schroeder, Professor at BTI and in Cornell’s Division of Chemistry and Chemical Biology, and Carla Gomes, Professor of Computing and Info Science and director of Cornell’s AI for Science Institute.
A distinct method to an outdated downside
Mass spectrometry is the workhorse for small molecule identification utilized in a variety of functions from toxicology to meals evaluation. When a compound is analyzed, the instrument fragments it and data the lots of the ensuing items. That sample of fragments, the tandem mass spectrum (or MS2 spectrum), capabilities as a molecular fingerprint. To determine an unknown compound, researchers examine its spectrum towards a reference library of spectra from identified compounds or develop hypotheses as to what the construction of a compound could also be primarily based on handbook evaluation of the fragmentation sample.
The issue is that experimental reference libraries stay sparse, whereas professional, one-by-one interpretation is labor-intensive and sluggish. Collectively, out there libraries cowl fewer than 1% of identified compounds, and resolving the construction of a single unknown can take days to months of iterative evaluation and experimental validation. Because of this, spectra with out shut library matches normally stay unannotated.
AIMe takes a special method. Reasonably than ready for experimental spectra to build up in libraries, it predicts spectra computationally—then organizes these predictions right into a searchable useful resource known as MS2KOSMOS. AIMe generated greater than 800 million predicted spectra protecting basically all identified small natural molecules in PubChem, the most important publicly out there chemical database. That represents roughly a thousandfold enlargement of searchable chemical house relative to present experimental libraries.
“On the core of AIMe is DeepMS2Reasoner, a mannequin that simulates how molecules fragment inside a mass spectrometer,” defined Gomes. “It builds fragmentation pathways step-by-step, utilizing symbolic chemical guidelines to enumerate bodily believable fragmentation steps and a neural community to assign likelihoods to every step. The result’s a predicted spectrum and an annotated map of how a molecule got here aside—a characteristic that makes AIMe’s outputs interpretable in chemical phrases, not simply computationally helpful.”
From mouse intestine to human biology
To show what AIMe can do in follow, the researchers utilized it to a comparative metabolomics dataset from mice. The experiment in contrast germ-free mice, animals raised with none intestine microbiota, towards mice with a standard complement of intestine micro organism. A number of thousand chemical options differed between the 2 teams, and a lot of the plentiful ones couldn’t be recognized utilizing commonplace strategies.
The workforce used AIMe to question MS2KOSMOS with spectra from the 111 most plentiful unidentified microbiota-dependent compounds. For roughly a 3rd, AIMe retrieved shut predicted spectral matches and associated structural candidates, offering chemically interpretable leads for follow-up. For the remainder, AIMe mapped the unknown spectra to molecular neighborhoods—units of structurally associated compounds whose shared fragmentation patterns might inform hypotheses about what the unknowns is likely to be.
“Two compounds particularly grew to become a case research in what AI-guided construction elucidation can accomplish,” mentioned Schroeder. “Each produced spectra that urged they had been polyamine derivatives, a well-studied class of molecules, however the fragment patterns did not match something beforehand described. Utilizing AIMe’s output as a information, our workforce assembled candidate constructions combinatorially, predicted spectra for every candidate, and used the comparability to slim the sphere.”
For the primary compound, the very best candidate was a linear putrescine by-product, which was then simply verified by synthesizing an genuine commonplace. For the second, the anticipated spectra of potential candidates persistently failed to elucidate two distinguished peaks—till the workforce expanded the search to incorporate cyclized variants.
Synthesis confirmed what the predictions urged. The second compound turned out to be a structurally uncommon macrocyclic polyamine—a ring-shaped variant in contrast to any beforehand reported from mouse or human biology. Searches of a giant public mass spectrometry database subsequently discovered the identical compound in samples of human origin, detected in 57 of 99 human fecal samples examined.
Polyamines occupy an vital place in biology. They sit on the intersection of food plan, the intestine microbiome, and immune operate—and these findings counsel that the catalog of microbiota-dependent polyamines is way much less full than scientists assumed.
Annotation at scale
The researchers additionally examined AIMe at repository scale, making use of it to greater than 7 million spectral clusters from the World Pure Merchandise Social Molecular Networking database, one of many largest publicly out there repositories of mass spectrometry information. Prior annotation efforts had annotated roughly 416,000 of these clusters. AIMe returned putative annotations for roughly 2.69 million, utilizing the identical similarity threshold utilized in that earlier effort.
AIMe is accessible at https://www.cs.cornell.edu/gomes/udiscoverit/kosmos/ and the supply code can be made publicly out there upon publication of the research.
Reference: Acikalin UU, Feng D, Ferber AM, et al. Charting the small-molecule universe from mass spectra with neuro-symbolic AI. bioRxiv. Preprint posted on-line August 6, 2026. doi: 10.64898/2026.08.05.743095
This text is predicated on analysis findings which are but to be peer-reviewed. Outcomes are due to this fact thought to be preliminary and must be interpreted as such. For additional data, please contact the cited supply.
This text has been republished from supplies linked above. Observe: materials might have been edited for size and content material. For additional data, please contact the cited supply.