New AI Tool Predicts Liquid Chromatography Retention Times
- The newly developed “2-step” software tool predicts the elution order and retention times of small metabolites in liquid chromatography without requiring system-specific retraining.
- Developed by a team led by bioinformatician Prof. Dr. Sebastian Böcker and co-first author Fleming Kretschmer, the method utilizes a machine learning model to calculate a Retention Order Index.
- The method was published in the journal Nature Methods, offering broad applications across drug discovery, natural product research, and environmental analytics.
What problem does liquid chromatography face in identifying small molecules?
Analyzing complex biological samples like blood, cell material, or bacterial cultures requires identifying myriad small molecules known as metabolites, which encompass metabolic products, natural substances, toxins, degradation products, and pharmaceuticals. To separate these substances, laboratories frequently use liquid chromatography, where a mixture passes through a separation column and individual molecules exit at varying times. This specific “retention time” provides vital clues about the substance inside a sample. However, predicting when exactly a specific molecule will exit the column has remained a challenge for decades.
Retention times depend heavily on experimental parameters such as the column type, solvent, gradient, pH value, temperature, and minor technical alterations. Even replacing a piece of tubing with a slightly longer alternative can significantly shift measured times. According to Prof. Dr. Sebastian Böcker, previous predictive models frequently required calibration using data from the exact measurement system where predictions would later occur. This demanded that researchers measure numerous standard substances beforehand, creating an expensive and time-consuming hurdle for routine laboratory work.
How does the new “2-step” software tool work?
The newly published approach circumvents system-specific training limitations by focusing on the most common reversed-phase liquid chromatography mode. Operating in two distinct phases, the method first uses a machine learning model to compute a Retention Order Index for any given molecule. This index establishes the molecule’s relative position within the anticipated elution sequence rather than assigning a rigid time value. In the second phase, this order index converts into concrete retention times using a small set of known reference points.
Fleming Kretschmer, co-first author of the study who worked on the procedure during his doctoral research, emphasized that the method surpasses existing techniques that require intensive tuning on the target system. The workflow allows precise out-of-the-box predictions even when applied to entirely new analytical systems and previously unknown molecules.
Where will researchers apply this predictive method?
The software addresses analytical bottlenecks across natural product research, environmental analytics, and pharmaceutical development. In natural product research, scientists analyze extracts from bacteria, fungi, and plants to isolate potential new antibiotics, cancer therapeutics, or other clinical agents. When examining complex matrices to determine whether a measured retention time matches a suspected chemical structure, the software aids in verifying molecular identities without exhaustive physical calibration runs.
Where is the software currently available for laboratories?
The research team published their findings in the journal Nature Methods. The “2-step” tool is available as both a software package and a web application via the GitHub repository hosted by the Böcker lab. The developers intend for the method to integrate directly into existing analytical software suites, automating retention time calculations for laboratories utilizing liquid chromatography coupled with mass spectrometry.