Decoding the Glycocode: AI-Powered Glycoproteomics & Cancer Research
- Photo: Concentrating on Chromatography: Decoding the Glycocode: AI-Powered Glycoproteomics & Cancer Research
- Video: Concentrating on Chromatography: Decoding the Glycocode: AI-Powered Glycoproteomics & Cancer Research
In this episode of Concentrating on Chromatography, host David Oliva is joined for the
first time by co-host Candice (Candi) Gokey — PhD candidate at UC San Diego / San Diego State University and expert in LC-MS/MS method development for untargeted metabolomics to interview Jacob Russell, a third-year PhD student in the Riley Research Group at the University of Washington (Seattle).
Jacob's research sits at the frontier of mass spectrometry-based intact glycoproteomics,
where the goal is to keep glycans attached to peptides during analysis — preserving the
biological information that is lost when glycans are enzymatically removed. His lab, led
by Prof. Nick Riley, develops new methods to tackle every stage of glycoproteomic
analysis, from sample preparation and enrichment through data acquisition and
bioinformatics.
WHAT WE COVER:
- What is the glycocalyx, and why does glycoproteomics matter for cancer, immunity,
and protein biology? - N-linked vs. O-linked glycosylation: why they require completely different
fragmentation strategies (HCD vs. ETD/EThcD) - Autonomous Dissociation-type Selection (ADS): how real-time library searching (RTLS)
of oxonium ion ratios (m/z 138 vs. 144) enables on-the-fly selection of the right
fragmentation method for N- vs. O-glycopeptides — published in J. Proteome Research
(Sutherland, Veth, Russell et al., 2024) - How XGBoost machine learning outperforms RTLS by classifying ~50 oxonium ions
simultaneously, recovering Core 2 O-glycopeptides that the simpler ratio-based
method misses - Casanovo Foundation: a transformer-based foundation model for tandem mass spectra
and how deep learning is reshaping glycopeptide classification
(Sanders et al., arXiv 2025) - LacNAc-ase enabled glycoproteomics: using endo-β-galactosidase (EBG) to "trim"
poly-LacNAc chains on N-glycans — uncovering previously undetectable glycoforms
(ASMS 2026 presentation) - Cutaneous vs. uveal melanoma: comparing poly-LacNAcylated glycoproteins (including
galectin-3 binding partners like CD63, LAMP-2, and basigin) between cell lines —
and why uveal melanoma has a ~50% distant metastasis rate vs. ~5% for cutaneous - The future of intelligent data acquisition and real-time mass spectrometry
- Candi's parallel world of untargeted metabolomics, coral-algae chemical communication,
and GNPS/MassQL — and the surprising overlap with glycoproteomics
Video Transcription
Why glycoproteomics is so challenging
Glycans play fundamental roles throughout biology. They are involved in immune processes, cancer, protein folding, protein stability, and many other cellular functions. Every living cell studied to date carries a glycan-rich surface layer, making glycosylation an important component of biological regulation and disease.
Studying glycoproteins, however, is considerably more complicated than conventional proteomics. One traditional way to simplify an experiment is to enzymatically remove glycans from proteins before analysis. While this makes the analytical problem easier, it also removes information about which amino acid carried a particular glycan.
The Riley laboratory instead works extensively with intact glycoproteomics, retaining the glycan on the peptide so that both the peptide sequence and glycosylation information can be investigated together. This introduces challenges at virtually every stage of the workflow, from enrichment and sample preparation through mass spectrometric acquisition to bioinformatic interpretation.
Russell emphasized that glycoproteomics still trails conventional proteomics in areas such as identification depth and standardized workflows. At the same time, these limitations provide opportunities for new analytical methods, including improved enrichment, acquisition strategies, instrumentation, and post-acquisition data processing.
N- and O-glycopeptides require different analytical strategies
One of the central complications is that N-linked and O-linked glycosylation behave differently.
N-glycosylation occurs at a relatively well-defined amino-acid sequence motif involving an asparagine residue. This predictable localization means that collision-based fragmentation can often provide sufficient information for identifying N-glycopeptides.
O-glycosylation presents a more difficult problem. O-glycans can occur at serine or threonine residues without an equivalent conserved sequence motif, and multiple O-glycans may be present on a single peptide. With conventional collision-based fragmentation, the glycan can be lost before sufficient peptide backbone information is generated, making it difficult to assign the modification to a specific amino acid.
To preserve localization information, researchers can use electron-based fragmentation such as electron transfer dissociation (ETD). ETD preferentially fragments the peptide backbone while leaving the glycan attached. The Riley group also employs supplemental collisional activation, combining the benefits of electron- and collision-based fragmentation. This can help separate non-covalently associated fragments while also producing glycan-derived fragments known as oxonium ions, which provide additional information about glycan composition.
The difficulty is therefore not simply detecting a glycopeptide. The instrument must ideally determine what kind of glycopeptide it is and select the most informative fragmentation strategy accordingly.
Making fragmentation decisions in real time
Earlier approaches to simultaneous N- and O-glycopeptide analysis often relied on product-dependent triggering. A collision-based MS/MS spectrum would first be acquired and examined for oxonium ions. If these characteristic glycan fragments were detected, the instrument would trigger an additional electron-based fragmentation event.
The limitation, according to Russell, is that this approach essentially asks only whether a precursor is a glycopeptide. It does not determine whether the spectrum is more likely to originate from an N- or O-glycopeptide. As a result, slower electron-based scans may be triggered unnecessarily, consuming valuable acquisition time during an LC-MS experiment.
To address this problem, the researchers developed an approach based on autonomous dissociation type selection and real-time spectral evaluation.
The strategy initially uses real-time library searching to compare an acquired spectrum with reference patterns characteristic of N- and O-glycopeptides. Particular attention is paid to two oxonium-ion signals at m/z 138 and 144. Their relative intensities provide information that can help distinguish the two glycopeptide classes.
If the acquired spectrum resembles an O-glycopeptide, the instrument can selectively trigger the more time-consuming electron-based fragmentation only where it is most useful. This makes more efficient use of instrument time and increases the amount of relevant information collected during a chromatographic run.
From two diagnostic ions to machine learning
Although the real-time library strategy improved acquisition efficiency, it still had limitations.
Some O-glycans contain structural features that can make their oxonium-ion pattern resemble that of N-glycans. A classification based primarily on the 138/144 ratio can therefore misclassify certain O-glycopeptides and fail to trigger the appropriate fragmentation experiment.
Machine learning offered a way to expand the decision beyond two diagnostic ions.
Rather than considering only a small number of manually selected signals, the new approach can evaluate approximately 50 oxonium-ion intensities simultaneously. Published, well-annotated glycoproteomics datasets can then be used to train models capable of recognizing more complex spectral patterns associated with different glycopeptide classes.
For this work, the group uses an in-house tool called GlyCounter to extract oxonium-ion intensities from spectra. These values are subsequently provided to an XGBoost machine-learning model.
According to Russell, offline extraction with GlyCounter may take several minutes for a large dataset, while XGBoost itself can train and perform classifications extremely quickly. More importantly for intelligent acquisition, the pre-trained model can classify spectra on a millisecond timescale, making it suitable for real-time operation during LC-MS analysis.
Machine learning improves glycopeptide classification
The machine-learning work was performed in collaboration with the laboratory of William Noble at the University of Washington.
When several machine- and deep-learning approaches were benchmarked against the earlier real-time library-searching strategy using a large published glycopeptide dataset, the learning-based methods substantially improved classification performance. Russell highlighted an especially important result: XGBoost was able to correctly recognize certain complex O-glycopeptide spectra that the simpler rule-based method had classified as N-glycopeptides.
Recovering these spectra means that information previously missed during acquisition can now be retained and subjected to the more appropriate fragmentation strategy.
Gokey drew parallels with challenges in untargeted metabolomics, where researchers also work with increasingly large spectral libraries and computational tools for interrogating characteristic fragments and mass-to-charge ratios. Their discussion highlighted a broader trend across mass spectrometry: analytical workflows are increasingly moving from acquiring everything first and interpreting it later toward using information generated during the experiment to guide what the instrument does next.
Getting more information from the mass spectrometer
Russell sees intelligent data acquisition as a way to make better use of modern instrumentation.
Rather than treating an LC-MS method as a fixed sequence of predetermined events, real-time processing allows the instrument to react to the data it is generating. Classification or spectral interpretation can therefore influence which precursor is selected, what fragmentation method is used, or which additional experiment should be performed.
He stressed that further development will depend on cooperation between academic researchers, industrial scientists, and instrument manufacturers, particularly through broader support for real-time mass spectrometry capabilities.
Gokey noted that similar concepts could eventually be valuable in areas such as native mass spectrometry, where researchers may benefit from observing molecular behavior and adjusting analytical parameters while an experiment is still running. Both researchers see computing resources, software integration, and access to real-time instrument control as important factors determining how quickly such strategies become routine.
Simplifying highly complex glycans
Another part of Russell's research focuses on particularly large N-glycans containing poly-N-acetyllactosamine (poly-LacNAc) structures.
These elongated glycans consist of repeating monosaccharide units and are biologically important, including in immune processes, tumor invasion, and metastasis. Their size and structural microheterogeneity, however, create major analytical problems. Large glycopeptides may ionize poorly, while the number of possible glycan configurations can make spectral interpretation difficult.
The researchers are exploring an enzymatic approach that removes the extended repeating region without removing the entire glycan from the peptide. Using endo-β-galactosidase, the poly-LacNAc chain can be shortened while retaining enough of the glycan structure to preserve biologically relevant information.
The resulting glycopeptides are easier to ionize and exhibit reduced microheterogeneity, simplifying intact glycoproteomic analysis. The concept had previously been demonstrated at the glycomics level, and the group has been investigating how to incorporate it into intact glycopeptide workflows.
Exploring glycosylation in melanoma
One application of this workflow is the comparison of cutaneous and uveal melanoma.
The researchers are interested in whether differences in poly-LacNAc glycosylation could contribute to the markedly different metastatic behavior of these diseases. Russell explained that poly-LacNAc structures have previously been associated with tumor-cell invasion and migration, making them an interesting target for investigating biological differences between the two melanoma types.
This work illustrates the wider motivation behind the development of new glycoproteomics technologies. Improvements in enrichment, fragmentation, intelligent acquisition, and computational interpretation are not only about generating more identifications. They can make previously inaccessible biological questions experimentally tractable.
The discussion between Russell and Gokey ultimately points toward a future in which mass spectrometers become increasingly adaptive analytical systems. By combining real-time spectral information with machine learning and carefully designed chemical workflows, researchers may be able to make better fragmentation decisions, recover information that would otherwise be lost, and analyze increasingly complex biological systems with greater depth and efficiency.
This text has been automatically transcribed from a video presentation using AI technology. It may contain inaccuracies and is not guaranteed to be 100% correct.
Concentrating on Chromatography Podcast
Dive into the frontiers of chromatography, mass spectrometry, and sample preparation with host David Oliva. Each episode features candid conversations with leading researchers, industry innovators, and passionate scientists who are shaping the future of analytical chemistry. From decoding PFAS detection challenges to exploring the latest in AI-assisted liquid chromatography, this show uncovers practical workflows, sustainability breakthroughs, and the real-world impact of separation science. Whether you’re a chromatographer, lab professional, or researcher you'll discover inspiring content!
You can find Concentrating on Chromatography Podcast in podcast apps:
and on YouTube channel

_s.webp)


