Unlocking the discriminative potential of MALDI-TOF profiling data with AI-enabled software for biomarker discovery, classification and typing

Posters | 2026 | Bruker | ASMSInstrumentation
LC/MS, LC/TOF, MALDI, Software
Industries
Clinical Research, Metabolomics
Manufacturer
Bruker

Significance of the topic


Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) is a high-throughput profiling technology that provides rapid access to broad analyte spaces (proteins, peptides, glycans) with minimal sample preparation. Combining MALDI-TOF profiling with AI-enabled data analysis unlocks increased discriminative power for biomarker discovery, sample classification and typing across life-science, biopharma and food/forensics applications. Efficient software workflows that integrate preprocessing, statistics and machine learning (ML) are essential to translate complex spectral data into robust, actionable results.

Objectives and overview of the study


The study demonstrates how a workflow-oriented, cloud-based software platform (Clover MSDA) leverages AI/ML methods to extract discriminative information from MALDI-TOF profiling datasets. Three representative application examples are used to illustrate capability and performance:
  • Biomarker analysis: detection of discriminative m/z features in human blood serum spiked with artificial peptides versus non-spiked controls.
  • Classification: unsupervised multivariate separation and clustering of research-grade nivolumab biosimilar samples using N-glycan MALDI-TOF spectra.
  • Identification/typing: supervised ML (support vector machine, SVM) model training and validation to verify closely related Salmonidae fish species from tissue-extract protein profiles.

Methods and data analysis workflow


Data acquisition and preprocessing:
  • Instruments: Bruker neofleX, autoFlex and microFlex series (data acquired in Bruker proprietary format).
  • Acquisition modes: protein profiles in positive linear mode; N-glycan profiles in positive reflector mode.
  • Preprocessing steps: smoothing, baseline correction, m/z alignment and generation of m/z feature tables (average-mass features for protein profiles; monoisotopic features for N-glycans where stated).
  • Typical dataset sizes and m/z ranges reported: biomarker study — 20 spectra per category, m/z 1,000–10,000; N-glycan classification — 30 spectra per sample, m/z 1,000–3,500; fish species typing — training 15–16 spectra/species, validation 6–8 spectra/species, m/z 2,000–30,000.
Statistical and machine-learning toolset (Clover MSDA):
  • Biomarker Analysis workflow: univariate tests (t-test, Mann–Whitney U) and ROC/AUC reporting with multiple-testing correction (q-values).
  • Classification workflow: unsupervised methods (principal component analysis, hierarchical clustering, k-means) and supervised multivariate/ML methods (PLS-DA, SVM, Random Forest, KNN, LightGBM). Models from supervised classifiers can be exported and reused in an Identification workflow.
  • Model validation: cross-validation (10-fold example shown), confusion matrices, prediction probabilities and balanced accuracy metrics used for performance assessment.

Used instrumentation


  • neofleX (Bruker)
  • autoFlex (Bruker)
  • microFlex (Bruker)

Main results and discussion


Biomarker analysis (spiked serum):
  • Univariate testing and ROC analysis identified a set of discriminative m/z features that matched the spiked peptides and their oxidized variants. Reported features showed highly significant p-values and q-values (many p < 1e-9) and high AUC values (several AUC = 1.00), indicating near-perfect separation between spiked and non-spiked groups for those peaks.
  • Complementary visualization (butterfly plots, peak heatmaps, PCA-based outlier detection and Pearson correlation maps) supported the robustness of these discriminative features.
Classification of nivolumab biosimilars (N-glycan profiles):
  • Hierarchical clustering and PCA separated the four biosimilar sample classes successfully. The fucosyltransferase (FT) knockout IgG1 biosimilar was clearly distinct due to the absence of core-fucosylated N-glycans, and PCA further discriminated samples by IgG subtype (IgG1 vs IgG4) along a secondary component.
SVM-based species verification (Salmonidae):
  • An SVM model trained on protein-profile spectra achieved excellent cross-validation performance and correctly identified 17 of 18 blinded/unknown spectra with high prediction probabilities (>75%). One sample was assigned with medium probability (55%–75%).
  • Model loadings and confusion-matrix analysis demonstrated clear discriminative spectral features between Arctic char, Atlantic salmon and king salmon, supporting reliable species-level typing using MALDI-TOF protein profiles.
Limitations and considerations:
  • Reported results are strong for the illustrated datasets but derive from controlled experimental conditions and moderate sample sizes; broader external validation is required before diagnostic or regulatory use.
  • Factors that influence robustness include sample preparation variability, instrument calibration, feature-detection parameters (average vs monoisotopic masses), and potential overfitting when many models or hyperparameters are considered.

Benefits and practical applications of the workflow


  • Rapid screening and targeted biomarker discovery in biofluids using automated preprocessing and statistical ranking of discriminative m/z features.
  • Quality control and comparability assessment of biopharmaceuticals (e.g., biosimilars) through glycan profiling and unsupervised clustering approaches.
  • Species authentication and food/forensics typing using supervised classifiers exported as reusable prediction models for routine identification workflows.
  • Cloud-based, workflow-oriented UI enables scalable cohort analysis, reproducible reporting, role-based user/project management and model reuse — practical for research labs, QC departments and service laboratories.

Future trends and potential applications


  • Deeper integration of advanced ML and deep-learning approaches (convolutional networks for spectral patterns, representation learning) to capture complex, nonlinear discriminative signals across datasets.
  • Transfer learning and domain-adaptation strategies to improve robustness across instruments, sample-prep protocols and labs, reducing the need for large retraining cohorts.
  • Federated learning and privacy-preserving model sharing to build broader spectral libraries while protecting sensitive data.
  • Standardized spectral reference libraries and validated ML pipelines to support regulatory acceptance for QC and clinical applications; automation of end-to-end workflows (sample-to-report) for high-throughput environments.
  • Hybrid multimodal workflows combining MALDI-TOF profiling with orthogonal measurements (LC-MS, top-down proteomics, glycomics) to increase confidence in biomarker and identity assignments.

Conclusion


Clover MSDA demonstrates that coupling MALDI-TOF profiling with a comprehensive, workflow-oriented AI toolbox enables efficient biomarker discovery, robust classification and accurate sample typing in diverse application domains. The presented examples—serum peptide biomarker detection, N-glycan-based biosimilar classification, and SVM-based species verification—illustrate practical strengths and typical performance metrics. For broader deployment, rigorous external validation, attention to preanalytical standardization and careful model governance are recommended.

References


  1. Huffman G., Mancera L., Asperger A. Bruker Technical Note TN-63. Bruker Daltonics; 2026.

Content was automatically generated from an orignal PDF document using AI and may contain inaccuracies.

Downloadable PDF for viewing
 

Similar PDF

Seafood Authenticity Testing System Using PCR-RFLP and Bioanalyzer Technology
Use of MALDI-TOF mass spectrometry and machine learning to detect the adulteration of extra virgin olive oils
GC-MS-IRMS: Addressing authenticity of fish oils by carbon and hydrogen isotope fingerprints
DLLME-Gas Chromatography-QuadrupoleTime-of-Flight Mass Spectrometry for Classification of Botanical Origin of Chinese Honey