Identifications Folder: Raw Files Created in MS/MS NIST26
- Photo: Mass Spec Interpretation Services/James Little: Identifications Folder: Raw Files Created in MS/MS NIST26
- Video: Mass Spec Interpretation Services/James Little: Identifications Folder: Raw Files Created in MS/MS NIST26
When NIST26 processes LC-MS/MS data in the Chromatogram Window, the final list of compound identifications is only one part of the story. Behind the results displayed in the software lies an entire group of intermediate files containing chromatographic traces, extracted MS/MS spectra, search results, isotope information, peak models and associations between chromatographic features and spectra. In a recent handout, James Little of Mass Spec Interpretation Services examined what these files represent, how they relate to the new XIC-centric processing workflow in NIST26, and what additional information they may offer to analysts. identifications-raw-files
The starting point is the Identifications folder created when an LC-MS/MS file is processed in the NIST26 Chromatogram Window. Rather than producing a single result file, the software generates a group of intermediate outputs associated with MS/MS deconvolution and library searching. These data are subsequently used to construct the information formally presented in the Chromatogram Window. The handout also explores whether some of these outputs could have value beyond the immediate identification workflow, for example as inputs for multivariate analysis or sample comparison. identifications-raw-files
An XIC-centric approach to LC-MS/MS processing
The workflow described in the presentation is built around extracted ion chromatograms (XICs). Instead of treating every MS/MS spectrum as an isolated identification event, NIST26 relates spectral information back to chromatographic behavior.
The workflow first extracts chromatographic information from MS1 data, groups spectra according to retention time, associates MS/MS spectra with XIC peaks and then uses this information to validate or refine identifications. The diagrams on pages 8–15 of the handout illustrate this as a multistep process rather than a simple “MS/MS spectrum → library match → identification” sequence. identifications-raw-files
This distinction is important because the chromatographic dimension provides additional evidence. A candidate identification can be evaluated not only from its MS/MS library match, but also according to whether the associated precursor behaves consistently in the chromatogram, whether its isotopic pattern makes sense and whether the MS/MS spectrum is correctly linked to the relevant chromatographic peak.
Parallel processing generates multiple intermediate files
One reason the Identifications folder contains so many files is that NIST26 divides the raw mzML dataset into multiple chunks for parallel processing. Each chunk generates its own intermediate outputs.
The presentation describes several important file types. The .tic files contain total ion chromatogram information based on MS1 intensity versus retention time and are used for processes such as peak detection, retention-time alignment and XIC windowing. The .mgf files contain peak-picked MS/MS spectra in Mascot Generic Format and serve as inputs for library searching. The .tsv files contain tabulated search results, including candidate identifications and associated scores and precursor information. identifications-raw-files
These intermediate files therefore correspond to different levels of the processing workflow: chromatographic information from MS1, spectral information from MS2, library-search results and the subsequent cross-validation of these layers.
From thousands of spectra to the final reported identifications
The handout gives a useful example of how strongly the dataset can be reduced during processing.
In the example discussed, the original mzML file contained 3,380 spectra, which were divided into four processing chunks. From these data, 216 tandem MS/MS spectra were extracted for library searching. Ultimately, 158 MS2 searches survived into the final displayed or reportable results. According to James Little's note, processing those 158 spectra took approximately 11 seconds. identifications-raw-files
This illustrates that the final result table represents only a selected subset of the spectral information originally present in the file. The intermediate outputs preserve information about how that subset was generated.
What is inside the “check” files?
A central point of the presentation is that the files informally referred to as “check” files are not simply temporary clutter. Conceptually, they form an intermediate evidence layer between the original MS/MS library search and the final identification.
According to the workflow shown in the handout, these files may contain XIC traces for individual precursor ions, isotope traces, peak models, retention-time associations, in-source ion relationships and grouping information describing which MS/MS spectra belong to which chromatographic peaks. identifications-raw-files
The workflow can therefore use chromatographic evidence to confirm or reject candidate identifications, refine precursor assignments and distinguish useful MS/MS spectra from background signals. The presentation contrasts this with a more traditional workflow in which an MS/MS library match might effectively represent the end of the identification process. In the NIST26 XIC-centric approach, the sequence becomes closer to MS/MS search → identification candidate → XIC validation → refinement and re-ranking. identifications-raw-files
Chromatographic information adds another layer of confidence
The practical value of this additional processing lies in combining information that would otherwise be considered separately.
An MS/MS spectrum may produce a convincing library match, but NIST26 can additionally examine whether the precursor generates a coherent chromatographic peak, whether the expected isotopic signals appear together and whether the retention-time relationships are consistent. The internal datasets can also help track in-source fragments, which could otherwise be interpreted as independent compounds.
The final slide of the handout summarizes these files as the core evidence layer of the new NIST approach, containing chromatographic information, isotope behavior, retention-time consistency and in-source relationships. According to the presentation, these data help correct precursor assignments, recognize in-source fragments and improve overall MS/MS identification confidence. identifications-raw-files
Can the outputs be used beyond compound identification?
Another interesting question raised in the presentation is whether the information generated during NIST26 processing could be reused for other data-analysis tasks.
The .tsv output generally contains one row per detected chromatographic feature and may include retention time, precursor m/z, peak area or intensity, library-match score, compound name, molecular formula and adduct information where available. identifications-raw-files
The handout suggests that this type of tabular output could provide a useful starting point for building a feature matrix for downstream analyses such as:
- principal component analysis (PCA),
- hierarchical clustering,
- comparison of “good” and “bad” samples,
- outlier detection,
- and marker discovery. identifications-raw-files
Importantly, this part of the handout is presented as an exploration of possible utility rather than as a formally demonstrated NIST26 workflow. James Little notes that ChatGPT was used to examine the relationships among the generated files and consider how they might potentially be used. identifications-raw-files
Why does NIST26 create so many files?
The large number of intermediate files is ultimately a consequence of two characteristics of the processing strategy: parallelization and multi-level analysis.
The raw data are split into several chunks to accelerate computation, while separate data products are generated for MS1 chromatographic processing, MS2 spectral searching and subsequent cross-validation. The workflow therefore preserves intermediate evidence at several stages rather than collapsing everything immediately into a single final result table.
The result is a folder that can initially appear unnecessarily complex, but which reflects the architecture of the analysis itself. As emphasized in the handout, the intermediate files are not simply processing leftovers; they document how chromatographic and mass-spectral information was combined to arrive at the final identification. identifications-raw-files
More than a library search
The main message of the presentation is that the new NIST26 LC-MS/MS workflow goes beyond conventional spectrum-to-library matching. By incorporating XIC behavior, isotope information, peak relationships, retention-time consistency and MS/MS evidence, the software can use multiple independent pieces of information to evaluate an identification.
For analysts, understanding what is stored in the Identifications folder provides a clearer picture of how the final results are generated. It also opens the possibility of inspecting intermediate processing steps when troubleshooting questionable identifications and, potentially, reusing some of the tabulated feature-level outputs for broader sample-comparison workflows.
The growing importance of these intermediate datasets also reflects a wider shift in LC-MS/MS analysis: confident identification increasingly depends not on a single score or spectrum, but on the combined interpretation of chromatographic, spectral and contextual evidence.




