Skip to content
matchmsPublic

About

Python library for processing (tandem) mass spectrometry data and for computing spectral similarities.

Topics

Resources

Contributing

Stars

273 stars

Watchers

10 watching

Forks

 
 

Repository files navigation

GitHub Badge License Badge Conda Badge Pypi Badge Research Software Directory Badge Zenodo Badge CII Best Practices Badge Howfairis badge

Code quality checks:

Continuous integration workflow Continuous integration workflow Documentation Status Sonarcloud Quality Gate Sonarcloud Coverage

matchms

matchms

Matchms is an open-source Python package for importing, processing, cleaning, exporting, and comparing tandem mass spectrometry data (MS/MS). It supports reproducible workflows that transform raw spectra from common file formats into cleaned, harmonized, and comparable spectral datasets.

The preferred way to work with matchms is now through SpectraCollection: a collection-level representation for complete MS/MS datasets. A SpectraCollection keeps metadata and fragment peak data synchronized, supports table-like inspection and filtering, and provides a natural basis for scalable dataset-level processing and similarity computation.

Matchms supports popular spectral data formats including mzML, mzXML, MSP, MGF, metabolomics-USI, JSON, and pickle. It provides tools for metadata harmonization, metadata validation, peak filtering, spectrum processing, collection processing, export, and large-scale spectral similarity calculations.

The classic Spectrum API, which was the default in matchms < 1.0, remains supported. Individual spectra are still represented as Spectrum objects, and existing workflows that process lists of spectra continue to work. For new workflows, however, SpectraCollection is recommended whenever a complete dataset is imported, cleaned, filtered, exported, or compared.

matchms code design

Citation

If you use matchms in your research, please cite the following software papers:

F Huber, S. Verhoeven, C. Meijer, H. Spreeuw, E. M. Villanueva Castilla, C. Geng, J.J.J. van der Hooft, S. Rogers, A. Belloum, F. Diblen, J.H. Spaaks, (2020). matchms - processing and similarity evaluation of mass spectrometry data. Journal of Open Source Software, 5(52), 2411, https://doi.org/10.21105/joss.02411

de Jonge NF, Hecht H, Michael Strobel, Mingxun Wang, van der Hooft JJJ, Huber F. (2024). Reproducible MS/MS library cleaning pipeline in matchms. Journal of Cheminformatics, 2024, https://jcheminf.biomedcentral.com/articles/10.1186/s13321-024-00878-1

Quick start: collection-first workflow

A typical matchms workflow starts by loading an MS/MS dataset directly as a SpectraCollection:

from matchms.importing import load_ms2_dataset

collection = load_ms2_dataset("my_spectra.mgf")  # you could here specify the required precision, default is: mz_precision=0.000001

print(collection)
print(collection.metadata.head())
print(collection.n_spectra)

Filters can then be applied directly to the collection:

from matchms.filtering import (
    harmonize_missing_entries,
    select_by_relative_intensity,
    require_minimum_number_of_peaks,
)

collection = harmonize_missing_entries(collection)
collection = select_by_relative_intensity(
    collection,
    intensity_from=0.01,
    intensity_to=1.0,
)
collection = require_minimum_number_of_peaks(collection, n_required=5)

The processed collection can be exported again:

collection.to_mgf("processed_spectra.mgf")
collection.to_msp("processed_spectra.msp")
collection.to_json("processed_spectra.json")

Similarity scores can be computed directly from the processed collection:

from matchms.similarity import ModifiedCosine

similarity = ModifiedCosine(tolerance=0.01)
scores = similarity.matrix(collection)

For repeated searches against a fixed reference library, build the library index once and reuse it:

library_index = similarity.build_index(collection)

scores = similarity.search(
    query_collection,
    library_index,
)

The index can also be saved and loaded again later:

similarity.save_index(
    library_index,
    "modified_cosine_index.npz",
)

library_index = similarity.load_index(
    "modified_cosine_index.npz"
)

Core concepts

SpectraCollection

SpectraCollection is the central matchms representation for complete MS/MS datasets. It stores many spectra in a synchronized collection layout and keeps spectrum-level metadata aligned with fragment peak data.

A collection separates dataset storage into:

  • MetadataCollection: a pandas-based table where each row corresponds to one spectrum.
  • FragmentCollection: a backend for storing all fragment peaks. The default backend is CSRFragmentCollection, which stores peaks in a sparse matrix with spectra as rows and binned m/z values as columns.

The central invariant is:

len(collection.metadata) == len(collection.fragments) == collection.n_spectra

Metadata row i and fragment row i always describe the same spectrum. Operations such as slicing, filtering, sorting, dropping, and deduplication preserve this alignment.

Example:

from matchms.importing import load_ms2_dataset

collection = load_ms2_dataset("my_spectra.mgf")

print(collection)
print(collection.metadata.head())
print(collection.describe())

first_spectrum = collection[0]

SpectraCollection supports dataset-level selection while keeping metadata and fragment data synchronized:

# Select spectra by metadata
positive = collection.filter(collection.metadata["ionmode"] == "positive")

# Sort spectra by metadata
sorted_collection = collection.sort("precursor_mz")

# Select spectra and restrict m/z range
selected = collection[:100, 50.0:500.0]

# Drop spectra without peaks
collection = collection.drop_empty_spectra()

# Drop duplicate spectra
collection = collection.drop_duplicates()

Individual rows can still be accessed as regular Spectrum objects:

spectrum = collection[0]

print(spectrum.peaks.mz)
print(spectrum.get("precursor_mz"))

It is also possible to run a simple for loop: for spectrum in collection:

Spectrum

Spectrum represents one mass spectrum. It contains:

  • Fragments: the m/z and intensity arrays of one spectrum.
  • Metadata: one spectrum-level metadata dictionary.

Spectrum is useful for individual spectra, custom spectrum-wise algorithms, and backward-compatible workflows (e.g., for projects partly using matchms < 1.0).

Example:

import numpy as np
from matchms import Spectrum

spectrum = Spectrum(
    mz=np.array([100.0, 150.0, 200.0]),
    intensities=np.array([0.1, 0.5, 1.0]),
    metadata={
        "precursor_mz": 201.1,
        "ionmode": "positive",
        "smiles": "CCCO",
    },
)

print(spectrum.peaks.mz)
print(spectrum.get("precursor_mz"))

Importing datasets

For new workflows, use load_ms2_dataset to import a complete dataset as a SpectraCollection:

from matchms.importing import load_ms2_dataset

collection = load_ms2_dataset("my_spectra.mgf")

The file type is detected automatically from the file extension. Supported file types include mzML, mzXML, mgf, msp, json, and pickle.

If needed, the file type can be specified explicitly:

collection = load_ms2_dataset("my_file.txt", ftype="mgf")

Metadata harmonization is enabled by default:

collection = load_ms2_dataset(
    "my_spectra.mgf",
    metadata_harmonization=True,
)

For lower-level workflows, individual importers remain available:

from matchms.importing import load_from_mgf

spectra = list(load_from_mgf("my_spectra.mgf"))

The helper load_spectra returns spectra as Spectrum objects or iterables of Spectrum objects:

from matchms.importing import load_spectra

spectra = list(load_spectra("my_spectra.mgf"))

Exporting datasets

SpectraCollection objects can be exported directly:

collection.to_mgf("processed_spectra.mgf")
collection.to_msp("processed_spectra.msp")
collection.to_json("processed_spectra.json")

MGF and MSP export support appending to an existing file:

collection.to_mgf("combined_spectra.mgf", append=True)
collection.to_msp("combined_spectra.msp", append=True)

The export style can be selected where supported:

collection.to_mgf("processed_spectra.mgf", export_style="gnps")
collection.to_msp("processed_spectra.msp", export_style="nist")

The classic save_spectra wrapper remains available for workflows that work with individual Spectrum objects or lists of spectra:

from matchms.exporting import save_spectra

save_spectra(spectra, "processed_spectra.mgf")

Filtering and processing

Many matchms filters support both Spectrum and SpectraCollection input. When a filter receives a collection, it operates on all rows while preserving the alignment between metadata and fragment data.

from matchms.filtering import (
    harmonize_missing_entries,
    select_by_relative_intensity,
    require_minimum_number_of_peaks,
)

collection = harmonize_missing_entries(collection)
collection = select_by_relative_intensity(collection, intensity_from=0.01)
collection = require_minimum_number_of_peaks(collection, n_required=5)

For filters that remove spectra, the behavior depends on the input type:

  • For Spectrum input, a failing spectrum returns None.
  • For SpectraCollection input, failing rows are removed from both metadata and fragments.

This keeps the collection synchronized throughout the workflow.

SpectraProcessor

SpectraProcessor provides a single processing pipeline for both individual Spectrum objects and complete SpectraCollection objects.

The processor uses the same ordered list of filters for both execution paths, but provides explicit methods for choosing how the pipeline is applied:

  • process_spectrum() applies all filters sequentially to one Spectrum.
  • process_collection() applies each filter to the complete SpectraCollection, allowing collection-native and vectorized filter implementations to be used.

For collection-based workflows:

from matchms import SpectraProcessor
from matchms.importing import load_ms2_dataset

collection = load_ms2_dataset("my_spectra.mgf")

processor = SpectraProcessor(
    filters=[
        "harmonize_missing_entries",
        ("select_by_relative_intensity", {"intensity_from": 0.01}),
        ("require_minimum_number_of_peaks", {"n_required": 5}),
    ]
)

processed_collection = processor.process_collection(collection)

For individual spectra, the same processor and filter configuration can be used:

from matchms import SpectraProcessor
from matchms.filtering.default_pipelines import BASIC_FILTERS
from matchms.importing import load_spectra

spectra = list(load_spectra("my_spectra.mgf"))
processor = SpectraProcessor(BASIC_FILTERS)

processed_spectra = [
    processed
    for spectrum in spectra
    if (processed := processor.process_spectrum(spectrum)) is not None
]

Both processing methods copy the input once before applying the pipeline, so the original Spectrum or SpectraCollection is not modified.

SpectraProcessor accepts filter descriptions in several forms:

filters = [
    "harmonize_missing_entries",
    ("select_by_intensity", {"intensity_from": 10.0, "intensity_to": 1000.0}),
    custom_filter_function,
    (custom_filter_with_parameters, {"parameter": "value"}),
]

Known matchms filters are ordered according to the matchms filter order. Custom filters are appended unless a specific position is supplied.

Processing reports

SpectraProcessor can optionally collect a processing report that summarizes how much each filter changed the data.

processor = SpectraProcessor(
    filters=[
        "harmonize_missing_entries",
        ("select_by_relative_intensity", {"intensity_from": 0.01}),
        ("require_minimum_number_of_peaks", {"n_required": 5}),
    ]
)

report = processor.create_processing_report()

processed_collection = processor.process_collection(
    collection,
    processing_report=report,
)

print(report.to_dataframe())

The report records, for each filter:

  • the number of input spectra,
  • the number of output spectra,
  • the number of removed spectra,
  • the number of spectra with changed metadata,
  • the number of spectra with changed fragments.

Changes are detected using metadata and fragment hashes, so reporting does not require keeping full copies of the data before every processing step.

The same reporting mechanism can also be used with process_spectrum():

report = processor.create_processing_report()

processed_spectrum = processor.process_spectrum(
    spectrum,
    processing_report=report,
)

Metadata handling

For collections, metadata is stored in MetadataCollection, a pandas-based table. Each row corresponds to one spectrum, and columns correspond to metadata fields.

print(collection.metadata.head())
print(collection.metadata.columns)

Metadata column names can be harmonized using matchms key conventions:

collection = collection.harmonize_metadata_columns()

For example:

"Precursor MZ" -> "precursor_mz"
"Compound Name" -> "compound_name"

Missing metadata values can be harmonized with:

from matchms.filtering import harmonize_missing_entries

collection = harmonize_missing_entries(collection)

By default, common aliases for missing values such as "", "N/A", "NA", "n/a", "NaN", "None", and "no data" are interpreted as missing entries.

Older specialized filters such as harmonize_undefined_inchi, harmonize_undefined_inchikey, and harmonize_undefined_smiles are kept for backward compatibility but are deprecated in favor of harmonize_missing_entries.

For individual spectra, metadata is stored in a Metadata object and can be accessed with:

spectrum.get("precursor_mz")
spectrum.set("compound_name", "example")

Peak data handling

For collections, peak data is stored in a FragmentCollection backend. The default backend is CSRFragmentCollection, which stores peaks in a sparse matrix.

This enables efficient dataset-level operations such as:

  • counting peaks per spectrum,
  • selecting peaks by intensity,
  • selecting peaks by relative intensity,
  • filtering spectra by number of peaks,
  • slicing m/z ranges,
  • computing fragment hashes,
  • computing summary statistics.

Example:

peak_counts = collection.fragments.count(axis=1)
intensity_sums = collection.fragments.sum(axis=1)

selected = collection[:, 50.0:500.0]

Because the default collection backend uses binned sparse storage, spectra reconstructed from a SpectraCollection may contain m/z values corresponding to bin centers rather than the exact original m/z values. This is important when testing for exact m/z equality; use numerical tolerances where appropriate.

For individual spectra, peak data is available through spectrum.peaks:

spectrum.peaks.mz
spectrum.peaks.intensities

Similarity measures

Matchms provides several similarity measures in matchms.similarity for comparing mass spectra, spectrum metadata, and molecular structures.

For most peak-based spectral comparisons, start with one of the high-level classes Cosine, ModifiedCosine, Entropy, or EntropySearch. These classes provide a common interface for comparing individual spectra, computing complete similarity matrices, and, where applicable, repeatedly searching a fixed reference library using a reusable index.

Similarity measures at a glance
Similarity Recommended class Typical use Specialized implementations
Cosine Cosine Standard peak-based spectral similarity with one-to-one peak matching. Supports reusable library indices for repeated searches. CosineGreedy, CosineHungarian, CosineLinear, CosineFlash, CosineBlink
Modified cosine ModifiedCosine Cosine similarity allowing both direct fragment matches and matches shifted by the difference in precursor m/z. Supports reusable library indices for repeated searches. ModifiedCosineGreedy, ModifiedCosineHungarian, ModifiedCosineLinear; CosineFlash with matching_mode="hybrid"
Spectral entropy Entropy General-purpose spectral entropy similarity with explicit one-to-one matching. Supports fragment, neutral-loss, and hybrid matching as well as reusable library indices. EntropyGreedy, EntropyFlash
Search-optimized spectral entropy EntropySearch High-throughput fragment-only entropy searches against large reference libraries. Requires sufficiently separated peaks and can optionally merge close peaks during preparation.
Neutral-loss cosine NeutralLossesCosine Compare spectra based on neutral-loss rather than fragment m/z patterns.
Binned spectra BinnedEmbeddingSimilarity Compare fixed-width binned spectrum representations using cosine or Euclidean similarity.
Molecular structure FingerprintSimilarity Compare molecular fingerprints derived from structure metadata.
Metadata MetadataMatch Compare arbitrary metadata fields using exact or tolerance-based matching.
Precursor or parent mass PrecursorMzMatch, ParentMassMatch Simple matching based on precursor m/z or parent mass.

Pair and matrix comparisons

Similarity matrices can be computed directly from a SpectraCollection:

from matchms.similarity import Entropy

similarity = Entropy(tolerance=0.02)
scores = similarity.matrix(collection)

The same API can be used to compare two collections:

from matchms.similarity import ModifiedCosine

similarity = ModifiedCosine(tolerance=0.01)
scores = similarity.matrix(references, queries)

Rows of the resulting matrix correspond to the first input collection and columns to the second.

Pairwise scoring of individual spectra is also supported:

from matchms.similarity import Cosine

similarity = Cosine(tolerance=0.01)
score = similarity.pair(spectrum_1, spectrum_2)

Repeated library searches with reusable indices

Cosine, ModifiedCosine, Entropy, and EntropySearch support repeated searches against a fixed reference library.

For this workflow, construct the library index once with build_index and then use search for one or more query collections:

from matchms.similarity import Cosine

similarity = Cosine(tolerance=0.01)

library_index = similarity.build_index(reference_library)

scores = similarity.search(
    query_collection,
    library_index,
)

Unlike matrix, search always treats the input spectra as queries. The returned score matrix therefore has query spectra as rows and reference library spectra as columns.

The same workflow is available for modified cosine:

from matchms.similarity import ModifiedCosine

similarity = ModifiedCosine(tolerance=0.01)

library_index = similarity.build_index(reference_library)
scores = similarity.search(query_collection, library_index)

and for general spectral entropy similarity:

from matchms.similarity import Entropy

similarity = Entropy(
    matching_mode="fragment",
    tolerance=0.01,
)

library_index = similarity.build_index(reference_library)
scores = similarity.search(query_collection, library_index)

Building the index separately is especially useful when many query batches are searched against the same reference library, because the reference spectra do not need to be prepared and indexed again for every search.

matchms similarity implementations and class inheritances

Saving and loading similarity indices

A built index can be stored on disk and reused in a later Python session:

similarity.save_index(
    library_index,
    "library_index.npz",
)

Load it again with a similarity object using the same scoring and preprocessing configuration:

from matchms.similarity import Cosine

similarity = Cosine(tolerance=0.01)

library_index = similarity.load_index(
    "library_index.npz"
)

scores = similarity.search(
    query_collection,
    library_index,
)

Index compatibility is checked when the index is loaded or used. An index should therefore be loaded with a similarity configuration matching the one used to create it.

The same build_index, save_index, load_index, and search workflow is available for Cosine, ModifiedCosine, Entropy, and EntropySearch.

For Cosine and ModifiedCosine, persistent indexed searching uses the default indexed greedy implementation. It is not available when use_hungarian=True.

Entropy and EntropySearch

Entropy and EntropySearch calculate spectral entropy similarity but make different assumptions about peak matching.

Entropy is the general-purpose implementation. It explicitly resolves competing one-to-one peak matches and supports "fragment", "neutral_loss", and "hybrid" matching. Use Entropy when general matching behavior is required, when tolerance windows may overlap, or when neutral-loss or hybrid matching is needed.

Example:

from matchms.similarity import Entropy

similarity = Entropy(
    matching_mode="hybrid",
    tolerance=0.01,
)

scores = similarity.matrix(collection)

EntropySearch is optimized for repeated fragment-only searches against large reference libraries. It requires peaks within each prepared spectrum to be sufficiently separated so that candidate matches do not compete for the same physical peak. This avoids the additional bookkeeping required by the general Entropy implementation.

By default, close peaks are merged during preparation:

from matchms.similarity import EntropySearch

similarity = EntropySearch(
    tolerance=0.01,
    peak_separation="merge",
)

library_index = similarity.build_index(reference_library)

scores = similarity.search(
    query_collection,
    library_index,
)

Merging close peaks changes the prepared spectrum and can therefore lead to scores that differ from Entropy.

If the input spectra are already sufficiently separated, use peak_separation="raise" to keep the input peak representation unchanged and raise an error when the separation requirement is violated:

similarity = EntropySearch(
    tolerance=0.01,
    peak_separation="raise",
)

EntropySearch currently supports fragment matching with an absolute Da tolerance. Use Entropy instead when ppm tolerances, neutral-loss matching, hybrid matching, or general one-to-one matching semantics are required.

Like the other indexed similarity classes, an EntropySearch index can be saved and reused:

similarity.save_index(
    library_index,
    "entropy_search_index.npz",
)

library_index = similarity.load_index(
    "entropy_search_index.npz"
)

Specialized similarity implementations

The specialized classes are useful when a particular scoring implementation is required.

CosineGreedy provides a simple pair-oriented greedy cosine implementation, while CosineHungarian uses optimal peak assignment. CosineLinear and CosineBlink provide alternative cosine implementations, and CosineFlash exposes the indexed Flash-based cosine implementation directly.

For spectral entropy, EntropyGreedy provides a compact pair-oriented implementation and EntropyFlash exposes the general indexed implementation used by Entropy.

For most applications, the high-level Cosine, ModifiedCosine, Entropy, and EntropySearch classes are preferred.

Installation

Prerequisites:

  • Python 3.11 - 3.14
  • Anaconda or another virtual environment manager is recommended

Install matchms with conda:

conda create --name matchms python=3.13
conda activate matchms
conda install --channel bioconda --channel conda-forge matchms

Documentation for users

For more extensive documentation, see:

matchms ecosystem

Additional packages can complement matchms functionality:

  • Spec2Vec: machine-learning spectral similarity scoring.
  • MS2DeepScore: supervised deep-learning-based spectral similarity scoring.
  • matchmsextras: additional tools for networks, PubChem search, and plotting.
  • MS2Query: MS/MS spectral analogue search.
  • memo: retention-time agnostic alignment of metabolomics samples.
  • RIAssigner: retention index calculation for GC-MS data.
  • MSMetaEnhancer: metadata enrichment using web services and computational chemistry packages.
  • SimMS: GPU-based implementations of common similarity classes.

If you know of another package that is compatible with matchms, let us know.

Ecosystem compatibility

NumPy Version spec2vec Status ms2deepscore Status ms2query Status
numpy https://img.shields.io/badge/spec2vec-0.9.1-green https://img.shields.io/badge/ms2deepscore-2.7.2-green https://img.shields.io/badge/ms2query-1.5.4-red
numpy https://img.shields.io/badge/spec2vec-0.9.1-green https://img.shields.io/badge/ms2deepscore-2.7.2-green https://img.shields.io/badge/ms2query-1.5.4-red

Documentation for developers

Development installation

git clone https://github.com/matchms/matchms.git
cd matchms

# Create environment using conda
conda create --name matchms-dev python=3.13
conda activate matchms-dev

# Or create environment using uv
uv venv --python 3.13
uv sync --group dev

# Or install with pip
pip install -r dev-requirements.txt
pip install --editable .

Code quality

Run the linter and formatter:

ruff check --fix matchms/YOUR-MODIFIED-FILE.py
ruff format matchms/YOUR-MODIFIED-FILE.py

Install pre-commit hooks:

pre-commit install

Run tests:

pytest

Developer notes: collection-first design

When adding new functionality, prefer SpectraCollection support from the start. New filters should ideally support both Spectrum and SpectraCollection inputs.

For filters with separate spectrum and collection implementations, use collection_filter:

def _my_filter_spectrum(spectrum_in, ..., clone=True):
    ...

def _my_filter_collection(collection, ..., clone=True):
    ...

my_filter = collection_filter(
    _my_filter_spectrum,
    collection_impl=_my_filter_collection,
)

For metadata-only filters, prefer metadata-level implementations and dispatch helpers:

def _my_metadata_filter(metadata, ...):
    ...

my_filter = metadata_update_filter(_my_metadata_filter)

For requirement filters that keep or remove spectra, use a predicate-style metadata implementation:

def _require_my_metadata(metadata, ...) -> bool:
    ...

require_my_metadata = metadata_requirement_filter(_require_my_metadata)

General recommendations:

  • Prefer collection-native implementations for new filters.
  • Metadata-only filters should operate on MetadataCollection or metadata rows.
  • Peak-only filters should operate on FragmentCollection.
  • Filters that drop spectra should return None for failing Spectrum inputs and should drop rows for SpectraCollection inputs.
  • Filters should preserve alignment between metadata rows and fragment rows.
  • Avoid introducing pandas-specific missing values into reconstructed Spectrum metadata.
  • Prefer ValueError over assert for user-facing validation.
  • Add tests for both Spectrum and SpectraCollection behavior whenever a public filter supports both input types.

Conda package

The conda packaging is handled by a recipe at Bioconda.

Publishing to PyPI will trigger the creation of a pull request on the Bioconda recipes repository. Once the pull request is merged, the new version of matchms will appear on Anaconda.

Support

To get support, join the public Slack channel.

Contributing

If you want to contribute to matchms development, see the contribution guidelines.

License

Copyright (c) 2026, Düsseldorf University of Applied Sciences

Licensed under the Apache License, Version 2.0. You may not use this file except in compliance with the License. You may obtain a copy of the License at:

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

About

Python library for processing (tandem) mass spectrometry data and for computing spectral similarities.

Topics

Resources

Contributing

Stars

273 stars

Watchers

10 watching

Forks

Releases

Packages

Used by

Contributors

Languages