Code quality checks:
Matchms is an open-source Python package for importing, processing, cleaning, exporting, and comparing tandem mass spectrometry data (MS/MS). It supports reproducible workflows that transform raw spectra from common file formats into cleaned, harmonized, and comparable spectral datasets.
The preferred way to work with matchms is now through SpectraCollection: a collection-level representation for complete MS/MS datasets. A SpectraCollection keeps metadata and fragment peak data synchronized, supports table-like inspection and filtering, and provides a natural basis for scalable dataset-level processing and similarity computation.
Matchms supports popular spectral data formats including mzML, mzXML, MSP, MGF, metabolomics-USI, JSON, and pickle. It provides tools for metadata harmonization, metadata validation, peak filtering, spectrum processing, collection processing, export, and large-scale spectral similarity calculations.
The classic Spectrum API, which was the default in matchms < 1.0, remains supported. Individual spectra are still represented as Spectrum objects, and existing workflows that process lists of spectra continue to work. For new workflows, however, SpectraCollection is recommended whenever a complete dataset is imported, cleaned, filtered, exported, or compared.
If you use matchms in your research, please cite the following software papers:
F Huber, S. Verhoeven, C. Meijer, H. Spreeuw, E. M. Villanueva Castilla, C. Geng, J.J.J. van der Hooft, S. Rogers, A. Belloum, F. Diblen, J.H. Spaaks, (2020). matchms - processing and similarity evaluation of mass spectrometry data. Journal of Open Source Software, 5(52), 2411, https://doi.org/10.21105/joss.02411
de Jonge NF, Hecht H, Michael Strobel, Mingxun Wang, van der Hooft JJJ, Huber F. (2024). Reproducible MS/MS library cleaning pipeline in matchms. Journal of Cheminformatics, 2024, https://jcheminf.biomedcentral.com/articles/10.1186/s13321-024-00878-1
A typical matchms workflow starts by loading an MS/MS dataset directly as a
SpectraCollection:
from matchms.importing import load_ms2_dataset
collection = load_ms2_dataset("my_spectra.mgf") # you could here specify the required precision, default is: mz_precision=0.000001
print(collection)
print(collection.metadata.head())
print(collection.n_spectra)Filters can then be applied directly to the collection:
from matchms.filtering import (
harmonize_missing_entries,
select_by_relative_intensity,
require_minimum_number_of_peaks,
)
collection = harmonize_missing_entries(collection)
collection = select_by_relative_intensity(
collection,
intensity_from=0.01,
intensity_to=1.0,
)
collection = require_minimum_number_of_peaks(collection, n_required=5)The processed collection can be exported again:
collection.to_mgf("processed_spectra.mgf")
collection.to_msp("processed_spectra.msp")
collection.to_json("processed_spectra.json")Similarity scores can be computed directly from the processed collection:
from matchms.similarity import ModifiedCosine
similarity = ModifiedCosine(tolerance=0.01)
scores = similarity.matrix(collection)For repeated searches against a fixed reference library, build the library index once and reuse it:
library_index = similarity.build_index(collection)
scores = similarity.search(
query_collection,
library_index,
)The index can also be saved and loaded again later:
similarity.save_index(
library_index,
"modified_cosine_index.npz",
)
library_index = similarity.load_index(
"modified_cosine_index.npz"
)SpectraCollection is the central matchms representation for complete MS/MS
datasets. It stores many spectra in a synchronized collection layout and keeps
spectrum-level metadata aligned with fragment peak data.
A collection separates dataset storage into:
MetadataCollection: a pandas-based table where each row corresponds to one spectrum.FragmentCollection: a backend for storing all fragment peaks. The default backend isCSRFragmentCollection, which stores peaks in a sparse matrix with spectra as rows and binned m/z values as columns.
The central invariant is:
len(collection.metadata) == len(collection.fragments) == collection.n_spectra
Metadata row i and fragment row i always describe the same spectrum.
Operations such as slicing, filtering, sorting, dropping, and deduplication
preserve this alignment.
Example:
from matchms.importing import load_ms2_dataset
collection = load_ms2_dataset("my_spectra.mgf")
print(collection)
print(collection.metadata.head())
print(collection.describe())
first_spectrum = collection[0]SpectraCollection supports dataset-level selection while keeping metadata
and fragment data synchronized:
# Select spectra by metadata
positive = collection.filter(collection.metadata["ionmode"] == "positive")
# Sort spectra by metadata
sorted_collection = collection.sort("precursor_mz")
# Select spectra and restrict m/z range
selected = collection[:100, 50.0:500.0]
# Drop spectra without peaks
collection = collection.drop_empty_spectra()
# Drop duplicate spectra
collection = collection.drop_duplicates()Individual rows can still be accessed as regular Spectrum objects:
spectrum = collection[0]
print(spectrum.peaks.mz)
print(spectrum.get("precursor_mz"))It is also possible to run a simple for loop: for spectrum in collection:
Spectrum represents one mass spectrum. It contains:
Fragments: the m/z and intensity arrays of one spectrum.Metadata: one spectrum-level metadata dictionary.
Spectrum is useful for individual spectra, custom spectrum-wise algorithms,
and backward-compatible workflows (e.g., for projects partly using matchms < 1.0).
Example:
import numpy as np
from matchms import Spectrum
spectrum = Spectrum(
mz=np.array([100.0, 150.0, 200.0]),
intensities=np.array([0.1, 0.5, 1.0]),
metadata={
"precursor_mz": 201.1,
"ionmode": "positive",
"smiles": "CCCO",
},
)
print(spectrum.peaks.mz)
print(spectrum.get("precursor_mz"))For new workflows, use load_ms2_dataset to import a complete dataset as a
SpectraCollection:
from matchms.importing import load_ms2_dataset
collection = load_ms2_dataset("my_spectra.mgf")The file type is detected automatically from the file extension. Supported file
types include mzML, mzXML, mgf, msp, json, and pickle.
If needed, the file type can be specified explicitly:
collection = load_ms2_dataset("my_file.txt", ftype="mgf")Metadata harmonization is enabled by default:
collection = load_ms2_dataset(
"my_spectra.mgf",
metadata_harmonization=True,
)For lower-level workflows, individual importers remain available:
from matchms.importing import load_from_mgf
spectra = list(load_from_mgf("my_spectra.mgf"))The helper load_spectra returns spectra as Spectrum objects or iterables
of Spectrum objects:
from matchms.importing import load_spectra
spectra = list(load_spectra("my_spectra.mgf"))SpectraCollection objects can be exported directly:
collection.to_mgf("processed_spectra.mgf")
collection.to_msp("processed_spectra.msp")
collection.to_json("processed_spectra.json")MGF and MSP export support appending to an existing file:
collection.to_mgf("combined_spectra.mgf", append=True)
collection.to_msp("combined_spectra.msp", append=True)The export style can be selected where supported:
collection.to_mgf("processed_spectra.mgf", export_style="gnps")
collection.to_msp("processed_spectra.msp", export_style="nist")The classic save_spectra wrapper remains available for workflows that work
with individual Spectrum objects or lists of spectra:
from matchms.exporting import save_spectra
save_spectra(spectra, "processed_spectra.mgf")Many matchms filters support both Spectrum and SpectraCollection input.
When a filter receives a collection, it operates on all rows while preserving
the alignment between metadata and fragment data.
from matchms.filtering import (
harmonize_missing_entries,
select_by_relative_intensity,
require_minimum_number_of_peaks,
)
collection = harmonize_missing_entries(collection)
collection = select_by_relative_intensity(collection, intensity_from=0.01)
collection = require_minimum_number_of_peaks(collection, n_required=5)For filters that remove spectra, the behavior depends on the input type:
- For
Spectruminput, a failing spectrum returnsNone. - For
SpectraCollectioninput, failing rows are removed from both metadata and fragments.
This keeps the collection synchronized throughout the workflow.
SpectraProcessor provides a single processing pipeline for both individual
Spectrum objects and complete SpectraCollection objects.
The processor uses the same ordered list of filters for both execution paths, but provides explicit methods for choosing how the pipeline is applied:
- process_spectrum() applies all filters sequentially to one
Spectrum. - process_collection() applies each filter to the complete
SpectraCollection, allowing collection-native and vectorized filter implementations to be used.
For collection-based workflows:
from matchms import SpectraProcessor
from matchms.importing import load_ms2_dataset
collection = load_ms2_dataset("my_spectra.mgf")
processor = SpectraProcessor(
filters=[
"harmonize_missing_entries",
("select_by_relative_intensity", {"intensity_from": 0.01}),
("require_minimum_number_of_peaks", {"n_required": 5}),
]
)
processed_collection = processor.process_collection(collection)For individual spectra, the same processor and filter configuration can be used:
from matchms import SpectraProcessor
from matchms.filtering.default_pipelines import BASIC_FILTERS
from matchms.importing import load_spectra
spectra = list(load_spectra("my_spectra.mgf"))
processor = SpectraProcessor(BASIC_FILTERS)
processed_spectra = [
processed
for spectrum in spectra
if (processed := processor.process_spectrum(spectrum)) is not None
]Both processing methods copy the input once before applying the pipeline, so the
original Spectrum or SpectraCollection is not modified.
SpectraProcessor accepts filter descriptions in several forms:
filters = [
"harmonize_missing_entries",
("select_by_intensity", {"intensity_from": 10.0, "intensity_to": 1000.0}),
custom_filter_function,
(custom_filter_with_parameters, {"parameter": "value"}),
]Known matchms filters are ordered according to the matchms filter order. Custom filters are appended unless a specific position is supplied.
SpectraProcessor can optionally collect a processing report that summarizes
how much each filter changed the data.
processor = SpectraProcessor(
filters=[
"harmonize_missing_entries",
("select_by_relative_intensity", {"intensity_from": 0.01}),
("require_minimum_number_of_peaks", {"n_required": 5}),
]
)
report = processor.create_processing_report()
processed_collection = processor.process_collection(
collection,
processing_report=report,
)
print(report.to_dataframe())The report records, for each filter:
- the number of input spectra,
- the number of output spectra,
- the number of removed spectra,
- the number of spectra with changed metadata,
- the number of spectra with changed fragments.
Changes are detected using metadata and fragment hashes, so reporting does not require keeping full copies of the data before every processing step.
The same reporting mechanism can also be used with process_spectrum():
report = processor.create_processing_report()
processed_spectrum = processor.process_spectrum(
spectrum,
processing_report=report,
)For collections, metadata is stored in MetadataCollection, a pandas-based
table. Each row corresponds to one spectrum, and columns correspond to metadata
fields.
print(collection.metadata.head())
print(collection.metadata.columns)Metadata column names can be harmonized using matchms key conventions:
collection = collection.harmonize_metadata_columns()For example:
"Precursor MZ" -> "precursor_mz"
"Compound Name" -> "compound_name"Missing metadata values can be harmonized with:
from matchms.filtering import harmonize_missing_entries
collection = harmonize_missing_entries(collection)By default, common aliases for missing values such as "", "N/A",
"NA", "n/a", "NaN", "None", and "no data" are interpreted
as missing entries.
Older specialized filters such as harmonize_undefined_inchi,
harmonize_undefined_inchikey, and harmonize_undefined_smiles are kept for
backward compatibility but are deprecated in favor of
harmonize_missing_entries.
For individual spectra, metadata is stored in a Metadata object and can be
accessed with:
spectrum.get("precursor_mz")
spectrum.set("compound_name", "example")For collections, peak data is stored in a FragmentCollection backend. The
default backend is CSRFragmentCollection, which stores peaks in a sparse
matrix.
This enables efficient dataset-level operations such as:
- counting peaks per spectrum,
- selecting peaks by intensity,
- selecting peaks by relative intensity,
- filtering spectra by number of peaks,
- slicing m/z ranges,
- computing fragment hashes,
- computing summary statistics.
Example:
peak_counts = collection.fragments.count(axis=1)
intensity_sums = collection.fragments.sum(axis=1)
selected = collection[:, 50.0:500.0]Because the default collection backend uses binned sparse storage, spectra
reconstructed from a SpectraCollection may contain m/z values corresponding
to bin centers rather than the exact original m/z values. This is important when
testing for exact m/z equality; use numerical tolerances where appropriate.
For individual spectra, peak data is available through spectrum.peaks:
spectrum.peaks.mz
spectrum.peaks.intensitiesMatchms provides several similarity measures in matchms.similarity for
comparing mass spectra, spectrum metadata, and molecular structures.
For most peak-based spectral comparisons, start with one of the high-level
classes Cosine, ModifiedCosine, Entropy, or EntropySearch.
These classes provide a common interface for comparing individual spectra,
computing complete similarity matrices, and, where applicable, repeatedly
searching a fixed reference library using a reusable index.
| Similarity | Recommended class | Typical use | Specialized implementations |
|---|---|---|---|
| Cosine | Cosine |
Standard peak-based spectral similarity with one-to-one peak matching. Supports reusable library indices for repeated searches. | CosineGreedy,
CosineHungarian,
CosineLinear,
CosineFlash,
CosineBlink |
| Modified cosine | ModifiedCosine |
Cosine similarity allowing both direct fragment matches and matches shifted by the difference in precursor m/z. Supports reusable library indices for repeated searches. | ModifiedCosineGreedy,
ModifiedCosineHungarian,
ModifiedCosineLinear;
CosineFlash with matching_mode="hybrid" |
| Spectral entropy | Entropy |
General-purpose spectral entropy similarity with explicit one-to-one matching. Supports fragment, neutral-loss, and hybrid matching as well as reusable library indices. | EntropyGreedy,
EntropyFlash |
| Search-optimized spectral entropy | EntropySearch |
High-throughput fragment-only entropy searches against large reference libraries. Requires sufficiently separated peaks and can optionally merge close peaks during preparation. | |
| Neutral-loss cosine | NeutralLossesCosine |
Compare spectra based on neutral-loss rather than fragment m/z patterns. | |
| Binned spectra | BinnedEmbeddingSimilarity |
Compare fixed-width binned spectrum representations using cosine or Euclidean similarity. | |
| Molecular structure | FingerprintSimilarity |
Compare molecular fingerprints derived from structure metadata. | |
| Metadata | MetadataMatch |
Compare arbitrary metadata fields using exact or tolerance-based matching. | |
| Precursor or parent mass | PrecursorMzMatch,
ParentMassMatch |
Simple matching based on precursor m/z or parent mass. |
Similarity matrices can be computed directly from a SpectraCollection:
from matchms.similarity import Entropy
similarity = Entropy(tolerance=0.02)
scores = similarity.matrix(collection)The same API can be used to compare two collections:
from matchms.similarity import ModifiedCosine
similarity = ModifiedCosine(tolerance=0.01)
scores = similarity.matrix(references, queries)Rows of the resulting matrix correspond to the first input collection and columns to the second.
Pairwise scoring of individual spectra is also supported:
from matchms.similarity import Cosine
similarity = Cosine(tolerance=0.01)
score = similarity.pair(spectrum_1, spectrum_2)Cosine, ModifiedCosine, Entropy, and EntropySearch support
repeated searches against a fixed reference library.
For this workflow, construct the library index once with build_index and
then use search for one or more query collections:
from matchms.similarity import Cosine
similarity = Cosine(tolerance=0.01)
library_index = similarity.build_index(reference_library)
scores = similarity.search(
query_collection,
library_index,
)Unlike matrix, search always treats the input spectra as queries.
The returned score matrix therefore has query spectra as rows and reference
library spectra as columns.
The same workflow is available for modified cosine:
from matchms.similarity import ModifiedCosine
similarity = ModifiedCosine(tolerance=0.01)
library_index = similarity.build_index(reference_library)
scores = similarity.search(query_collection, library_index)and for general spectral entropy similarity:
from matchms.similarity import Entropy
similarity = Entropy(
matching_mode="fragment",
tolerance=0.01,
)
library_index = similarity.build_index(reference_library)
scores = similarity.search(query_collection, library_index)Building the index separately is especially useful when many query batches are searched against the same reference library, because the reference spectra do not need to be prepared and indexed again for every search.
A built index can be stored on disk and reused in a later Python session:
similarity.save_index(
library_index,
"library_index.npz",
)Load it again with a similarity object using the same scoring and preprocessing configuration:
from matchms.similarity import Cosine
similarity = Cosine(tolerance=0.01)
library_index = similarity.load_index(
"library_index.npz"
)
scores = similarity.search(
query_collection,
library_index,
)Index compatibility is checked when the index is loaded or used. An index should therefore be loaded with a similarity configuration matching the one used to create it.
The same build_index, save_index, load_index, and search
workflow is available for Cosine, ModifiedCosine, Entropy, and
EntropySearch.
For Cosine and ModifiedCosine, persistent indexed searching uses the
default indexed greedy implementation. It is not available when
use_hungarian=True.
Entropy and EntropySearch calculate spectral entropy similarity but
make different assumptions about peak matching.
Entropy is the general-purpose implementation. It explicitly resolves
competing one-to-one peak matches and supports "fragment",
"neutral_loss", and "hybrid" matching. Use Entropy when general
matching behavior is required, when tolerance windows may overlap, or when
neutral-loss or hybrid matching is needed.
Example:
from matchms.similarity import Entropy
similarity = Entropy(
matching_mode="hybrid",
tolerance=0.01,
)
scores = similarity.matrix(collection)EntropySearch is optimized for repeated fragment-only searches against
large reference libraries. It requires peaks within each prepared spectrum to
be sufficiently separated so that candidate matches do not compete for the
same physical peak. This avoids the additional bookkeeping required by the
general Entropy implementation.
By default, close peaks are merged during preparation:
from matchms.similarity import EntropySearch
similarity = EntropySearch(
tolerance=0.01,
peak_separation="merge",
)
library_index = similarity.build_index(reference_library)
scores = similarity.search(
query_collection,
library_index,
)Merging close peaks changes the prepared spectrum and can therefore lead to
scores that differ from Entropy.
If the input spectra are already sufficiently separated, use
peak_separation="raise" to keep the input peak representation unchanged
and raise an error when the separation requirement is violated:
similarity = EntropySearch(
tolerance=0.01,
peak_separation="raise",
)EntropySearch currently supports fragment matching with an absolute Da
tolerance. Use Entropy instead when ppm tolerances, neutral-loss matching,
hybrid matching, or general one-to-one matching semantics are required.
Like the other indexed similarity classes, an EntropySearch index can be
saved and reused:
similarity.save_index(
library_index,
"entropy_search_index.npz",
)
library_index = similarity.load_index(
"entropy_search_index.npz"
)The specialized classes are useful when a particular scoring implementation is required.
CosineGreedy provides a simple pair-oriented greedy cosine implementation,
while CosineHungarian uses optimal peak assignment. CosineLinear and
CosineBlink provide alternative cosine implementations, and CosineFlash
exposes the indexed Flash-based cosine implementation directly.
For spectral entropy, EntropyGreedy provides a compact pair-oriented
implementation and EntropyFlash exposes the general indexed implementation
used by Entropy.
For most applications, the high-level Cosine, ModifiedCosine,
Entropy, and EntropySearch classes are preferred.
Prerequisites:
- Python 3.11 - 3.14
- Anaconda or another virtual environment manager is recommended
Install matchms with conda:
conda create --name matchms python=3.13
conda activate matchms
conda install --channel bioconda --channel conda-forge matchmsFor more extensive documentation, see:
Additional packages can complement matchms functionality:
- Spec2Vec: machine-learning spectral similarity scoring.
- MS2DeepScore: supervised deep-learning-based spectral similarity scoring.
- matchmsextras: additional tools for networks, PubChem search, and plotting.
- MS2Query: MS/MS spectral analogue search.
- memo: retention-time agnostic alignment of metabolomics samples.
- RIAssigner: retention index calculation for GC-MS data.
- MSMetaEnhancer: metadata enrichment using web services and computational chemistry packages.
- SimMS: GPU-based implementations of common similarity classes.
If you know of another package that is compatible with matchms, let us know.
| NumPy Version | spec2vec Status | ms2deepscore Status | ms2query Status |
|---|---|---|---|
git clone https://github.com/matchms/matchms.git
cd matchms
# Create environment using conda
conda create --name matchms-dev python=3.13
conda activate matchms-dev
# Or create environment using uv
uv venv --python 3.13
uv sync --group dev
# Or install with pip
pip install -r dev-requirements.txt
pip install --editable .Run the linter and formatter:
ruff check --fix matchms/YOUR-MODIFIED-FILE.py
ruff format matchms/YOUR-MODIFIED-FILE.pyInstall pre-commit hooks:
pre-commit installRun tests:
pytestWhen adding new functionality, prefer SpectraCollection support from the
start. New filters should ideally support both Spectrum and
SpectraCollection inputs.
For filters with separate spectrum and collection implementations, use
collection_filter:
def _my_filter_spectrum(spectrum_in, ..., clone=True):
...
def _my_filter_collection(collection, ..., clone=True):
...
my_filter = collection_filter(
_my_filter_spectrum,
collection_impl=_my_filter_collection,
)For metadata-only filters, prefer metadata-level implementations and dispatch helpers:
def _my_metadata_filter(metadata, ...):
...
my_filter = metadata_update_filter(_my_metadata_filter)For requirement filters that keep or remove spectra, use a predicate-style metadata implementation:
def _require_my_metadata(metadata, ...) -> bool:
...
require_my_metadata = metadata_requirement_filter(_require_my_metadata)General recommendations:
- Prefer collection-native implementations for new filters.
- Metadata-only filters should operate on
MetadataCollectionor metadata rows. - Peak-only filters should operate on
FragmentCollection. - Filters that drop spectra should return
Nonefor failingSpectruminputs and should drop rows forSpectraCollectioninputs. - Filters should preserve alignment between metadata rows and fragment rows.
- Avoid introducing pandas-specific missing values into reconstructed
Spectrummetadata. - Prefer
ValueErroroverassertfor user-facing validation. - Add tests for both
SpectrumandSpectraCollectionbehavior whenever a public filter supports both input types.
The conda packaging is handled by a recipe at Bioconda.
Publishing to PyPI will trigger the creation of a pull request on the Bioconda recipes repository. Once the pull request is merged, the new version of matchms will appear on Anaconda.
To get support, join the public Slack channel.
If you want to contribute to matchms development, see the contribution guidelines.
Copyright (c) 2026, Düsseldorf University of Applied Sciences
Licensed under the Apache License, Version 2.0. You may not use this file except in compliance with the License. You may obtain a copy of the License at:
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.