A basic BioIO reader plugin for imzML mass spectrometry imaging (MSI) data, read with pyimzML.
pip install bioio-imzmlRequires a sibling .imzML + .ibd file pair (the standard imzML layout).
from bioio import BioImage
img = BioImage("sample.imzML")
img.dims.order # "TCZYX" -- C is the m/z axis
img.channel_names # "<m/z>±<tolerance>" strings, e.g. "798.5400±0.0000"
img.data # (T, C, Z, Y, X) numpy arrayimzML-specific options (mz, mz_step, n_bins, mz_tolerance_absolute,
mz_tolerance_relative, mz_agg, add_tic) work the same way through BioImage,
since it forwards unrecognized keyword arguments straight to the reader.
Pass reader=bioio_imzml.Reader to skip plugin auto-detection (useful when
more than one installed plugin could claim the file):
import bioio_imzml
from bioio import BioImage
# "processed" mode files (one m/z axis per pixel) need target channels:
img = BioImage("sample.imzML", reader=bioio_imzml.Reader, mz=[798.54, 826.57, 885.55])
# reject a target with no real peak nearby instead of returning whatever
# peak happens to be closest, however far away. mz_tolerance_absolute and
# mz_tolerance_relative are both in the same units as mz (m/z) -- relative
# is a plain fraction, not a ppm count, so convert yourself (3 ppm = 3e-6).
# They combine per channel as: tolerance = absolute + m/z * relative
img = BioImage(
"sample.imzML",
reader=bioio_imzml.Reader,
mz=[798.54, 826.57],
mz_tolerance_absolute=0.005,
mz_tolerance_relative=3e-6, # 3 ppm
)
img.reader.mz_tolerance # the resulting per-channel tolerance, e.g. [0.0074, 0.0075]
img.channel_names # ["798.5400±0.0074", "826.5700±0.0075"]
# leave both unset and "processed" mode files get a tolerance for free:
# half the distance to each target's nearest neighboring target, so windows
# never overlap (a lone target with no neighbor is left unbounded).
img = BioImage("sample.imzML", reader=bioio_imzml.Reader, mz=[798.54, 826.57, 885.55])
img.reader.mz_tolerance # e.g. [14.015, 14.015, 29.49] (half the gaps above/below)
# or let the reader pick evenly spaced channels across the file's m/z range,
# either a fixed count (n_bins) or a fixed step (mz_step) in m/z units:
img = BioImage("sample.imzML", reader=bioio_imzml.Reader, n_bins=512)
img = BioImage("sample.imzML", reader=bioio_imzml.Reader, mz_step=0.1)
# mz_agg controls how peaks within a channel's tolerance window combine.
# Default is "sum" -- every measured peak in the window is added up, matching
# how tools like Lipostar/MetaboScape aggregate signal in a window. Pass
# "nearest" instead to take only the single closest measured peak per
# channel (dropping the rest):
img = BioImage(
"sample.imzML", reader=bioio_imzml.Reader, mz=[798.54, 826.57], mz_agg="nearest"
)
# add_tic appends one extra channel named "TIC" (Total Ion Count): each pixel's
# value is the sum of every peak in that pixel's full raw spectrum -- computed
# before channel extraction, so it includes signal outside the target m/z grid
# and signal dropped by tolerance windows (not the same as summing the
# extracted channels). Works in both continuous and processed mode.
img = BioImage(
"sample.imzML", reader=bioio_imzml.Reader, mz=[798.54, 826.57], add_tic=True
)
img.reader.channel_names # [..., "TIC"] -- one more entry than mz_values
tic = img.data[0, -1, 0] # (Y, X) Total Ion Count map (last channel)mz_values and mz_tolerance keep describing only the m/z channels, so with
add_tic=True the TIC channel is the extra trailing one and channel_names
has one more entry than mz_values.
Don't know which m/z channels a file actually has signal at? auto_pick_peaks
finds candidate peaks on the file's mean spectrum, and can then drop candidates
that are too rare across pixels or spatially unstructured (noise/matrix
artifacts rather than real signal). Both quality filters are opt-in: with
min_pixel_frequency and max_spatial_chaos both left at None the slow
per-pixel pass is skipped entirely and result.pixel_frequency /
result.spatial_chaos come back as NaN.
import bioio_imzml
from bioio import BioImage
# pin a tolerance and reuse it for picking and extraction, so extraction
# matches what pixel_frequency/spatial_chaos actually scored. Size it to
# bin_width, not to ppm mass-accuracy precision: candidates come from a
# bin_width-binned mean spectrum, so a candidate's reported m/z can be off
# from the true peak by up to ~bin_width/2.
bin_width = 0.05
tol_abs = bin_width
result = bioio_imzml.auto_pick_peaks(
"sample.imzML",
min_mz=650,
max_mz=850,
bin_width=bin_width,
mz_tolerance_absolute=tol_abs,
)
result.mzs # candidate m/z values, sorted by descending intensity
result.pixel_frequency # fraction of pixels with signal, one per mz
result.spatial_chaos # 0 (structured) .. 1 (spatially random), one per mz
# LipostarMSI-equivalent culling -- pass these explicitly to turn the
# per-pixel quality filters on (`max_spatial_chaos=0.4` is LipostarMSI's
# "min spatial chaos 0.6" under this module's opposite convention):
result = bioio_imzml.auto_pick_peaks(
"sample.imzML",
min_mz=650,
max_mz=850,
bin_width=bin_width,
mz_tolerance_absolute=tol_abs,
min_pixel_frequency=0.01,
max_spatial_chaos=0.4,
)
if len(result.mzs) == 0:
# those thresholds can reject every candidate on data with sparse
# per-pixel peak-picking (e.g. single-cell-resolution processed-mode
# files) -- loosen or disable a filter rather than pass an empty mz
# list on to Reader/BioImage:
result = bioio_imzml.auto_pick_peaks(
"sample.imzML",
min_mz=650,
max_mz=850,
bin_width=bin_width,
mz_tolerance_absolute=tol_abs,
min_pixel_frequency=0.01,
max_spatial_chaos=None,
)
img = BioImage(
"sample.imzML",
reader=bioio_imzml.Reader,
mz=result.mzs,
mz_tolerance_absolute=tol_abs,
)Tune detection sensitivity (snr_threshold, min_relative_intensity) and the
quality filters (min_pixel_frequency, max_spatial_chaos) as keyword
arguments; see the docstring for defaults. Three independent m/z windows govern
picking, and physically they should satisfy
bin_width <= mz_tolerance <= 0.5 * min_separation:
bin_width-- the mean-spectrum grid step (detection resolution).mz_tolerance_absolute/mz_tolerance_relative-- the extraction/scoring window half-width±tol(mass accuracy).min_separation_absolute/min_separation_relative-- the minimum gap between two accepted candidates (instrument resolving power). Combined asabsolute + m/z * relative, so it can grow with m/z. Defaults to2 * bin_widthwhen both are unset, so dedup is always enforced; pass0for both to disable it.
Separation is deliberately decoupled from mz_tolerance: setting it equal to
the half-width tol would let two accepted peaks sit tol apart with
50%-overlapping extraction windows and double-count intensity under
mz_agg="sum". auto_pick_peaks emits a UserWarning (it never raises) when
these windows are set inconsistently.
snr_threshold, min_pixel_frequency, and max_spatial_chaos each accept
None to disable that filter -- passing None for both quality filters
also skips the per-pixel pass over the file entirely (the slow part),
leaving result.pixel_frequency/result.spatial_chaos as NaN.
bioio_imzml.peak_picking also exposes the individual steps --
mean_spectrum, find_peaks_in_spectrum, and
pixel_frequency_and_spatial_chaos -- to inspect intermediate results or why
a candidate was dropped before committing to thresholds.
| Parameter | Default | Description |
|---|---|---|
image |
(required) | Path to the imzML file. |
min_mz |
None |
Lower bound of the m/z range to scan (whole range if None). |
max_mz |
None |
Upper bound of the m/z range to scan (whole range if None). |
bin_width |
0.05 |
Bin width (m/z) of the mean spectrum candidates are detected on. |
smooth |
True |
Apply Savitzky-Golay smoothing before detection (detection only; not applied to the returned raw spectrum). |
savgol_window |
7 |
Savitzky-Golay window length; widen to suppress jagged/spurious candidates. |
savgol_polyorder |
2 |
Savitzky-Golay polynomial order. |
snr_threshold |
None |
Minimum signal-to-noise ratio; None disables the SNR filter. |
min_relative_intensity |
0.0 |
Minimum intensity relative to the tallest peak. |
min_pixel_frequency |
None |
Minimum fraction of pixels with signal; None (default) disables the filter. LipostarMSI-equivalent: 0.01. |
max_spatial_chaos |
None |
Maximum spatial chaos (0 structured .. 1 random); None (default) disables the filter. LipostarMSI-equivalent: 0.4 (its "min spatial chaos 0.6", opposite convention). With both quality filters None the slow per-pixel pass is skipped and both scores come back NaN. |
top_n_peaks |
None |
Cap on channels returned after filtering (all if None). |
mz_tolerance_absolute |
None |
Absolute extraction/scoring window half-width (m/z), used to build each channel and score per-pixel frequency/chaos. |
mz_tolerance_relative |
None |
Relative component of the same window (fraction), combining as absolute + m/z * relative. |
min_separation_absolute |
None |
Absolute minimum gap between accepted candidates (m/z, resolving power). Both separation components None defaults to 2 * bin_width; pass 0 for both to disable dedup. |
min_separation_relative |
None |
Relative component of the separation gap (fraction), combining as absolute + m/z * relative; makes the gap scale with m/z. |
max_rows |
None |
Image rows in flight during the per-pixel pass, its peak-memory bound. None (default) fits it to the memory free right now; pass a number to cap it yourself. It never adds a pass over the file. |
fs_kwargs |
{} |
Extra kwargs forwarded to the underlying file reader. |
The per-pixel pass (min_pixel_frequency/max_spatial_chaos) streams one
image row at a time, and max_rows decides how many run concurrently:
row_bytes = n_candidates * width * 8
peak_RAM ~ 6 * max_rows * row_bytes + ~2GB
It defaults to None, which fits the row count to the memory free right now,
capped at the machine's usable core count and at 8 -- so normally there is
nothing to tune. Pass a number to override, as a hard cap on a shared machine
or to deliberately trade speed for memory:
result = bioio_imzml.auto_pick_peaks(
"sample.imzML",
bin_width=bin_width,
mz_tolerance_absolute=tol_abs,
min_pixel_frequency=0.01,
max_spatial_chaos=0.4,
max_rows=2, # omit the argument to size automatically
)Two things worth knowing about the shape of the trade:
- Speed saturates at the physical core count. Rows in flight beyond it cost
RAM and return nothing -- on an SMT machine the extra hardware threads only
contend for memory bandwidth. Concurrency is therefore capped at the physical
core count (via
psutilwhen it is installed, logical cores otherwise) and at 8, which is as far as this has been measured. - Memory scales with the row count, speed does not. Going from one row in flight to the core count is close to a linear speedup; past that the curve is flat while RAM keeps climbing.
In a memory-constrained runtime (a WASM notebook, say) the automatic sizing degrades to a single row when it cannot read how much memory is free -- slow, but bounded.
| Attribute | Description |
|---|---|
mzs |
Candidate m/z values, sorted by descending mean-spectrum intensity. |
pixel_frequency |
Fraction of pixels with signal, one per mzs (NaN if both quality filters disabled). |
spatial_chaos |
Spatial chaos 0 (structured) .. 1 (random), one per mzs (NaN if both quality filters disabled). |
mean_spectrum_mz |
m/z axis of the full (raw) mean spectrum candidates were detected from. |
mean_spectrum_intensity |
Raw intensities of that mean spectrum. |
imzML stores spectra in one of two ways:
- continuous: every pixel shares one m/z axis, so intensities already line up across pixels. Detected automatically (identical m/z byte offset and length for every spectrum) and read directly -- no resampling, no channel arguments needed.
- processed: each pixel has its own m/z axis (typical for high-resolution
profile data). There's no single true channel set, so this reader resamples
every spectrum onto shared target m/z values, given via
mz=or auto-generated withn_bins=, summing peaks within each channel's tolerance window by default (mz_agg="sum";mz_agg="nearest"takes the single closest peak instead).
reader.is_continuous reports which case applies to a given file.
uv sync
uv run pytest
uv run ruff check .
uv run ruff format .
uv run ty checkBump the version (updates pyproject.toml) and tag a release to publish to
PyPI via CI:
uv version --bump patch # or minor / major
git commit -am "Bump version"
git tag "v$(uv version --short)"
git push --tags