Skip to content

Readers

spxtacular provides six format-specific reader classes — MzmlReader, DReader, ThermoReader, MgfReader, Ms2Reader, and MspReader — plus a format-agnostic Reader that picks the right one from the file extension. All of them expose a uniform interface for iterating over MsnSpectrum objects regardless of the underlying file format.

Every reader yields MsnSpectrum instances populated with as much instrument metadata as the format provides. All spectrum-processing methods (.filter(), .denoise(), .deconvolute(), etc.) are immediately available on each yielded object.

MzmlReader, DReader, and ThermoReader need an optional extra (mzmlpy / tdfpy / fisher-py); MgfReader, Ms2Reader, and MspReader are pure standard library and always available, as are the matching writers write_mgf, write_ms2, and write_msp.

Lookup objects

.ms1 and .ms2 are not generators. Each property returns a small lookup object that is both iterable (for spec in reader.ms1:) and indexable (reader.ms2[42]):

Reader .ms1 type .ms2 type
MzmlReader MzmlSpectraLookup MzmlSpectraLookup
DReader DReaderMs1Lookup DReaderMs2Lookup
ThermoReader ThermoScanLookup ThermoScanLookup
MgfReader / Ms2Reader / MspReader PeakListLookup (always empty) PeakListLookup
Reader whichever the detected backend provides whichever the detected backend provides

Because they are iterables and not iterators, next(reader.ms2) raises TypeError — use next(iter(reader.ms2)) to pull the first spectrum:

first_ms2 = next(iter(reader.ms2))

Index semantics differ per backend:

Expression Meaning
MzmlReader.ms1[key] / .ms2[key] Spectrum by overall 0-based index or native ID string. No MS-level filtering is applied on random access, so reader.ms2[0] is the first spectrum in the file, not the first MS2 spectrum
MzmlReader[key] Same as above, straight off the reader (reader[0], reader["scan=19"])
DReader.ms1[frame_id] MS1 spectrum by tdfpy frame_id
DReader.ms2[precursor_id] MS2 spectrum by tdfpy precursor_id — DDA only. DIA and PRM raise NotImplementedError (their MS2 records are not keyed by a single id); use get_by_native_id or iterate
ThermoReader.ms1[scan] / .ms2[scan] Spectrum by native 1-based scan number; KeyError if the scan does not exist or is not of that MS level. ThermoReader[scan] fetches any level
MgfReader[key] / Ms2Reader[key] / MspReader[key] Spectrum by 0-based position in the file, or by native_id string (first match wins — library files repeat names across collision energies). Each lookup streams the file from the start, so it is O(n) — iterate when you want them all

DReader and ThermoReader lookups raise SpxtacularError if the reader has not been opened.

Lookup by scan number or native id

Every reader, and Reader, has the same three methods for fetching one spectrum. They follow the scan-number rule: a spectrum has a scan_number only when that number identifies it alone, and a lookup that cannot name exactly one spectrum raises instead of guessing.

reader.get_by_scan(scan_number: int, *, ms_level: int | None = None) -> MsnSpectrum
reader.get_by_native_id(native_id: str) -> MsnSpectrum
reader.get_by_sage_scannr(scannr: str | int) -> MsnSpectrum    # DReader: also precursor_offset=1
import pandas as pd
from spxtacular import Reader

psms = pd.read_csv("results.sage.tsv", sep="\t")
with Reader("run.mzML") as reader:
    spectra = [reader.get_by_sage_scannr(s) for s in psms.scannr]
    ms2 = reader.get_by_scan(20, ms_level=2)
    same = reader.get_by_native_id("controllerType=0 controllerNumber=1 scan=20")

Errors. A key that is well formed but not in the file raises KeyError. A key that cannot identify one spectrum raises SpxtacularError:

Situation Raises
No spectrum has this scan number or native id KeyError
The spectrum exists but is not of ms_level KeyError
get_by_native_id with an id of a form the reader never produces (Thermo, Bruker .d) KeyError
The file carries no scan numbers at all: MSP, MGF without SCANS, mzML with Bruker frame=…, Waters function=… or SCIEX cycle=… ids SpxtacularError, pointing to get_by_native_id
Several spectra share the scan number or native id SpxtacularError
Bruker DDA: the number is both an MS1 frame id and a precursor id, and ms_level is None SpxtacularError; pass ms_level
Bruker DIA / PRM: the frame holds several MS2 spectra (one per window or target) SpxtacularError; use the native id
Bruker .d not opened, or a non-integer scan number, ms_level or scannr SpxtacularError

Per format:

Reader get_by_scan key get_by_native_id key Cost
MzmlReader The number in a scan=, Thermo, index= or spectrum= id (not the 0-based position: use reader[i]) spectrum/@id First call lists the file's ids once; then mzmlpy's random access
ThermoReader 1-based scan number (same as reader[n]) controllerType=0 controllerNumber=1 scan=N, or scan=N Direct
DReader MS1 frame_id, DDA precursor_id, DIA / PRM MS2 frame_id frame=F (MS1), precursor=P (DDA), F@wI (DIA window), F@tT (PRM target) Direct; DIA / PRM group their windows once per open(). Needs an open reader (with DReader(...)), like the rest of DReader; unopened it raises SpxtacularError
MgfReader SCANS TITLE First call parses the whole file once and keeps the byte offset of each record; later calls seek to it. The index is rebuilt when the file's size or mtime changes. A gzipped file still decompresses up to the record
Ms2Reader S line I NativeID, else scan=N As MGF
MspReader always SpxtacularError Name As MGF

Sage scannr. Sage writes a different value per input format, and get_by_sage_scannr reads each:

  • mzML: the native id (results.sage.tsv), or only the number from scan=N (.pin).
  • MGF: the TITLE (results.sage.tsv), or the number from scan=N in it (.pin).
  • Bruker .d: a bare integer, timsrust's 0-based spectrum index. Upstream Sage (checked against Sage 0.15.0-beta.1, built on timsrust 0.4.2) counts DDA precursors from 0, so scannr N is precursor N + 1: the default precursor_offset=1. Sage builds on timsrust 0.6 or later write the precursor id itself; pass precursor_offset=0. DIA and PRM raise SpxtacularError: Sage numbers them by timsrust's expanded window list, which spxtacular does not reproduce.

Everywhere but .d, the value is tried as a native id and, if it is a bare integer, as a scan number too. If both match and name different spectra (an MGF with TITLE=12 on one spectrum and SCANS=12 on another), it raises SpxtacularError rather than guess. A .pin scannr from a file without scan numbers (Bruker mzML) therefore raises SpxtacularError; use results.sage.tsv, which keeps the full id.

The ms1 / ms2 indexing below is unchanged and still works.

polarity, activation_type, im_type, and analyzer are populated as plain strings straight from the underlying format (including raw PSI-MS accessions such as "MS:1002481") — activation_type, im_type and analyzer also accept the ActivationType, IMType, and Analyzer enums documented in API reference — Metadata enums if you want to set or compare them with autocomplete/typo-safety.


Reader

Reader is the format-agnostic entry point: it inspects the path suffix and delegates to DReader (.d), MzmlReader (.mzml), ThermoReader (.raw), MgfReader (.mgf), Ms2Reader (.ms2), or MspReader (.msp) — case-insensitive, and a trailing .gz is stripped before matching so .mzML.gz, .mgf.gz, .ms2.gz, and .msp.gz all dispatch correctly. Anything else raises ValueError. Usage is identical regardless of the underlying format.

class Reader:
    def __init__(
        self,
        path: str | Path,
        *,
        centroid_config: CentroidConfig | None = None,
        mzml_gzip_mode: Literal["auto", "indexed", "stream"] = "auto",
        mzml_in_memory: bool = False,
    ): ...

    def open(self) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> Reader: ...
    def __exit__(self, exc_type, exc_val, exc_tb) -> None: ...

    @property
    def ms1(self) -> DReaderMs1Lookup | MzmlSpectraLookup: ...

    @property
    def ms2(self) -> DReaderMs2Lookup | MzmlSpectraLookup: ...

    @property
    def access_strategy(self) -> str | None: ...

    def get_by_scan(self, scan_number: int, *, ms_level: int | None = None) -> MsnSpectrum: ...
    def get_by_native_id(self, native_id: str) -> MsnSpectrum: ...
    def get_by_sage_scannr(self, scannr: str | int, *, precursor_offset: int | None = None) -> MsnSpectrum: ...
from spxtacular import Reader

with Reader("run.mzML") as r:      # or Reader("/data/sample.d")
    for spec in r.ms1:
        print(spec)

with Reader("/data/sample.d") as r:
    ms2 = r.ms2[42]                # DDA precursor_id

centroid_config is only meaningful for .d inputs. The mzml_* options are forwarded only to MzmlReader. Reader exposes ms1, ms2, access_strategy, open, close, and the lookup methods; precursor_offset is for .d only and raises SpxtacularError for other formats. access_strategy is None for non-mzML inputs.


MzmlReader

Reads standard .mzML files using mzmlpy. No context manager is required, but one is strongly recommended — see File handles below.

class MzmlReader:
    def __init__(
        self,
        mzml_path: str | Path,
        *,
        gzip_mode: Literal["auto", "indexed", "stream"] = "auto",
        in_memory: bool = False,
    ): ...

    def open(self) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> MzmlReader: ...
    def __exit__(self, exc_type, exc_val, exc_tb) -> None: ...

    @property
    def ms1(self) -> MzmlSpectraLookup: ...

    @property
    def ms2(self) -> MzmlSpectraLookup: ...

    @property
    def access_strategy(self) -> str | None: ...

    def __getitem__(self, key: int | str) -> MsnSpectrum: ...

For gzipped mzML, the default gzip_mode="auto" uses a self-indexed gzip file when there is one, else reads the compressed file in place through rapidgzip (reusing complete sidecars, never writing new ones), else decompresses into memory. It never writes next to your file. The selected route is reported by access_strategy. Use gzip_mode="stream" when a service intentionally reads sequentially:

with MzmlReader("large-run.mzML.gz", gzip_mode="stream", in_memory=False) as r:
    first_ms1 = next(iter(r.ms1))

The "indexed" mode builds a random-access gzip index and requires rapidgzip. Streaming avoids an index, but later index lookups must scan forward from the start.

Create the self-indexed format through the thin mzMLPy-backed helper:

from spxtacular import write_indexed_mzml_gzip

write_indexed_mzml_gzip("run.mzML", "run.indexed.mzML.gz")

Properties

Property Contents
ms1 All MS1 spectra in scan order (iteration); index/native-ID access is unfiltered
ms2 All MS2 spectra in scan order, including parsed precursor information (iteration); index/native-ID access is unfiltered
access_strategy Concrete mzMLPy route such as embedded, rapidgzip, memory, stream, or plain

Metadata populated from mzML

Field Source
scan_number The number in the native id when it identifies the spectrum on its own: scan=19 or Thermo controllerType=0 controllerNumber=1 scan=19 -> 19; index=5 / spectrum=5 -> 5. None for every other id (Bruker frame=… scan=…, Waters function=… scan=…, SCIEX cycle=…), whose scan value repeats or is missing; native_id keeps the full id
ms_level msLevel CV param
native_id Raw spectrum id attribute
rt scan start time (converted to seconds)
mz_range scan window lower/upper limit
polarity positive scan / negative scan CV params
spectrum_type centroid spectrum / profile spectrum CV params
charge (array) Per-peak charge array when present in the binary data
im (array) First ion mobility binary array when present
precursors MS2 only: selected ion m/z, intensity, charge, and activation info
collision_energy MS2 only: from activation element
activation_type MS2 only: from activation element

Examples

Iterate MS1 spectra:

from spxtacular import MzmlReader

reader = MzmlReader("run.mzML")
for spec in reader.ms1:
    print(spec)
    # MsnSpectrum(scan=0, ms_level=1, rt=1.23s, polarity=positive, n_peaks=4521)

Filter and denoise each MS1 scan:

from spxtacular import MzmlReader

reader = MzmlReader("run.mzML")
for spec in reader.ms1:
    processed = spec.filter(min_mz=200, max_mz=1600).denoise(method="mad")
    print(f"Scan {spec.scan_number}: {len(processed)} peaks after denoise")

Iterate MS2 spectra with precursor info:

from spxtacular import MzmlReader

reader = MzmlReader("run.mzML")
for spec in reader.ms2:
    if not spec.precursors:
        continue
    prec = spec.precursors[0]
    print(
        f"Scan {spec.scan_number} | "
        f"Precursor {prec.precursor_mz:.4f} m/z, z={prec.charge} | "
        f"CE={spec.collision_energy} eV"
    )

Full deconvolution pipeline on MS1:

from spxtacular import MzmlReader

reader = MzmlReader("run.mzML")
for spec in reader.ms1:
    neutral = (
        spec
        .filter(min_mz=300, min_intensity=1000)
        .denoise(method="mad")
        .deconvolute(charge_range=(1, 5), tolerance=10, tolerance_unit="ppm")
        .decharge()
    )
    for peak in neutral.top_peaks(10):
        print(f"  mass={peak.mz:.4f} Da  intensity={peak.intensity:.2e}")
    break  # first scan only

Fetch a single spectrum by index or native ID:

from spxtacular import MzmlReader

with MzmlReader("run.mzML") as reader:
    first = reader[0]              # by overall 0-based index
    named = reader["scan=19"]      # by native ID
    first_ms2 = next(iter(reader.ms2))   # first MS2 spectrum

File handles

MzmlReader.open() opens a persistent mzmlpy handle and close() releases it; __enter__ / __exit__ call them, so __exit__ is not a no-op. While a handle is open, iteration and index access reuse it (the fast path). Without one, every .ms1 / .ms2 / reader[...] operation opens and closes the file again — correct, but considerably slower. Prefer the context manager:

with MzmlReader("run.mzML") as reader:
    for spec in reader.ms1:
        ...

DReader

Reads Bruker timsTOF .d directories using tdfpy. Must be opened before use — either with open() / close() or (preferred) as a context manager. The underlying tdfpy handle is opened on __enter__ and closed on __exit__; touching .ms1 / .ms2 before open() raises SpxtacularError.

class DReader:
    def __init__(
        self,
        analysis_dir: str | Path,
        centroid_config: CentroidConfig | None = None,
    ): ...

    def open(self) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> DReader: ...
    def __exit__(self, exc_type, exc_val, exc_tb) -> None: ...

    @property
    def ms1(self) -> DReaderMs1Lookup: ...

    @property
    def ms2(self) -> DReaderMs2Lookup: ...

The acquisition type (DDA, DIA, PRM) is detected automatically from the .d directory at construction time and stored as reader.acquisition_type (AcquisitionType enum).

CentroidConfig

timsTOF frames arrive as raw (frame, scan) points, so DReader centroids them via tdfpy. CentroidConfig holds the parameters forwarded to that step for MS1, DIA and PRM spectra. DDA MS2 spectra use tdfpy's per-precursor merged peaks and ignore it, as does MzmlReader. All fields are keyword-only.

from spxtacular import CentroidConfig, DReader

cfg = CentroidConfig(mz_tolerance=8.0, min_peaks=3, noise_filter="mad")
with DReader("/data/sample.d", centroid_config=cfg) as reader:
    ...
Field Default Description
mz_tolerance 8.0 m/z merge window
mz_tolerance_unit "ppm" "ppm" or "da"
im_tolerance 0.1 Ion mobility merge window
im_tolerance_unit "relative" "relative" or "absolute"
min_peaks 3 Minimum raw points required to emit a centroid
noise_filter None "mad", "percentile", "histogram", "baseline", "iterative_median", a float threshold, or None

Passing centroid_config=None (the default) uses CentroidConfig() with the values above.

Properties

Property Contents
ms1 All MS1 frames, centroided and merged by tdfpy; ms1[frame_id] fetches one frame
ms2 All MS2 spectra (DDA: per-precursor; DIA: per isolation window; PRM: per transition); ms2[precursor_id] fetches one DDA spectrum

Metadata populated from timsTOF

MS1:

Field Source
scan_number frame_id
native_id "frame={frame_id}"
ms_level Always 1
rt Frame acquisition time (seconds)
injection_time Frame accumulation time (ms)
total_ion_current Frame summed intensity
mz_range Instrument acquisition range from metadata
im_range 1/K0 acquisition range from metadata
im (array) Per-peak 1/K0 values
analyzer Always "TOF"
ramp_time Frame ramp time (ms)
polarity From frame polarity field
spectrum_type Always CENTROID (timsTOF data arrives centroided)

MS2 (DDA):

Field Source
scan_number precursor_id
native_id "precursor={precursor_id}"
ms_level Always 2
rt Retention time (seconds)
isolation_mz_range Precursor isolation window
isolation_ook0_range 1/K0 range of precursor
precursors Single Precursor with monoisotopic m/z (or largest peak m/z if unavailable), intensity, charge, and 1/K0 (im_type="ook0")
collision_energy From precursor record
activation_type "MS:1002481" (PASEF)

MS2 (DIA):

Field Source
scan_number frame_id
native_id "{frame_id}@w{window_index}"
ms_level Always 2
rt Retention time (seconds)
isolation_mz_range Isolation window m/z range
isolation_ook0_range Isolation window 1/K0 range
im (array) Per-peak 1/K0 values
collision_energy From window record
precursors None — DIA windows have no defined precursor

MS2 (PRM):

Field Source
scan_number frame_id
native_id "{frame_id}@t{target_id}"
ms_level Always 2
rt Retention time (seconds)
isolation_mz_range / isolation_ook0_range Transition isolation windows
precursors Single Precursor built from the PRM target (monoisotopic m/z, charge, 1/K0); intensity is the summed MS2 peak intensity, since PRM targets carry no measured precursor intensity
collision_energy From transition record

Examples

Iterate MS1 frames (DDA or DIA):

from spxtacular import DReader

with DReader("/data/sample.d") as reader:
    print(f"Acquisition type: {reader.acquisition_type}")
    for spec in reader.ms1:
        print(spec)
        # MsnSpectrum(scan=1, ms_level=1, rt=0.42s, polarity=positive, n_peaks=8234)
        break

MS1 with ion mobility filtering:

from spxtacular import DReader

with DReader("/data/sample.d") as reader:
    for spec in reader.ms1:
        # Keep only peaks in a specific 1/K0 window
        filtered = spec.filter(min_im=0.8, max_im=1.2, min_intensity=500)
        if len(filtered) == 0:
            continue
        neutral = (
            filtered
            .deconvolute(charge_range=(1, 5), tolerance=15, tolerance_unit="ppm")
            .decharge()
        )
        print(f"Frame {spec.scan_number}: {len(neutral)} neutral masses")

MS2 DDA — inspect precursors:

from spxtacular import DReader

with DReader("/data/sample_dda.d") as reader:
    for spec in reader.ms2:
        if not spec.precursors:
            continue
        prec = spec.precursors[0]
        print(
            f"Precursor {spec.scan_number}: "
            f"m/z={prec.precursor_mz:.4f}, z={prec.charge}, "
            f"1/K0={prec.im:.3f}, "
            f"monoisotopic={prec.is_monoisotopic}"
        )
        break

MS2 DIA — iterate isolation windows:

from spxtacular import DReader

with DReader("/data/sample_dia.d") as reader:
    for spec in reader.ms2:
        print(
            f"{spec.native_id}: "
            f"mz_range={spec.mz_range}, "
            f"CE={spec.collision_energy}"
        )
        break

ThermoReader

Reads Thermo .raw files using fisher-py, which wraps Thermo's official RawFileReader .NET assemblies. Must be opened before use — either with open() / close() or (preferred) as a context manager; touching .ms1 / .ms2 before open() raises SpxtacularError.

pip install spxtacular[thermo]

Beyond the Python extra, .raw reading needs a .NET runtime (8 or later) on the machine — install it from dotnet.microsoft.com/download and make sure dotnet is on PATH (or DOTNET_ROOT points at it). Without one, constructing a ThermoReader raises an ImportError explaining exactly that; import spxtacular itself never touches the runtime. A .raw directory is the Waters format, not Thermo's, and is rejected with a ValueError suggesting mzML conversion.

class ThermoReader:
    def __init__(
        self,
        raw_path: str | Path,
        prefer_vendor_centroid: bool = True,
    ): ...

    def open(self) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> ThermoReader: ...
    def __exit__(self, exc_type, exc_val, exc_tb) -> None: ...

    @property
    def ms1(self) -> ThermoScanLookup: ...

    @property
    def ms2(self) -> ThermoScanLookup: ...

    def __getitem__(self, scan_number: int) -> MsnSpectrum: ...

Vendor centroids vs. profile

Thermo FTMS scans acquired in profile mode also carry the instrument's own centroid ("label") stream, complete with per-peak charge annotations. By default (prefer_vendor_centroid=True) ThermoReader yields those centroids as a CENTROID spectrum with a charge array — Thermo's "charge unknown" (0) is remapped to spxtacular's unassigned marker (-1), since 0 is reserved for decharged spectra. Pass prefer_vendor_centroid=False to get scans exactly as acquired: the full PROFILE trace for profile-mode scans (centroid it yourself with .centroid() or deconvolute the vendor centroids instead). Scans acquired in centroid mode — typical for ion-trap detectors — are unaffected by the flag and always come back as CENTROID without charges.

Metadata populated from .raw

Field Source
scan_number Native 1-based scan number
native_id "controllerType=0 controllerNumber=1 scan=<n>" (mzML-compatible)
ms_level Scan filter MS order
rt Scan start time, converted from RawFileReader's minutes to seconds
mz_range Scan window from the scan header
polarity Scan filter polarity
analyzer Scan filter mass analyzer; FTMS resolves to orbitrap (or ft_icr on LTQ FT instruments), ITMS to ion_trap, etc.
injection_time Trailer Ion Injection Time (ms)
resolution Trailer Orbitrap Resolution / FT Resolution
total_ion_current Scan header TIC
charge (array) Vendor centroid stream charge annotations (default mode only); 0 → -1
precursors MS2+: isolation target m/z (trailer Monoisotopic M/Z when set, flagged is_monoisotopic), trailer Charge State; intensity is the summed product-ion intensity, since the scan records no precursor intensity
isolation_mz_range Reaction isolation width (+ offset) around the target m/z
collision_energy Reaction collision energy
activation_type Reaction activation, mapped to ActivationType (CID, HCD, ETD, …); ETD plus a supplemental-activation reaction becomes EThcD / ETciD; unrecognised vendor modes pass through as raw strings

Examples

from spxtacular import ThermoReader

with ThermoReader("run.raw") as reader:
    for spec in reader.ms2:
        prec = spec.precursors[0]
        print(
            f"Scan {spec.scan_number} | {spec.activation_type} "
            f"{spec.collision_energy} eV | precursor {prec.precursor_mz:.4f} z={prec.charge}"
        )

    scan_42 = reader[42]           # any MS level, by native scan number
    ms1_scan = reader.ms1[41]      # KeyError if scan 41 is not MS1
# The data exactly as acquired — profile scans stay profile:
with ThermoReader("run.raw", prefer_vendor_centroid=False) as reader:
    profile = reader.ms2[2]        # SpectrumType.PROFILE
    centroided = profile.centroid()

MGF / MS2 / MSP

MGF (Mascot Generic Format), MS2, and MSP (the NIST spectral-library format) are plain-text peak lists. All three are handled by spxtacular.peaklist, which is pure standard library — no optional extra, nothing to install, always importable. All three formats hold fragmentation spectra only, so every spectrum read back is an MsnSpectrum with ms_level=2 and spectrum_type=SpectrumType.CENTROID, and reader.ms1 is a valid but always empty walk.

class MgfReader:            # and Ms2Reader / MspReader — identical interface
    def __init__(self, path: str | Path): ...

    def open(self) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> Self: ...
    def __exit__(self, exc_type, exc_val, exc_tb) -> None: ...

    def __iter__(self) -> Iterator[MsnSpectrum]: ...
    def __len__(self) -> int: ...
    def __getitem__(self, key: int | str) -> MsnSpectrum: ...

    @property
    def ms1(self) -> PeakListLookup: ...   # always empty

    @property
    def ms2(self) -> PeakListLookup: ...
from spxtacular import MgfReader, Ms2Reader, write_mgf, write_ms2

with MgfReader("run.mgf") as reader:        # run.mgf.gz works too
    print(len(reader))                      # spectra in the file
    for spec in reader:                     # or: for spec in reader.ms2
        prec = spec.precursors[0]
        print(f"{spec.scan_number}: {prec.precursor_mz:.4f} z={prec.charge} rt={spec.rt}")

write_ms2(MgfReader("run.mgf"), "run.ms2")  # readers are iterables of spectra

Gzip is handled transparently on read — detected by magic bytes, so a compressed file works under any name — and on write, where a .gz suffix compresses the output.

File handles

Unlike MzmlReader, these readers hold no handle between walks: every iteration streams a fresh one, so iterations are independent and may be nested. open() only checks the file exists (raising FileNotFoundError early) and close() is a no-op; the context manager is supported for symmetry with the other readers, not because anything needs releasing.

len(reader) makes one pass that counts record-start lines without parsing peaks, then caches the result.

Metadata populated from MGF

Field Source
mz / intensity Ion lines inside BEGIN IONS … END IONS
charge (array) Optional third column on the ion lines — kept only when every peak has one
scan_number SCANS (a range such as 1024-1030 collapses to its first scan)
ms_level Always 2
native_id TITLE, verbatim
rt RTINSECONDS, or the non-standard RTINMINUTES × 60 when that is all the file has
polarity Implied by the sign of CHARGE (2+ → positive, 2- → negative)
precursors One Precursor from PEPMASS (m/z, optional intensity) and CHARGE
spectrum_type Always CENTROID

Metadata populated from MS2

Field Source
mz / intensity Ion lines following an S record
scan_number First scan field of the S line
ms_level Always 2
native_id I NativeID when present (spxtacular writes it for ids that are not scan=<n>), else synthesised as "scan=<scan_number>"
rt I RTime / I RetTime × 60 — those values are minutes in the wild, rt is seconds
injection_time I IonInjectionTime
total_ion_current I TIC
activation_type I ActivationType
polarity Implied by the sign of the Z charge
precursors One Precursor: m/z from the S line, charge from the first Z line, intensity from I PrecursorInt
spectrum_type Always CENTROID

H header lines and D analysis lines are read and skipped. Unmapped I keys are skipped too.

Metadata populated from MSP

MSP has no formal spec; MspReader handles both dialects found in the wild — NIST/SpectraST peptide libraries and metabolomics exports (MoNA, GNPS, MS-DIAL). Header keys are matched case-insensitively, ignoring spaces, underscores, and hyphens, so PrecursorMZ, PRECURSORMZ, and Precursor_mz are all the same key. A record is header lines up to Num Peaks: N followed by exactly N peaks — the count is both metadata and the record terminator.

Field Source (first hit wins)
mz / intensity Peak lines after Num Peaks; several peaks may share a line, semicolon-separated; a trailing annotation column ("b2/0.001") is ignored
native_id Name, verbatim
ms_level Always 2
precursors One Precursor from PrecursorMZ, falling back to Parent= inside Comment
precursor charge Charge header, Charge= comment pair, or a trailing /2 on a peptide Name
polarity Ion_mode / Polarity (P/N/Positive/Negative), else implied by the charge sign
rt RetentionTime / RT header or RT= comment pair — verbatim, no unit conversion: MSP has no unit convention (NIST peptide libraries write seconds, most metabolomics exporters minutes), so no guess is made. Know your library
collision_energy Collision_energy header or CE= comment pair; the first number is used, so 35, 35 eV, and HCD 35% all parse
spectrum_type Always CENTROID

Comment: is parsed as space-separated Key=value pairs (values optionally quoted) and the known keys above are extracted. Everything else — Formula:, SMILES:, InChIKey:, Precursor_type: adducts, Synon:, Mods= — is skipped: MsnSpectrum has no fields for compound identity, and per-peak annotation strings have no per-peak storage. Count mismatches in either direction, a record with no Num Peaks line, and unparsable numbers raise ValueError naming the file and line, in keeping with the leniency table above.

Leniency, and where it stops

Real peak lists are written by a long tail of tools, so parsing is deliberately forgiving:

Input Behaviour
Unknown KEY=VALUE MGF headers Ignored
Headers outside any BEGIN IONS block Ignored (global MGF headers such as SEARCH=MIS)
Blank lines, and comments opening with #, ;, !, or / Skipped anywhere in the file
CHARGE=2+ and 3+, CHARGE=2+,3+ First state only — Precursor.charge is a single value
Repeated MS2 Z lines All parsed, first used — same reason
SCANS=1024-1030, RTINSECONDS=120-130 First number used
Comma-separated MGF ion lines Accepted alongside whitespace
A block with no ion lines Yields an empty spectrum (length-0 arrays), not an error
Undecodable bytes Replaced, so one mangled TITLE cannot make a file unreadable

Structural damage is not tolerated, and every such error names the file and the line:

ValueError: run.mgf:412: 'END IONS' without a matching 'BEGIN IONS'
ValueError: run.mgf:87: expected 'mz intensity' on an ion line, got '843.4102'
ValueError: run.ms2:1: '100.0 5.0' appears before any 'S' scan line

Unterminated blocks, nested BEGIN IONS, short S/Z lines, and unparsable numbers all raise ValueError the same way.

Writing

write_mgf(spectra, path) -> Path
write_ms2(spectra, path, *, ionization_model=None) -> Path
write_msp(spectra, path) -> Path

All three take an iterable of Spectrum/MsnSpectrum (a lone spectrum is accepted too) and return the path written. mz and intensity go out at repr precision, so a write → read round trip reproduces them exactly. The mapped metadata in the tables above round-trips with them.

Writers build a temporary file beside the destination and replace the destination only after the entire iterable has been written and closed successfully. A failure preserves an existing file and removes the temporary output. Reading and writing the same path is supported, including gzip files. Neutral masses are rejected because these formats describe m/z values. Use JSON or NPZ to preserve a decharged spectrum.

Before using an MSP library for chromatogram extraction or exporting its retention times to another format, convert its rt values to seconds according to the library's source convention. The reader cannot infer an MSP time unit.

Rule Detail
Profile data is refused A SpectrumType.PROFILE spectrum raises ValueError — peak lists are centroid data. Call .centroid() first
Polarity rides on the charge sign Neither format has a polarity field. A negative-polarity spectrum is written with a negative charge (CHARGE=2-, Z -2) and reads back with charge = -2
Missing metadata is omitted A plain Spectrum writes just its peaks. MS2's S line has no optional fields, so an absent precursor m/z becomes 0.0
Missing scan number An absent scan_number (a plain Spectrum, or mzML ids such as Bruker frame=… that carry no unique one) leaves MGF SCANS out, so a position never collides with a real scan number. MS2's S line needs one, so there it is the 1-based position in the input. The native id still goes out as MGF TITLE and as an MS2 I NativeID line
MGF TITLE / MSP Name native_id, falling back to scan=<scan_number>
MS2 Z mass Derived from the precursor m/z and charge (singly protonated mass). It is regenerated on write and ignored on read
rt in MS2 Written as minutes (I RTime), so it returns to within floating-point noise rather than bit-exact. MGF's RTINSECONDS is exact
rt in MSP Written verbatim under RetentionTime: (spxtacular's rt is seconds; MSP has no unit convention), so it round-trips exactly
Polarity in MSP Written as an explicit Ion_mode: P/N line — MSP has its own field, so unlike MGF/MS2 the charge keeps its sign

Things that do not survive a round trip, by design: im, iso_score, mz_range, isolation_mz_range, collision_energy, resolution, analyzer, and MsnSpectrum.ms_level for anything other than MS2 — no peak-list format has a field for them. A DECONVOLUTED spectrum written to MGF comes back as CENTROID with a per-peak charge array.


AcquisitionType

class AcquisitionType(StrEnum):
    DDA = "DDA"
    DIA = "DIA"
    PRM = "PRM"
    UNKNOWN = "UNKNOWN"

Detected automatically by DReader from the .d directory. Accessible as reader.acquisition_type. UNKNOWN falls back to the DDA reader; PRM has its own MS2 iteration path (per transition).