Repository navigation
[models] Add SeizureTransformer - #1236
Conversation
SeizureTransformer (Wu, Zhao and Yener, 2025) labels every sample of an EEG window with a U-shaped network: a convolutional encoder, residual convolution blocks, a Transformer encoder and a convolutional decoder with skip connections. It won the 2025 SzCORE seizure detection challenge. The braindecode class is written from the paper and returns per-sample logits of shape (batch, n_outputs, n_times). The authors' released checkpoint converts to it: all 216 tensors the network uses load strictly, and the outputs match the reference code to 1e-14 in float64. Closes braindecode#1199
|
Integration gate (braindecode maintainers) Thanks @raghav-rathi for a careful port and an unusually complete PR body: the checkpoint hash, key map and parity table made this quick to check. Target. The paper's own cells (TUSZ, SeizeIT1, Dianalund) sit behind DUAs or are private. So the gate is the one public score of the released competition checkpoint ( Protocol. Inference only, all 41 recordings, run on Voyager from one
The port and the original wrote byte-identical annotation TSVs for 41/41 recordings. The current-evaluator numbers match the benchmark to every digit. On a random 4×19×15360 input, the CPU probability difference against the authors' architecture is 7.5e-8 in fp32. Verdict: REPLICATED (gap 0 %, accepted gap 5 %). Not blocking: the six red CI jobs were cancelled by runner starvation, not failed. The recomputed positional table differs from the stored one by 1.5e-5 (float32 |
|
Thanks a lot for running the whole replication on your side! You're right about the positional table. On my machine it came out exactly equal but float32 sin isn't the same on every platform so I updated that line in the description |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1236 +/- ##
==========================================
+ Coverage 87.99% 88.04% +0.05%
==========================================
Files 151 152 +1
Lines 17628 17713 +85
==========================================
+ Hits 15511 15596 +85
Misses 2117 2117 🚀 New features to boost your workflow:
|
bruAristimunha
left a comment
There was a problem hiding this comment.
Thanks for the port! Replication on Voyager via NeuralBench: the released SzCORE checkpoint through this port scores event-based F1 0.7060 on Siena (41 recordings), identical to the authors' code and to the published SzCORE benchmark to full float precision; weight parity 7.5e-8 (fp32), strict load of all 216 tensors. I bumped versionadded to 1.9.0. CI green.
Model information
SIENAandCHBMITbut no seizure detection model yet. SeizureTransformer won the 2025 SzCORE challenge and is first on the list in Epilepsy Benchmarks + Braindecode #828. It gives a prediction for every time sample, so seizure onsets and offsets come straight out of the model. Closes Add SeizureTransformer, the 2025 SzCORE seizure detection winner #1199Implementation fidelity
Reference implementation: https://github.com/keruiwu/SeizureTransformer at cf83f59 under MIT, plus the competition checkpoint from the authors' Docker image
yujjio/seizure_transformer:latest, filewu_2025/model.pthwith SHA-25679b14e4715fef055ba252ae3f7a072325907e3d5d11bf91a8070d4220995f65d. I wrote the class from the layer table in the paper (Table 13) and checked it against their code. No code is copied from their repo.Deviations:
(batch, n_outputs, n_times)instead of sigmoid probabilities of shape(batch, n_times)drop_probfor the residual blocks, the positional encoding and the Transformer. Upstream uses 0.1 for all threeceil_mode=Trueand a crop in the decoder instead of padding with -1e10 and precomputed crops. The outputs are the same, see below, and inputs shorter thann_timesalso workbraindecode.functional.sinusoidal_positional_encoding. On my machine it equals the stored one exactly but on others it can differ by up to 1.5e-5 because float32 sin varies between platforms. This does not change the outputsnn.TransformerEncoderdeep-copies the layer it gets. This model leaves it out, so it has 37,848,897 parameters instead of 41,001,281Checkpoint/parity evidence: After renaming, all 216 tensors the network uses load with
strict=True. Largest absolute difference of the logits against the reference code on the same inputOn the GPU, PyTorch's default TF32 convolutions shift the logits of both implementations away from the float64 result, so the GPU column has TF32 off.
Checklist
Implementation (
braindecode/models/seizure_transformer.py)EEGModuleMixinbeforenn.Module,license="mit", attribution in the header and inNOTICE.txtsuper().__init__(...)(batch, n_chans, n_times)to(batch, n_outputs, n_times)per-sample logitsself.final_layeris the last child module and there is no final sigmoid or softmaxactivation=nn.ELUandactivation_res=nn.ReLUclass defaults anddrop_prob[Wu2025]_referencesRegistration and documentation
braindecode/models/__init__.pymodels_mandatory_parameterswithn_times=1024to keep CI fast, and an entry innon_classification_modelsbraindecode/models/summary.csvwith 37,848,897 parametersdocs/api.rstdocs/_static/model/seizure_transformer_arch.pngfrom the authors' repo under MIT, listed inNOTICE.txtdocs/whats_new.rstValidation and compatibility
test_models.py. They check one prediction per input sample for even, odd and shorter inputs, and errors for too-long inputs and invalid settingspre-commit run --from-ref origin/master --to-ref HEADpasses with all hookstest_reset_head_updates_configandtest_reset_head_model_reloads_after_savingpass, andsave_pretrainedthenfrom_pretrainedof a trained model gives identical outputsreset_headreplaces the output convolution, no sentinel valuesValidation evidence
pytest test/unit_tests/models/ -k "SeizureTransformer or seizure_transformer"on CPU gives 25 passed and 5 skipped. The skips are the non-classification exemptions and the embedding-parameter test. This includestorch.compile,torch.export, TorchScript and the new contract tests from Add registry-wide model contract gates #1208test_model_contract.pypasses with this model registered, 146 passed and 2 skippedpytest -vv --durations=0 -n 16 --dist worksteal test/on CPU with Linux, Python 3.12 and torch 2.14 gave 4084 passed, 261 skipped and 0 failed. That run was on master from earlier today. After rebasing on the latest master I reran the tests above and pre-commitEEGRegressor,BCEWithLogitsLossand RAdam at lr 1e-4 and weight decay 2e-5 as in the paper. From scratch on 160 Siena windows from 11 subjects, 12 epochs on one RTX 4090 in 12 s. Training loss went from 0.714 to 0.356 and per-sample AUROC on 34 windows from 3 held-out subjects went from 0.588 to 0.871, or 0.855 for the same run on CPU. This only shows that training works, it is not a benchmarkGradScaler, follows the float32 loss curve, and the largest activation in the network is about 50, far below the float16 limitBenchmark reproduction
braindecode.datasets.SIENAresults/yujjio-seizure-transformer-latest/siena.jsonin esl-epfl/szcoreAll 41 Siena recordings were loaded with
braindecode.datasets.SIENAand prepared the way the authors do it. Each channel is z-scored over the whole recording and cut into 60 s windows with the last one zero padded. Each window then gets a causal order 3 Butterworth band-pass from 0.5 to 120 Hz and notch filters at 1 Hz and 60 Hz. This class with the converted checkpoint predicted on the CPU and on one RTX 4090. After the authors' post-processing, which is a 0.8 threshold, opening and closing with 5 samples and removing events shorter than 2 s,szcore-evaluationscored the result the same way the leaderboard does.All eight values match the published ones on both devices, event precision only differs in the 16th decimal. Run under the same settings the reference code and this class also write identical event files for all 41 recordings, on the CPU and on the GPU with TF32 off.
Siena was part of the training data for this checkpoint, so this shows that braindecode reproduces the released model and not how well it generalises. The paper's held-out results use TUSZ, which needs a TUH data agreement, plus SeizeIT1 and Dianalund. I did not run those.
Notes for reviewers
braindecode/like the other ports? I'd ask the authors first and then add thefrom_pretrainedexample and the Hub test entry, here or in a follow-up PR.. versionadded:: 1.8.2like DIVER1 and BaRISTA