Repository navigation
[models] Add NeuroRVQ EEG foundation model - #1218
Conversation
61d744e to
b849701
Compare
|
@bruAristimunha this is ready for maintainer review when convenient. The current head has exact CPU parity against the released NeuroRVQ checkpoint (263/263 shared encoder tensors; max logit and input-gradient error 0.0), and I clarified the licensing boundary in the PR body: no checkpoint bytes are vendored, the adapted module is explicitly CC BY-NC 4.0 in NOTICE.txt, and weights are fetched from a pinned upstream revision at runtime. I’ll hold the head stable unless review identifies a concrete issue. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #1218 +/- ##
==========================================
- Coverage 87.99% 87.98% -0.01%
==========================================
Files 151 152 +1
Lines 17628 17847 +219
==========================================
+ Hits 15511 15702 +191
- Misses 2117 2145 +28 🚀 New features to boost your workflow:
|
Resolved docs/whats_new.rst conflict by keeping both the NeuroRVQ entry and master's new llms.txt/docs entry. All other paths auto-merged, including training/losses.py and test/unit_tests/training/test_losses.py, which now exactly match origin/master (the braindecode#1198 CroppedLoss fix and its regression test are restored).
- Replace NeuroRVQ's private _DropPath/_Mlp with braindecode.modules DropPath/MLP; load_pretrained_weights remaps the released checkpoint's mlp.fc1/mlp.fc2 keys to mlp.0/mlp.2 to match the new module. Parity verified at the _Block level with a state-dict-mapped random init: max-abs diff 0.0 in both eval and train mode (drop_prob=0.0 by default). - Fold test_neurorvq.py into test_models.py, keeping every case. - Trim the NOTICE.txt NeuroRVQ entry to the file-list-only convention and add the upstream LICENSE link to the model docstring.
Resolved docs/whats_new.rst by keeping both entries.
|
Integration gate (braindecode maintainers), covering #1218 (model) and #1223 (standalone tokenizer)
Result (2026-10-06 03:56 Paris, 5 % gate): REPLICATED: 5-seed mean 0.8370 ± 0.0028 (SD over seeds; mean fold-SD 0.0324) vs paper 0.869 ± 0.026, band [0.8256, 0.9125], rel gap 3.7 %. Numbers recomputed from the per-cell files; full record in #1223 tokenizer (updated 2026-10-06 21:30 Paris, head
|
# Conflicts: # braindecode/models/summary.csv # test/unit_tests/models/test_models.py
bruAristimunha
left a comment
There was a problem hiding this comment.
Thanks for the port! Replication on Voyager via NeuralBench, paper Table 1(A) Eyes (PhysioNet eyes open vs closed), released fine-tuning recipe, 5 seeds x 10 folds: balanced accuracy 0.837 +/- 0.003 vs paper 0.869 +/- 0.026 (gap 3.7%, inside the 5% gate). Tokenizer parity vs the reference code 2.98e-8. I merged master in (summary.csv / test_models.py keep-both). CI green.
whats_new: keep both; refs for NeuroRVQ (braindecode#1218), BrainTokenizer (braindecode#1230), BrainOmni (braindecode#1231) point to the merged PRs.
Summary
Closes #1090
Implementation fidelity
Reference: https://github.com/KonstantinosBarmpas/NeuroRVQ (CC BY-NC 4.0; attribution and noncommercial terms are recorded in the module header and
NOTICE.txt). The checkpoint is fetched from the pinned Hugging Face revisiond944b87f44ae0ba2923b2f10d0518f23f6803b76.A CPU parity probe loads the released checkpoint into the reference model and this implementation, copies the task head, and compares logits plus input gradients for a 3-channel, 4-patch input. All 263 shared encoder tensors load; maximum absolute logit and input-gradient differences are both 0.0. The checkpoint loader also rejects missing or unexpected encoder tensors outside the documented pretraining-only heads.
The port also preserves the upstream zero-based spatial-slot convention:
create_embedding_ixmaps the first electrode to slot 0 and the reference forward path separately pads the CLS position with another 0; a focused regression locks this unusual but checkpoint-relevant behavior.Licensing / distribution boundary
braindecode/models/neurorvq.pymodule is explicitly marked CC BY-NC 4.0 and is listed inNOTICE.txt, following Braindecode's existing per-file third-party licensing pattern.Validation
pytest test/unit_tests/models/test_neurorvq.py test/unit_tests/models/test_integration.py -k NeuroRVQ -q— 21 passed, 3 skipped.pytest test/unit_tests/models/test_return_features.py test/unit_tests/models/test_models.py -k 'NeuroRVQ or completeness_summary_table' -q— 4 passed.pytest test/unit_tests/models/test_integration.py -k 'completeness_summary_table and NeuroRVQ' -q— 1 passed.ruff check braindecode/models/neurorvq.py test/unit_tests/models/test_neurorvq.py— passed.ruff format --check braindecode/models/neurorvq.py test/unit_tests/models/test_neurorvq.py braindecode/models/util.py— passed.git diff --check— passed.The paper's dataset benchmark was not rerun here. The parity result validates implementation agreement with the released checkpoint, not reproduction of published downstream benchmark scores.
Channel provenance
Released pretrained weights require explicit
channel_namesorchs_info. Randomly initialized models may use the first-N reference-channel fallback for integration/training-from-scratch, butload_pretrained_weights()rejects that ambiguous mapping so pretrained spatial embeddings cannot be silently assigned to the wrong electrodes.