Repository navigation
[models] Add NeuroRVQ residual tokenizer - #1223
bruAristimunha merged 22 commits into
Conversation
Resolved docs/whats_new.rst by keeping both entries.
…he tokenizer, fold its tests into test_models
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #1223 +/- ##
==========================================
+ Coverage 89.02% 89.16% +0.14%
==========================================
Files 157 158 +1
Lines 19398 19641 +243
==========================================
+ Hits 17269 17513 +244
+ Misses 2129 2128 -1 🚀 New features to boost your workflow:
|
…nd document its outputs
…construction error
d6fef7e to
718dffb
Compare
|
Additional local verification at PR head
These targeted runs passed locally; they supplement but do not replace the repository's full cross-platform CI matrix. |
|
I inspected the Windows Python 3.12 check annotation. It reports the runner failed while finalizing its paging log because the runner disk was full ( |
|
@bruAristimunha I re-ran the focused checks on the current PR head The branch includes the upstream review follow-ups and is synced to current |
…r options and helpers; outputs unchanged
…ph and output-compare helper; deflake the tokenize test
|
Follow-up on the maintainer-updated head |
|
I independently reran the pinned released-checkpoint parity on the current maintainer-updated PR head The standalone script, exact source/checkpoint hashes, result JSON, method, reproduction instructions, and limitations are published on my fork: https://github.com/lindicaphxag-tech/braindecode/tree/research/neurorvq-tokenizer-parity/research/neurorvq_tokenizer_parity. This is deterministic implementation parity on one synthetic input shape; it is not a paper benchmark reproduction or task-performance result. |
|
CI update for maintainer-updated head |
|
Methodology clarification for the public evidence branch: the paper reports High Gamma raw-signal MSE 0.084 in Table 10, but does not specify enough preprocessing and aggregation detail to prove an exact reproduction. The public EEG-Benchmarking repository documents a compatible protocol, not evidence that it is identical to the authors' Table 10 pipeline. I updated the evidence README to label the pending all-subject run as a protocol-aligned replication attempt and to report any difference without claiming an exact reproduction: https://github.com/lindicaphxag-tech/braindecode/blob/research/neurorvq-tokenizer-parity/research/neurorvq_tokenizer_parity/README.md Could you clarify whether reproducing 0.084 under this public protocol is the intended |
|
CI unblock note for current exact PR head |
Conflict: docs/whats_new.rst (kept the NeuroRVQTokenizer entry from braindecode#1223 and the AXON entry).
…gh spectral_input
Model information
NeuroRVQTokenizerImplementation fidelity
926e770d9d16b6aa308404280fa0cc0211a6f9fb; CC BY-NC 4.0. Released checkpoint is pinned to Hugging Face revisiond944b87f44ae0ba2923b2f10d0518f23f6803b76.7.11e-15. The script asserts every trainable parameter receives a gradient. This is implementation parity, not reproduction of paper benchmark scores.Checklist
Implementation
EEGModuleMixinintegration, source attribution, license declaration andNOTICE.txtentryfinal_layerconvention: N/A; tokenizer returns reconstructed patches and codes, and is registered as a non-classification modelRegistration and documentation
summary.csvrowdocs/whats_new.rstValidation and compatibility
reset_headand mixed classifier head modes: N/A; tokenizer has no classification headValidation evidence
pytest test/unit_tests/models/test_neurorvq_tokenizer.py -q— 7 passed.pytest test/unit_tests/models/test_integration.py -k NeuroRVQTokenizer -q— 6 passed, 6 skipped, 911 deselected.python scripts/validate_neurorvq_tokenizer_parity.py --neurorvq-source <NeuroRVQ clone> --checkpoint <pinned checkpoint>— exact reference outputs/codes/input gradients/EMA state; max parameter-gradient error7.11e-15over 345 trainable parameters.pre-commit run --files <changed files>— passed, including ruff, mypy, docstrfmt, sphinx-lint and isort;git diff --checkpassed.Benchmark reproduction
The paper's public-dataset benchmark is not claimed as reproduced. This contribution integrates the released tokenizer and verifies implementation parity; benchmark reproduction requires the paper's complete datasets and training/evaluation protocol.
Current upstream gate status — 2026-10-09
Release-ready review surface: the model integration prerequisite
Braindecode #1218
has already been merged and independently replicated on public EEG
by the upstream maintainers (5 seeds × 10 subject-independent folds,
BAcc 0.8370 ± 0.0028 against the published 0.869 ± 0.026, inside
their predeclared 5% band). That is the upstream EEG foundation-model
port's maintainer-run replication, not a tokenizer Table-10 benchmark or
a new task improvement claimed for this PR.
For the current tokenizer PR head
f5307aa08242e3bb9f5d84866f84decb90ac5ca5, GitHub Actions reports:Python 3.12 and 3.13, including 2 real-data acceptance jobs).
mark documentation green until that particular upstream job succeeds.
at maintainer discretion.
Current-source CI checks.
The public author-run reference-parity artifact
was frozen at earlier maintainer-refactored head
61194d5and compares checkpoint/reconstruction/discrete IDs/EMA and gradients.
The current
f5307aahead passes the full upstream test matrix;the author parity run is not claimed to be freshly rerun against
every intervening upstream commit. No full 14-subject High Gamma
reconstruction benchmark has been independently audited.
Suggested minimal review decision: assess whether the self-contained
tokenizer's checkpoint-loading, source license, quantizer EMA and public
API are acceptable. A future EEG benchmark reproduction can be a
separate research deliverable rather than an implied current PR result.
Notes for reviewers
The tokenizer is self-contained with respect to Braindecode: it does not import from or require another open PR. The original NeuroRVQ license is noncommercial (CC BY-NC 4.0), recorded in the model and repository notices.