Skip to content

[models] Add NeuroRVQ residual tokenizer - #1223

Merged
bruAristimunha merged 22 commits into
braindecode:masterfrom
lindicaphxag-tech:feat/neurorvq-tokenizer-standalone
Oct 8, 2026
Merged

bruAristimunha merged 22 commits into
braindecode:masterfrom
lindicaphxag-tech:feat/neurorvq-tokenizer-standalone

Conversation

@lindicaphxag-tech

@lindicaphxag-tech lindicaphxag-tech commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Model information

Implementation fidelity

  • Reference implementation: KonstantinosBarmpas/NeuroRVQ, source commit 926e770d9d16b6aa308404280fa0cc0211a6f9fb; CC BY-NC 4.0. Released checkpoint is pinned to Hugging Face revision d944b87f44ae0ba2923b2f10d0518f23f6803b76.
  • Deviations: Adds Braindecode mixin/API integration, strict checkpoint loading and cold-codebook initialization. The released architecture and preprocessing contract are preserved (200 Hz, 200-sample patches, released 104-channel montage). Required transformer and multiscale embedding layers are included as private helpers, so this PR has no dependency on another open PR. In training, EMA counts and sums follow the reference distributed all-reduce path. Eval/tokenize inference is state-pure; the reference increments a usage-statistics buffer in eval, but that buffer does not affect weights, codes, reconstruction, or gradients.
  • Checkpoint/parity evidence: CPU parity against the released source and pinned checkpoint: exact eval/train targets and reconstructions, exact token indices, input gradients, and EMA state after one training update; maximum absolute difference across 345 trainable parameter gradients is 7.11e-15. The script asserts every trainable parameter receives a gradient. This is implementation parity, not reproduction of paper benchmark scores.

Checklist

Implementation

  • EEGModuleMixin integration, source attribution, license declaration and NOTICE.txt entry
  • Documented and tested task-specific inputs/outputs: standardized patches and discrete codes, not class logits
  • Model docstring describes architecture, parameters and paper
  • No new runtime dependency; Hub loading uses existing optional dependency
  • Classifier final_layer convention: N/A; tokenizer returns reconstructed patches and codes, and is registered as a non-classification model

Registration and documentation

  • Public export and model registry integration
  • Mandatory-parameter integration and summary.csv row
  • API entry, model guide, and docs/whats_new.rst
  • Architecture figure: N/A; the tokenizer is documented with its patch/token data flow and does not introduce a separate figure

Validation and compatibility

  • Model-specific and integration regression tests
  • Checkpoint save/load round trip covered by model tests
  • Shared helpers are private to this port; no existing model/API behavior is changed
  • reset_head and mixed classifier head modes: N/A; tokenizer has no classification head
  • Released checkpoint parity script covers train/eval forward, gradients, discrete codes and EMA update

Validation evidence

  • pytest test/unit_tests/models/test_neurorvq_tokenizer.py -q — 7 passed.
  • pytest test/unit_tests/models/test_integration.py -k NeuroRVQTokenizer -q — 6 passed, 6 skipped, 911 deselected.
  • python scripts/validate_neurorvq_tokenizer_parity.py --neurorvq-source <NeuroRVQ clone> --checkpoint <pinned checkpoint> — exact reference outputs/codes/input gradients/EMA state; max parameter-gradient error 7.11e-15 over 345 trainable parameters.
  • pre-commit run --files <changed files> — passed, including ruff, mypy, docstrfmt, sphinx-lint and isort; git diff --check passed.
  • Full repository test suite was not run. A full Sphinx build was attempted earlier but was blocked by unrelated optional MOABB and local TensorBoard/protobuf environment issues; the tokenizer documentation checks pass.

Benchmark reproduction

The paper's public-dataset benchmark is not claimed as reproduced. This contribution integrates the released tokenizer and verifies implementation parity; benchmark reproduction requires the paper's complete datasets and training/evaluation protocol.

Current upstream gate status — 2026-10-09

Release-ready review surface: the model integration prerequisite
Braindecode #1218
has already been merged and independently replicated on public EEG
by the upstream maintainers (5 seeds × 10 subject-independent folds,
BAcc 0.8370 ± 0.0028 against the published 0.869 ± 0.026, inside
their predeclared 5% band). That is the upstream EEG foundation-model
port's maintainer-run replication
, not a tokenizer Table-10 benchmark or
a new task improvement claimed for this PR.

For the current tokenizer PR head
f5307aa08242e3bb9f5d84866f84decb90ac5ca5, GitHub Actions reports:

  • Full test matrix: 8/8 jobs green (Ubuntu, Windows, macOS;
    Python 3.12 and 3.13, including 2 real-data acceptance jobs).
  • Code style: green; What's New check: green.
  • Sphinx docs job: still in progress when checked; do not
    mark documentation green until that particular upstream job succeeds.
  • PR state: open, non-draft and mergeable; actual merge remains
    at maintainer discretion.

Current-source CI checks.
The public author-run reference-parity artifact
was frozen at earlier maintainer-refactored head 61194d5
and compares checkpoint/reconstruction/discrete IDs/EMA and gradients.
The current f5307aa head passes the full upstream test matrix;
the author parity run is not claimed to be freshly rerun against
every intervening upstream commit. No full 14-subject High Gamma
reconstruction benchmark has been independently audited.

Suggested minimal review decision: assess whether the self-contained
tokenizer's checkpoint-loading, source license, quantizer EMA and public
API are acceptable. A future EEG benchmark reproduction can be a
separate research deliverable rather than an implied current PR result.

Notes for reviewers

The tokenizer is self-contained with respect to Braindecode: it does not import from or require another open PR. The original NeuroRVQ license is noncommercial (CC BY-NC 4.0), recorded in the model and repository notices.

@lindicaphxag-tech
lindicaphxag-tech marked this pull request as ready for review October 4, 2026 11:42
@bruAristimunha bruAristimunha added model Adds a new model needs-replication Model PR: paper number must be replicated (NeuralBench) before merge labels Oct 5, 2026
Resolved docs/whats_new.rst by keeping both entries.
@codecov

codecov Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.70782% with 8 lines in your changes missing coverage. Please review.
✅ Project coverage is 89.16%. Comparing base (6abd308) to head (f5307aa).

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1223      +/-   ##
==========================================
+ Coverage   89.02%   89.16%   +0.14%     
==========================================
  Files         157      158       +1     
  Lines       19398    19641     +243     
==========================================
+ Hits        17269    17513     +244     
+ Misses       2129     2128       -1     
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lindicaphxag-tech
lindicaphxag-tech force-pushed the feat/neurorvq-tokenizer-standalone branch from d6fef7e to 718dffb Compare October 7, 2026 18:23
@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

Additional local verification at PR head 718dffbc7474d2a1dcce0ac938a49baacd5bb89f (Python 3.12.4):

  • python -m pytest test/unit_tests/models/test_models.py -k neurorvq -q — 31 passed, 832 deselected.
  • python -m pytest test/unit_tests/models/test_integration.py -k neurorvq -q — 15 passed, 9 skipped, 1,000 deselected.

These targeted runs passed locally; they supplement but do not replace the repository's full cross-platform CI matrix.

@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

I inspected the Windows Python 3.12 check annotation. It reports the runner failed while finalizing its paging log because the runner disk was full (There is not enough space on the disk); this is an infrastructure failure, not a NeuroRVQ assertion or import failure. The check metadata still shows the full-suite step as in progress, and the overall workflow has not reached a terminal state. The focused model and integration test results on this PR head are already recorded in my earlier comment.

@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

@bruAristimunha I re-ran the focused checks on the current PR head 718dffbc7474d2a1dcce0ac938a49baacd5bb89f (Python 3.12.4): model tests 31 passed; integration tests 15 passed, 9 skipped. The latest full CI run is still reported in progress; its Windows 3.12 job annotation says the runner ran out of disk while finalizing logs, not a tokenizer assertion failure.

The branch includes the upstream review follow-ups and is synced to current master. Could you take a look at the current state and let me know if you want any remaining code changes? If the runner is still wedged, a fresh CI run may be needed for a final full-suite result.

@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

Follow-up on the maintainer-updated head 61194d513e0c5234a2f55923bf157643354e4b42: I re-ran the focused checks after the shared-layer refactor and loss cleanup. test_models.py -k neurorvq passes (26 passed); test_integration.py -k neurorvq passes (15 passed, 9 skipped). On GitHub, pre-commit, acceptance, Codacy, and CircleCI are green; the full cross-platform suite and docs build are still pending.

@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

I independently reran the pinned released-checkpoint parity on the current maintainer-updated PR head 61194d513e0c5234a2f55923bf157643354e4b42, after the shared-layer refactor and loss cleanup. Against official NeuroRVQ source 926e770d9d16b6aa308404280fa0cc0211a6f9fb and HF checkpoint revision d944b87f44ae0ba2923b2f10d0518f23f6803b76, evaluation and training outputs matched exactly, token IDs and EMA state matched, and max gradient error across 345 trainable parameters was 7.11e-15.

The standalone script, exact source/checkpoint hashes, result JSON, method, reproduction instructions, and limitations are published on my fork: https://github.com/lindicaphxag-tech/braindecode/tree/research/neurorvq-tokenizer-parity/research/neurorvq_tokenizer_parity. This is deterministic implementation parity on one synthetic input shape; it is not a paper benchmark reproduction or task-performance result.

@lindicaphxag-tech

Copy link
Copy Markdown
Contributor Author

CI update for maintainer-updated head 61194d513e0c5234a2f55923bf157643354e4b42: Ubuntu 3.13 and Windows 3.13 full test jobs have completed successfully, in addition to the earlier acceptance, pre-commit, Codacy, and CircleCI checks. Ubuntu 3.12, Windows 3.12, macOS, and docs are still running. The pinned released-checkpoint parity evidence is also linked in my earlier comment. Could you review the current PR when the remaining matrix finishes?

Copy link
Copy Markdown
Contributor Author

Methodology clarification for the public evidence branch: the paper reports High Gamma raw-signal MSE 0.084 in Table 10, but does not specify enough preprocessing and aggregation detail to prove an exact reproduction. The public EEG-Benchmarking repository documents a compatible protocol, not evidence that it is identical to the authors' Table 10 pipeline. I updated the evidence README to label the pending all-subject run as a protocol-aligned replication attempt and to report any difference without claiming an exact reproduction: https://github.com/lindicaphxag-tech/braindecode/blob/research/neurorvq-tokenizer-parity/research/neurorvq_tokenizer_parity/README.md

Could you clarify whether reproducing 0.084 under this public protocol is the intended needs-replication gate, or point me to the exact Table 10 preprocessing/aggregation protocol you expect? The full Kaggle run is still in progress; no benchmark score is claimed yet.

Copy link
Copy Markdown
Contributor Author

CI unblock note for current exact PR head f5307aa08242e3bb9f5d84866f84decb90ac5ca5 (no code changes): I checked the upstream Actions job list rather than inferring from the green summary. All eight full test/acceptance jobs succeeded (Ubuntu, macOS, Windows; Python 3.12/3.13), and code-style/What's New checks are green. The remaining docs workflow 37788241377 has build_docs job 113348911701 still in_progress, specifically step Create Docs; previous installation/import/cache steps passed, and later docs steps have not started. I cannot treat that pending job as a pass, and I do not have upstream Actions permission to cancel or rerun it. If this is an orphaned runner rather than an active Sphinx build, could a maintainer cancel/retry that workflow/job or advise whether a docs-only rerun is required for review? I will leave the tested source unchanged until there is a concrete failure or maintainer-requested change. This is infrastructure triage, not new tokenizer benchmark evidence.

@bruAristimunha
bruAristimunha merged commit 83a1e67 into braindecode:master Oct 8, 2026
16 checks passed
bruAristimunha added a commit to mahirjain01/braindecode that referenced this pull request Oct 8, 2026
Conflict: docs/whats_new.rst (kept the NeuroRVQTokenizer entry from braindecode#1223 and the AXON entry).
bruAristimunha added a commit to bruAristimunha/braindecode that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Adds a new model needs-replication Model PR: paper number must be replicated (NeuralBench) before merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants