Skip to content

Integration checks for the error patterns that break models on accelerators and low precision (CPU-only); fixes for what they catch - #1253

Merged
bruAristimunha merged 16 commits into
braindecode:masterfrom
bruAristimunha:test/model-integration
Oct 9, 2026
Merged

bruAristimunha merged 16 commits into
braindecode:masterfrom
bruAristimunha:test/model-integration

Conversation

@bruAristimunha

Copy link
Copy Markdown
Collaborator

CPU-only integration checks, run on every model in models_dict, for the error patterns that broke models on accelerators (Intel Gaudi) and in low precision. No accelerator is needed: the forward and backward ops are recorded on CPU and checked.

Checks (each guards a bug class we already hit):

  • No complex tensor outside braindecode.functional.spectral_input (Gaudi has no complex dtype).
  • No host sync (.item(), bool(tensor)) or data-dependent shape (nonzero, boolean-mask indexing) in forward, except listed init-time cases with a reason (they break lazy graphs).
  • No op from a short list of known kernel gaps: avg_pool3d in bf16/fp16, low-precision cdist, ELU/SELU backward with scale != 1, roll on a non-contiguous input, an LSTM fed by a permuted Conv1d output.
  • Tensors follow the model: forward on the meta device touches no CPU tensor, and after model.to(torch.float64) every floating tensor created in forward is float64.
  • A training step after a forward under torch.inference_mode(); deepcopy and pickle; strict state_dict round trip.
  • For the pretrained models, model(x) and model.forward(x) agree under a channel strategy and the config keeps the input montage.

Model fixes (root cause → equivalent op):

  • BrainOmni, BrainTokenizer: nn.SELU backward has no Gaudi kernel (ELU scale != 1) → scale * x / F.elu(x, alpha * scale).
  • EEGSym: avg_pool3d has no CPU bf16/fp16 kernel → avg_pool2d over the time axis.
  • FBCNet, FBMSNet, FBLightConvNet (StatLayer): clamp at 1e6 overflows float16 → clamp and log in at least float32.
  • LUNA: the 1e-8 in the channel-position min-max normalisation underflows in float16 (NaN) → normalise in at least float32.
  • EEGMiner (GeneralizedGaussianFilter): a non-persistent filters buffer built with grad and overwritten in forward broke deepcopy after training → local tensor.
  • AttnSleep: attention weights stashed on the module (held the graph, not moved by .to) → dropped.
  • MLP (Labram, NeuroRVQ): stored a lambda, so the models did not pickle → no lambda.
  • SignalJEPA heads: accept channel_strategy like SignalJEPA.
  • CodeBrain, TCFormer, LUNA, MVPFormer, BrainOmni/BrainTokenizer, ZUNA, DIVER1: tensors built in forward without the input's dtype/device → follow the input; EEGDINO: host-side one_hot → on device.

Float32 CPU outputs, gradients and state dicts are bit-identical to master for every touched model. With the #1249 fixes reverted, the new checks fail on exactly those models (EEGSym, BrainOmni/BrainTokenizer, EMG2QwertyNet, MetaNeuromotorHand, AttnSleep, EEGDINO and the dtype-follow models).

Copilot AI balanced review requested due to automatic review settings October 8, 2026 14:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI balanced review requested due to automatic review settings October 8, 2026 14:07

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@codecov

codecov Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 89.25%. Comparing base (cb1ca88) to head (17d116e).
⚠️ Report is 1 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1253      +/-   ##
==========================================
+ Coverage   89.23%   89.25%   +0.02%     
==========================================
  Files         158      158              
  Lines       19604    19627      +23     
==========================================
+ Hits        17493    17518      +25     
+ Misses       2111     2109       -2     
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Copilot AI balanced review requested due to automatic review settings October 8, 2026 14:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI balanced review requested due to automatic review settings October 8, 2026 18:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI balanced review requested due to automatic review settings October 8, 2026 20:30

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

# Conflicts:
#	braindecode/models/neurorvq_tokenizer.py
Copilot AI balanced review requested due to automatic review settings October 8, 2026 20:38

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@bruAristimunha
bruAristimunha merged commit 62ca9f2 into braindecode:master Oct 9, 2026
16 checks passed
bruAristimunha added a commit to lindicaphxag-tech/braindecode that referenced this pull request Oct 9, 2026
…raindecode#1253) into EEG-CLIP; EEGCLIP text side in _UNUSED_IN_FORWARD
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants