Hello, I was looking through the code of labram.py to replicate it's forward pass for my own related work, and I feel as if the addition of the temporal embeddings to the patch embeddings may be being handled incorrectly.
In the case of self.neural_tokenizer being set to False, we get patch embeddings as:
x = self.patch_embed(x) and then we set
batch_size, n_patch, temporal = x.shape
In this set up, I believe that the n_patch variable refers to all of the patch embeddings for some multi channel EEG input, and not
the number of patches per channel. This is important because then the variable n_time_tokens is assigned the value...
n_time_tokens = min(n_patch, self.temporal_embedding.shape[1] - 1)
To my understanding of the LaBram architecture, each temporal embedding is an absolute positional encoding that refers to a particular slice of time across all channels, so then n_time_tokens I feel should be set to something more like,
n_time_tokens = min(n_patch_per_chan,self.temporal_embedding.shape[1] - 1)
Hello, I was looking through the code of labram.py to replicate it's forward pass for my own related work, and I feel as if the addition of the temporal embeddings to the patch embeddings may be being handled incorrectly.
In the case of
self.neural_tokenizerbeing set to False, we get patch embeddings as:x = self.patch_embed(x)and then we setbatch_size, n_patch, temporal = x.shapeIn this set up, I believe that the
n_patchvariable refers to all of the patch embeddings for some multi channel EEG input, and notthe number of patches per channel. This is important because then the variable
n_time_tokensis assigned the value...n_time_tokens = min(n_patch, self.temporal_embedding.shape[1] - 1)To my understanding of the LaBram architecture, each temporal embedding is an absolute positional encoding that refers to a particular slice of time across all channels, so then
n_time_tokensI feel should be set to something more like,