Repository navigation
rbd: fix healer staging path for Block volumeMode - #6257
Conversation
|
Thanks for the PR! Please correct the subject of the commit, so that it says: Ideally an e2e test is added for this as well. Please check e2e/rbd.go for a similar test that uses rbd-nbd and a filesystem volume, |
7575ab1 to
79e58f3
Compare
|
Thanks for the review! I've addressed both points:
|
79e58f3 to
717e127
Compare
|
@Mergifyio rebase |
1 similar comment
|
@Mergifyio rebase |
☑️ Command
|
717e127 to
3b6fced
Compare
|
Deprecation notice: This pull request comes from a fork and was rebased using |
✅ Branch has been successfully rebased |
|
/test ci/centos/mini-e2e/k8s-1.35 |
|
@Mergifyio rebase |
🛑 The pull request rule doesn't match anymoreDetailsThis action has been cancelled. |
Block volumeMode PVCs using rbd-nbd mounter permanently lose IO after CSI plugin pod restart due to two bugs in rbd_healer.go: 1. formatStagingTargetPath computes a filesystem-format path (.../csi/<driver>/<sha256>/globalmount) for all volumes, but Block volumes use a completely different path: .../csi/volumeDevices/staging/<pv-name>/ ValidateNodeStageVolumeRequest calls checkDirExists before any healer logic, so the wrong path causes an immediate InvalidArgument error and attachRBDImage never runs. 2. callNodeStageVolume detects Block volumes using CSI.FSType == "block", but Block volumeMode PVs have an empty FSType. This is always false, causing the healer to send a Mount capability instead of a Block capability. Filesystem PVCs are unaffected because their sha256 path is correct. Block PVCs have no fallback recovery path (kubelet does not re-call NodeStageVolume when the bind mount is still in mountinfo), so IO is permanently lost until the node is restarted. Fix by: - Adding a Block branch to formatStagingTargetPath using pv.Name as the staging subdirectory, matching kubelet actual path layout - Changing Block detection to use pv.Spec.VolumeMode - Removing the unused fsTypeBlockName constant Add an e2e test that verifies IO on a Block volumeMode PVC using rbd-nbd mounter is restored after nodeplugin restart. Signed-off-by: YLShiJustFly <[email protected]>
3b6fced to
f6f9c9d
Compare
|
/test ci/centos/k8s-e2e-external-storage/1.35 |
|
/test ci/centos/k8s-e2e-external-storage/1.36 |
|
/test ci/centos/mini-e2e-helm/k8s-1.35 |
|
/test ci/centos/upgrade-tests-cephfs |
|
/test ci/centos/mini-e2e-helm/k8s-1.36 |
|
/test ci/centos/mini-e2e/k8s-1.35 |
|
/test ci/centos/upgrade-tests-rbd |
|
/test ci/centos/mini-e2e/k8s-1.36 |
|
/test ci/centos/k8s-e2e-external-storage/1.34 |
|
/test ci/centos/mini-e2e-helm/k8s-1.34 |
|
/test ci/centos/mini-e2e/k8s-1.34 |
|
/retest ci/centos/k8s-e2e-external-storage/1.34 |
Deployment failed logs |
|
Hi @nixpanic, Just a gentle reminder regarding this PR. It has already received 2 approvals. Regarding the current CI failures:
Since these failures are completely unrelated to this RBD healer patch (which only refactors the internal path string formatting), could you please help apply the 'approved' label to bypass the CI and manually merge this, just like #6356? Thank you for your time and help! |
|
/retest ci/centos/mini-e2e/k8s-1.35 |
|
/retest ci/centos/mini-e2e/k8s-1.36 |
1 similar comment
|
/retest ci/centos/mini-e2e/k8s-1.36 |
|
Queued — the merge queue status continues in this comment ↓. |
|
Deprecation notice: This pull request comes from a fork and was queued with |
Merge Queue Status
This pull request spent 11 seconds in the queue, including 2 seconds running CI. Required conditions to merge
|
rbd: fix healer staging path for Block volumeMode
Block volumeMode PVCs using rbd-nbd mounter permanently lose IO after
CSI plugin pod restart due to two bugs in rbd_healer.go:
formatStagingTargetPath computes a filesystem-format path
(.../csi///globalmount) for all volumes, but Block
volumes use a completely different path:
.../csi/volumeDevices/staging//
ValidateNodeStageVolumeRequest calls checkDirExists before any healer
logic, so the wrong path causes an immediate InvalidArgument error and
attachRBDImage never runs.
callNodeStageVolume detects Block volumes using
CSI.FSType == "block", but Block volumeMode PVs have an empty FSType.
This is always false, causing the healer to send a Mount capability
instead of a Block capability.
Filesystem PVCs are unaffected because their sha256 path is correct.
Block PVCs have no fallback recovery path (kubelet does not re-call
NodeStageVolume when the bind mount is still in mountinfo), so IO is
permanently lost until the node is restarted.
Fix by:
staging subdirectory, matching kubelet actual path layout