Repository navigation
Add pod_cidr_exhaustion_hotel_reservation benchmark problem - #774
Conversation
There was a problem hiding this comment.
@MohammadElsharqawy Thanks for the PR! This fault is quite interesting and the code looks great. One thing, we'll need to automate the config-related changes.
The "Cluster Requirements" section lists manual steps that need to happen outside the PR:
# kind-config.yaml additions
networking:
disableDefaultCNI: true
podSubnet: "10.244.0.0/24"kubectl apply -f https://raw.githubusercontent.com/projectcalico/calico/v3.27.0/manifests/calico.yaml
kubectl patch ipamconfig default --type=merge -p '{"spec":{"strictAffinity":true}}'
rm ~/cache_dir/cluster_baseline_state.json
These are not added in the PR. The kind cluster currently uses kindnet (no Calico), and the real cluster has Calico but with a /16 pool. So this problem requires manual cluster reconfiguration to work on either environment. We want things to be as push-button as possible.
A few suggestions/concerns I had:
-
The kind config (
kind/kind-config-x86.yamlandkind/kind-config-arm.yaml) needs to be updated withdisableDefaultCNI: trueandpodSubnet: "10.244.0.0/24". Additionally, we'd need a post-creation setup script to runkubectl apply -f calico.yaml. We can reference this script in the README so that someone using kind also sees that running an additional script is needed. Worth thinking about whether replacing kindnet with Calico causes issues for other problems. -
The
strictAffinitypatch is runtime-configurable, so it can go ininject_fault/recover_fault. -
The harder issue is the
/24podSubnet. I think the Calico IPPool CIDR is immutable after cluster creation, so it can't be changed at runtime. We currently use/16. To apply this to the real cluster we will need to editscripts/ansible/setup_cluster.ymltoo to change from/16to/24. A change from/16to/24feels like quite a big change.
One alternative approach that could potentially avoid the /24 requirement entirely is instead of shrinking the cluster CIDR, we could disable the default /16 pool at runtime and create a tiny pool that's easy to exhaust. This way both kind and real cluster can stay on /16. This is just something I am hypothesizing based on the docs I skimmed. So it could be totally unviable. Could you please check it out if it's suitable or not?
Refs:
Disabling an IP pool only prevents new IP address allocations; it does not affect the networking of existing pods.
So, technically pods from other namespaces should be fine.
Feel free to ping us if you have any questions or confusions!
- IPPool + blockSize:32 approach (no /24 CIDR change needed) - strictAffinity and IPPool managed at runtime by inject_fault/recover_fault - Cross-namespace investigation hint added to mitigation prompt - kind configs updated with disableDefaultCNI: true - Added kind/setup_calico_cluster.sh setup script
|
Thanks for the review @Saadmrp1038! The runtime IPPool approach worked! I've updated the PR with the new implementation. No cluster CIDR changes needed — everything is handled at runtime. Details are in the updated PR description. Let me know if you have any questions or need any changes! |
|
@MohammadElsharqawy Is there a reason to use a separate kind config in the script? My idea was to create the cluster with the existing config (with the slight modification of disabling default CNI already baked in), and the script would just install Calico as the CNI. From this point forward, when someone is going to create a kind cluster they just run the script. There's no point in making the config in the script different from the checked-in config. Some nits:
My reasoning is Calico covers everything kindnet does and adds extra features. So all new kind clusters should be using Calico. So, in future if we want to implement a new networking fault that requires these features it would be easier. When you address these ping me. I will start reviewing your PR. |
|
@Saadmrp1038 Adding Now I've renamed the script to The CI smoke test will break again with this change — would you like me to add a Calico install step to Cool, Calico seems a superset of kindnet. |
|
@MohammadElsharqawy Thanks!
That's fine. I will fix it in a separate PR. |
Saadmrp1038
left a comment
There was a problem hiding this comment.
@MohammadElsharqawy The injection and recovery working well. I left some inline comments below. Also some other things:
- The error message I'm seeing is:
Failed to create pod sandbox: rpc error: code = Unknown desc = failed to set up sandbox container "5503345a71 │
│ 4a2a0bacb1b240bbf97f4368bab93725fb04376b2b3a09ac57b65c" network for pod "mongodb-geo-8cb7d9f5b-mmqlx": networkPlugin cni failed to set up pod "mongodb-geo-8cb7d9f5b-mmqlx_hotel-reservation" network: plugin type="calico │
│ " failed (add): cannot allocate new block due to per host block limit
This is a per-host block limit error, not global IP/CIDR exhaustion. The PR description mentions failed to request IPv4 addresses. The fault seems to hit the block limit before actual IP exhaustion. This might cause the agent to score lower in diagnosis since the root cause description mentions CIDR exhaustion but the actual signal says block limit. Is the block size manipulation really needed? After removing the blockSize: 32 part fault injection still works correctly and now expected error message is showing:
rpc error: code = Unknown desc = failed to set up sandbox container "2a0a38347dc569b140745f351635e3a3d7135c9ba73d │
│ ca5e126c41a8e5ecdff1" network for pod "wrk2-job-25vp4": networkPlugin cni failed to set up pod "wrk2-job-25vp4_hotel-reservation" network: plugin type="calico" failed (add): failed to request IPv4 addresses: Assigned 0 │
│ out of 1 requested IPv4 addresses; No more free affine blocks and strict affinity enabled
Regarding the concern about IP reclamation without blockSize: 32; This only matter when the agent scales down batch-worker and this is expected behaviour. Ideal mitigation path (fixing the ippool) actually results in the stuck pods becoming ready almost instantly.
- We should add a global cleanup for the IPPool/strictAffinity manipulation. If someone exits mid-run, the app cleanup won't fix the IPPool state. The entire cluster becomes unusable (no pods can get IPs). I actually ran into this issue while testing. We have something similar for kubelet crash. This needs to run on conductor startup. Check that for reference.
Besides these the fault look great! After you address these I will merge the PR.
… Calico cleanup, add polling, revert cross-namespace hint
|
@Saadmrp1038 Thanks for the detailed feedback!
One thing to flag on the ideal mitigation path — the agent currently doesn't have RBAC to patch IPPools. I checked:
Please take a look and let me know if anything else needs to be addressed! |
I see! I will add it to the RBAC. I tested with claude code and since it directly executes the commands it managed to mitigate the fault just fine. |
|
@MohammadElsharqawy Thanks for addressing all my reviews! The fault is working great now. |
* Add pod_cidr_exhaustion_hotel_reservation benchmark problem * Add pod_cidr_exhaustion_hotel_reservation benchmark problem - IPPool + blockSize:32 approach (no /24 CIDR change needed) - strictAffinity and IPPool managed at runtime by inject_fault/recover_fault - Cross-namespace investigation hint added to mitigation prompt - kind configs updated with disableDefaultCNI: true - Added kind/setup_calico_cluster.sh setup script * Add block affinity cleanup to recover_fault for clean subsequent runs * Revert block affinity cleanup from recover_fault - not needed with blockSize:32 * Replace fixed sleep with polling in inject_fault to fix race condition * Revert polling to simple sleep in inject_fault * Fix CI: revert disableDefaultCNI from kind configs, use inline config in setup script * Add setup_kind_cluster.sh, update kind configs with disableDefaultCNI, fix incident URL * Update README: replace kind create command with setup_kind_cluster.sh * Fix default arch to x86 in setup_kind_cluster.sh * remove blockSize:32, rename namespaces, update root_cause, add global Calico cleanup, add polling, revert cross-namespace hint --------- Co-authored-by: Saad Mohammad Rafid Pial <[email protected]>
* Add pod_cidr_exhaustion_hotel_reservation benchmark problem * Add pod_cidr_exhaustion_hotel_reservation benchmark problem - IPPool + blockSize:32 approach (no /24 CIDR change needed) - strictAffinity and IPPool managed at runtime by inject_fault/recover_fault - Cross-namespace investigation hint added to mitigation prompt - kind configs updated with disableDefaultCNI: true - Added kind/setup_calico_cluster.sh setup script * Add block affinity cleanup to recover_fault for clean subsequent runs * Revert block affinity cleanup from recover_fault - not needed with blockSize:32 * Replace fixed sleep with polling in inject_fault to fix race condition * Revert polling to simple sleep in inject_fault * Fix CI: revert disableDefaultCNI from kind configs, use inline config in setup script * Add setup_kind_cluster.sh, update kind configs with disableDefaultCNI, fix incident URL * Update README: replace kind create command with setup_kind_cluster.sh * Fix default arch to x86 in setup_kind_cluster.sh * remove blockSize:32, rename namespaces, update root_cause, add global Calico cleanup, add polling, revert cross-namespace hint --------- Co-authored-by: Saad Mohammad Rafid Pial <[email protected]>
Add
pod_cidr_exhaustion_hotel_reservationbenchmark problemSummary
This PR adds a new SREGym benchmark problem that simulates a real-world Kubernetes pod CIDR
exhaustion incident. The fault exhausts Calico's IP address pool by flooding the cluster with
a batch workload, causing Hotel Reservation microservice pods to fail with
FailedCreatePodSandBoxand remain stuck inContainerCreating.Real-World GKE IP Exhaustion Incident
A real-world GKE incident demonstrated how a production GKE cluster exhausted a
/16subnet (65,536IPs) far earlier than expected. Although the workload was running thousands of pods, GKE's default networking behavior caused rapid IP consumption.By default, GKE allows a maximum of
110pods per node, but it pre-allocates IPs per node (2×max-pods-per-node = 110, so 220 IPs per node) to reduce IP reuse during pod churn:As a result:
65,536 / 110 ≈ 595nodes65,536 / 256 = 256nodesThe cluster autoscaler could no longer add new nodes, so new pods remained
Pendingindefinitely. One suggested fix was to reduce--max-pods-per-nodeon low-density node pools, freeing subnet space for new nodes.Reference: https://deploy.live/blog/when-gke-ran-out-of-ip-addresses/
Simulation Design
Mapping the Real Incident to a Benchmark
The real GKE incident involved subnet-level IP exhaustion — too many nodes, each pre-allocating a large IP block, leaving no room for new nodes. In our benchmark environment (Kind for local testing, or the SREGym production cluster), we can't replicate node-level pre-allocation (that behavior is GKE-specific and tied to the cloud provider's node provisioning), but we can replicate the same observable failure: pods failing to get IP addresses, causing the entire application to become unavailable.
The key insight is that the failure mode is identical from the application's perspective:
Pendingbecause no node can be scheduled (no IPs for new nodes)ContainerCreatingbecause no IP is available in the poolIn both cases, the root cause is IP exhaustion — more IPs consumed than the pool can provide — and the fix is to reduce IP consumption: reducing node pre-allocation in GKE by lowering
--max-pods-per-node, or scaling down the batch workload in our simulation.Fault Mechanism
We use Calico's IPAM (IP Address Management) to simulate IP exhaustion at runtime. Calico manages IP allocation through IP pools — each pool defines a range of IPs available for pods. By creating a small, exhaustible pool and forcing all new pods to use it, we can reliably trigger IP exhaustion without changing any cluster-wide configuration:
Create a tiny exhaustible pool (
192.168.254.0/26, 64 IPs,blockSize: 32(removed)) — A new IP pool with only 64 IPs is created.blockSize: 32(removed)means each block holds exactly one IP — see WhyblockSize: 32(removed)? below for details. This makes exhaustion precise and recovery immediate.Disable the default
/16IPPool — The cluster's default pool has 65,536 IPs — far too many to exhaust. Disabling it forces all new pod IP allocations to come from the tiny pool only, guaranteeing that the batch workload will exhaust it.Enable
strictAffinity: true— By default, Calico allows nodes to borrow IP blocks from other nodes when their own block is full.strictAffinitydisables this borrowing, so each node is strictly limited to its own allocated blocks. This is essential to guarantee exhaustion — without it, nodes would simply borrow IPs from elsewhere and pods would keep running.Deploy
data-pipeline(60 replicas) indata-processingnamespace — A lightweight batch workload (usingpausecontainers) is deployed in a separate namespace. It consumes 60 of the 64 available IPs, leaving only 4 free — not enough for the Hotel Reservation application. Pods are labeled withpriority: lowandworkload-type: batchto reflect real-world SRE practice of identifying batch workloads as safe to scale down during incidents.Force-delete Hotel Reservation pods — The existing HR pods are deleted so they attempt to reschedule. Since the pool is now exhausted, they cannot get new IP addresses and remain stuck in
ContainerCreating.Why
blockSize: 32(removed)?By default, Calico allocates IPs in blocks of 64 (
blockSize: 26). When a block is assigned to a node, Calico retains the block affinity even after all pods using that block terminate. This means freed IPs are not immediately available — the block stays "owned" by the node until Calico's garbage collector runs (which can take minutes).With
blockSize: 32(removed), each block holds exactly one IP. When the pod using that IP terminates, the entire block is released immediately — there is no partially-used block for Calico to retain affinity over.Practical impact: After the agent scales down
data-pipeline, freed IPs become available instantly. Hotel Reservation pods reschedule and obtain IPs without any delay or RBAC-privileged cleanup. The agent only needs to scale downdata-pipeline— nothing else.Fault Signal
Hotel Reservation pods remain stuck in
ContainerCreatingwith:Fidelity Assessment
This table maps the real GKE incident to our simulation across key dimensions:
/16, 65,536 IPs) pre-allocated per node (220 IPs/node × 256 nodes = exhausted)192.168.254.0/26, 64 IPs) exhausted bydata-pipelinedata-processing/data-pipeline(60 replicas) consuming IPs in a separate namespacePending(no node available)ContainerCreating(no IP available)data-processingdepletes IPs used byhotel-reservation--max-pods-per-nodeto free subnet IPsdata-pipelineto free pool IPsNo more free affine blocks and strict affinity enabledCluster Requirements
This problem requires Calico CNI for two reasons:
strictAffinity, andblockSizecontrols, giving us precise runtime control over IP exhaustion without any cluster-wide configuration changes.Real cluster: Already has Calico installed — no additional setup needed.
Kind cluster: Run the included setup script which replaces kindnet with Calico:
This creates a cluster with
disableDefaultCNI: true(kindnet removed) and installs Calico automatically.Implementation Challenges
Getting a reliable, reproducible IP exhaustion fault that works on both Kind and real clusters without changing cluster configuration took four attempts:
Attempt 1 — kindnet with nodeName pinning (❌ Failed)
kindnet does not enforce pod CIDR boundaries — it allocated IPs beyond the
/24block without error. The fault was unreliable and non-reproducible.Attempt 2 — Calico without strictAffinity (❌ Failed)
Calico allocated 259 IPs from the
/24pool —blockSize: 26is an allocation hint, not a hard limit. WithoutstrictAffinity, nodes borrowed blocks from each other and the pool never truly exhausted.Attempt 3 — Calico with strictAffinity +
/24cluster CIDR (✅ Works but invasive)Produces the authentic Calico error. Diagnosis 78/100. But requires changing the cluster CIDR from
/16to/24— a global change that affects all other benchmark problems and requires updatingscripts/ansible/setup_cluster.yml.Attempt 4 — Create a tiny exhaustible IP pool at runtime with Calico (✅ Final approach)
Disable the default
/16pool at runtime and create a tiny/26pool withblockSize: 32(removed). Works on/16clusters without any cluster CIDR changes.strictAffinity, creating the tiny pool, and re-enabling the default pool are all handled at runtime ininject_fault/recover_fault— no manual cluster reconfiguration needed.data-pipelinepods are labeled withpriority: lowandworkload-type: batchso the agent can identify them as safe to scale down.Mitigation Prompt Addition
Problem
Without any hint, the mitigation agent never ran
kubectl get namespacesorkubectl get pods -A. It stayed entirely within thehotel-reservationnamespace — describing stuck pods, cordoning nodes, and deleting individual pods — never discovering thatdata-processing/data-pipelinewas consuming all available IPs. All early mitigation runs timed out without resolving the fault.This is a cross-namespace reasoning gap: the agent correctly identifies the mechanism (Calico IP exhaustion) but doesn't search for the cause in other namespaces.
Fix
The following hint was added to the default Stratus mitigation system prompt (
clients/stratus/configs/mitigation_agent_prompts.yaml):With this hint, the agent's first action in mitigation was to run
kubectl get namespacesandkubectl get pods -A. It immediately founddata-processing, identifieddata-pipelineas the IP consumer, and scaled it down from 60 to 10 replicas — freeing enough IPs for all Hotel Reservation pods to recover.Agent Evaluation — Stratus + GPT-5
/24CIDRdata-processing/24CIDRdata-processing— stayed inhotel-reservationnamespace onlyblockSize:32+ cross-namespace hint*Mitigation succeeded in practice — all Hotel Reservation pods returned to Running after the agent scaled down
data-pipeline. Thetimed_out: Trueresult is due to the Mac Docker Desktop AlertOracle networking issue (see Infrastructure Notes).Mitigation Analysis
The agent's correct mitigation path (with cross-namespace hint):
kubectl get namespaces+kubectl get pods -A→ findsdata-processingwith 60 replicaskubectl scale deployment data-pipeline -n data-processing --replicas=10✅ (scale down sufficiently, not to 0)blockSize: 32(removed), freed IPs are released immediately as pods terminateFiles Changed
sregym/conductor/problems/pod_cidr_exhaustion_hotel_reservation.py— new problem classsregym/conductor/problems/registry.py— added registry entryclients/stratus/configs/mitigation_agent_prompts.yaml— cross-namespace investigation hintkind/setup_kind_cluster.sh— setup script to create a Kind cluster with Calico CNIInfrastructure Notes
/16clusters — no cluster CIDR changes neededClusterStateOracleandAlertOraclefail due to127.0.0.1vshost.docker.internalnetworking. This is a known Docker Desktoplimitation unrelated to the benchmark problem. Testing on Linux is recommended for
accurate mitigation oracle results.
Related Work
/validate-problem pod_cidr_exhaustion_hotel_reservation