Skip to content

Add pod_cidr_exhaustion_hotel_reservation benchmark problem - #774

Merged
Saadmrp1038 merged 13 commits into
SREGym:mainfrom
MohammadElsharqawy:main
May 27, 2026
Merged

Saadmrp1038 merged 13 commits into
SREGym:mainfrom
MohammadElsharqawy:main

Conversation

@MohammadElsharqawy

@MohammadElsharqawy MohammadElsharqawy commented May 24, 2026 •

Copy link
Copy Markdown
Contributor

Add pod_cidr_exhaustion_hotel_reservation benchmark problem

Summary

This PR adds a new SREGym benchmark problem that simulates a real-world Kubernetes pod CIDR
exhaustion incident. The fault exhausts Calico's IP address pool by flooding the cluster with
a batch workload, causing Hotel Reservation microservice pods to fail with
FailedCreatePodSandBox and remain stuck in ContainerCreating.


Real-World GKE IP Exhaustion Incident

A real-world GKE incident demonstrated how a production GKE cluster exhausted a /16 subnet (65,536 IPs) far earlier than expected. Although the workload was running thousands of pods, GKE's default networking behavior caused rapid IP consumption.

By default, GKE allows a maximum of 110 pods per node, but it pre-allocates IPs per node (2× max-pods-per-node = 110, so 220 IPs per node) to reduce IP reuse during pod churn:

"By having approximately twice as many available IP addresses as the number of pods that can be created on a node, Kubernetes is able to mitigate IP address reuse as Pods are added to and removed from a node."

As a result:

"This meant that each of our 110 maximum pod nodes was consuming 256 IP addresses, which aligns exactly with the observed behaviour: 65'536 / 256 = 256."

  • Expected: 65,536 / 110 ≈ 595 nodes
  • Actual: 65,536 / 256 = 256 nodes

The cluster autoscaler could no longer add new nodes, so new pods remained Pending indefinitely. One suggested fix was to reduce --max-pods-per-node on low-density node pools, freeing subnet space for new nodes.

Reference: https://deploy.live/blog/when-gke-ran-out-of-ip-addresses/


Simulation Design

Mapping the Real Incident to a Benchmark

The real GKE incident involved subnet-level IP exhaustion — too many nodes, each pre-allocating a large IP block, leaving no room for new nodes. In our benchmark environment (Kind for local testing, or the SREGym production cluster), we can't replicate node-level pre-allocation (that behavior is GKE-specific and tied to the cloud provider's node provisioning), but we can replicate the same observable failure: pods failing to get IP addresses, causing the entire application to become unavailable.

The key insight is that the failure mode is identical from the application's perspective:

  • Real incident: new pods stay Pending because no node can be scheduled (no IPs for new nodes)
  • Simulation: new pods stay ContainerCreating because no IP is available in the pool

In both cases, the root cause is IP exhaustion — more IPs consumed than the pool can provide — and the fix is to reduce IP consumption: reducing node pre-allocation in GKE by lowering --max-pods-per-node, or scaling down the batch workload in our simulation.

Fault Mechanism

We use Calico's IPAM (IP Address Management) to simulate IP exhaustion at runtime. Calico manages IP allocation through IP pools — each pool defines a range of IPs available for pods. By creating a small, exhaustible pool and forcing all new pods to use it, we can reliably trigger IP exhaustion without changing any cluster-wide configuration:

  1. Create a tiny exhaustible pool (192.168.254.0/26, 64 IPs, blockSize: 32 (removed)) — A new IP pool with only 64 IPs is created. blockSize: 32 (removed) means each block holds exactly one IP — see Why blockSize: 32 (removed)? below for details. This makes exhaustion precise and recovery immediate.

  2. Disable the default /16 IPPool — The cluster's default pool has 65,536 IPs — far too many to exhaust. Disabling it forces all new pod IP allocations to come from the tiny pool only, guaranteeing that the batch workload will exhaust it.

  3. Enable strictAffinity: true — By default, Calico allows nodes to borrow IP blocks from other nodes when their own block is full. strictAffinity disables this borrowing, so each node is strictly limited to its own allocated blocks. This is essential to guarantee exhaustion — without it, nodes would simply borrow IPs from elsewhere and pods would keep running.

  4. Deploy data-pipeline (60 replicas) in data-processing namespace — A lightweight batch workload (using pause containers) is deployed in a separate namespace. It consumes 60 of the 64 available IPs, leaving only 4 free — not enough for the Hotel Reservation application. Pods are labeled with priority: low and workload-type: batch to reflect real-world SRE practice of identifying batch workloads as safe to scale down during incidents.

  5. Force-delete Hotel Reservation pods — The existing HR pods are deleted so they attempt to reschedule. Since the pool is now exhausted, they cannot get new IP addresses and remain stuck in ContainerCreating.

Why blockSize: 32 (removed)?

By default, Calico allocates IPs in blocks of 64 (blockSize: 26). When a block is assigned to a node, Calico retains the block affinity even after all pods using that block terminate. This means freed IPs are not immediately available — the block stays "owned" by the node until Calico's garbage collector runs (which can take minutes).

With blockSize: 32 (removed), each block holds exactly one IP. When the pod using that IP terminates, the entire block is released immediately — there is no partially-used block for Calico to retain affinity over.

Practical impact: After the agent scales down data-pipeline, freed IPs become available instantly. Hotel Reservation pods reschedule and obtain IPs without any delay or RBAC-privileged cleanup. The agent only needs to scale down data-pipeline — nothing else.

Fault Signal

Hotel Reservation pods remain stuck in ContainerCreating with:

Failed to create pod sandbox: plugin type="calico" failed (add):
failed to request IPv4 addresses: Assigned 0 out of 1 requested IPv4 addresses;
No more free affine blocks and strict affinity enabled

Note: Earlier versions used blockSize: 32 which produced a different error (cannot allocate new block due to per host block limit). This has been removed — the correct error above is now shown.

Fidelity Assessment

This table maps the real GKE incident to our simulation across key dimensions:

Dimension Real GKE Incident This Simulation Fidelity
IP pool exhausted VPC subnet (/16, 65,536 IPs) pre-allocated per node (220 IPs/node × 256 nodes = exhausted) Calico tiny pool (192.168.254.0/26, 64 IPs) exhausted by data-pipeline Same mechanism — a bounded IP pool runs out
What triggers failure No IPs left for new nodes → autoscaler can't add nodes No IPs left in pool → Calico IPAM can't assign pod IPs Same root cause — IP exhaustion blocks workload scheduling
Unexpected consumer GKE node pre-allocation consuming 2× expected IPs per node data-processing/data-pipeline (60 replicas) consuming IPs in a separate namespace Both involve a consumer the operator didn't anticipate
Affected workload state New pods stay Pending (no node available) Existing pods stuck in ContainerCreating (no IP available) Slightly different — Pending vs ContainerCreating, but same observable effect: application down
Cross-namespace causation Node pool config affects all namespaces cluster-wide data-processing depletes IPs used by hotel-reservation Fault originates outside the affected application's namespace
Fix action Reduce --max-pods-per-node to free subnet IPs Scale down data-pipeline to free pool IPs Same pattern — reduce the consuming workload
Error message Scheduler: no nodes available Calico: No more free affine blocks and strict affinity enabled Different layer — scheduler vs CNI, but both mean "no IP available"

Cluster Requirements

This problem requires Calico CNI for two reasons:

  • kindnet (Kind's default CNI) does not enforce IP boundaries — pods can get IPs beyond their allocated block without error, making reliable IP exhaustion impossible (confirmed in Attempt 1).
  • Calico provides a full IPAM system with IPPools, strictAffinity, and blockSize controls, giving us precise runtime control over IP exhaustion without any cluster-wide configuration changes.

Real cluster: Already has Calico installed — no additional setup needed.

Kind cluster: Run the included setup script which replaces kindnet with Calico:

bash kind/setup_kind_cluster.sh arm   # for ARM (Apple Silicon)
bash kind/setup_kind_cluster.sh x86   # for x86

This creates a cluster with disableDefaultCNI: true (kindnet removed) and installs Calico automatically.

Implementation Challenges

Getting a reliable, reproducible IP exhaustion fault that works on both Kind and real clusters without changing cluster configuration took four attempts:

Attempt 1 — kindnet with nodeName pinning (❌ Failed)
kindnet does not enforce pod CIDR boundaries — it allocated IPs beyond the /24 block without error. The fault was unreliable and non-reproducible.

Attempt 2 — Calico without strictAffinity (❌ Failed)
Calico allocated 259 IPs from the /24 pool — blockSize: 26 is an allocation hint, not a hard limit. Without strictAffinity, nodes borrowed blocks from each other and the pool never truly exhausted.

Attempt 3 — Calico with strictAffinity + /24 cluster CIDR (✅ Works but invasive)
Produces the authentic Calico error. Diagnosis 78/100. But requires changing the cluster CIDR from /16 to /24 — a global change that affects all other benchmark problems and requires updating scripts/ansible/setup_cluster.yml.

Attempt 4 — Create a tiny exhaustible IP pool at runtime with Calico (✅ Final approach)
Disable the default /16 pool at runtime and create a tiny /26 pool with blockSize: 32 (removed). Works on /16 clusters without any cluster CIDR changes. strictAffinity, creating the tiny pool, and re-enabling the default pool are all handled at runtime in inject_fault/recover_fault — no manual cluster reconfiguration needed. data-pipeline pods are labeled with priority: low and workload-type: batch so the agent can identify them as safe to scale down.



Mitigation Prompt Addition

Problem

Without any hint, the mitigation agent never ran kubectl get namespaces or kubectl get pods -A. It stayed entirely within the hotel-reservation namespace — describing stuck pods, cordoning nodes, and deleting individual pods — never discovering that data-processing/data-pipeline was consuming all available IPs. All early mitigation runs timed out without resolving the fault.

This is a cross-namespace reasoning gap: the agent correctly identifies the mechanism (Calico IP exhaustion) but doesn't search for the cause in other namespaces.

Fix

The following hint was added to the default Stratus mitigation system prompt (clients/stratus/configs/mitigation_agent_prompts.yaml):

**IMPORTANT - Cross-namespace investigation:** Faults are often caused by workloads in a
DIFFERENT namespace than the affected application. Always start by running
`kubectl get namespaces` and `kubectl get pods -A` to identify ALL workloads running in
the cluster. A workload in another namespace may be consuming shared cluster resources
(CPU, memory, IP addresses, storage) that your target application needs. Do not limit
your investigation to the application's own namespace.

With this hint, the agent's first action in mitigation was to run kubectl get namespaces and kubectl get pods -A. It immediately found data-processing, identified data-pipeline as the IP consumer, and scaled it down from 60 to 10 replicas — freeing enough IPs for all Hotel Reservation pods to recover.

Note: After editing mitigation_agent_prompts.yaml, the Stratus Docker image must be rebuilt for the changes to take effect:

uv run python main.py --agent stratus --force-build

Agent Evaluation — Stratus + GPT-5

Run Approach Diagnosis TTL Mitigation Notes
Run 1 Unmodified prompt, /24 CIDR 78/100 ✅ 127.5s ❌ Never found data-processing
Run 2 Cross-namespace hint added, /24 CIDR 78/100 ✅ 166.6s ❌ (timeout) AlertOracle polling exhausted budget
Run 3 Tiny Calico pool (no cross-namespace hint) 78/100 ✅ 70.7s ❌ Agent never found data-processing — stayed in hotel-reservation namespace only
Run 4 Tiny Calico pool + blockSize:32 + cross-namespace hint 78/100 ✅ 91.2s ✅* Agent scaled data-pipeline 60→10, all HR pods recovered. Oracle timed out due to Mac AlertOracle issue, not agent failure

*Mitigation succeeded in practice — all Hotel Reservation pods returned to Running after the agent scaled down data-pipeline. The timed_out: True result is due to the Mac Docker Desktop AlertOracle networking issue (see Infrastructure Notes).

Note on Run 3: Although the agent correctly found data-processing in Run 3, we did not continue with that approach (the /24 CIDR) because making it work on the real cluster would require shrinking the cluster CIDR from /16 to /24 — a global change that affects all other benchmark problems. Run 4 solves this by creating a tiny pool at runtime, avoiding any cluster-level changes.

Mitigation Analysis

The agent's correct mitigation path (with cross-namespace hint):

  1. kubectl get namespaces + kubectl get pods -A → finds data-processing with 60 replicas
  2. kubectl scale deployment data-pipeline -n data-processing --replicas=10 ✅ (scale down sufficiently, not to 0)
  3. With blockSize: 32 (removed), freed IPs are released immediately as pods terminate
  4. All Hotel Reservation pods reschedule, obtain IPs, and return to Running ✅ — full application recovery confirmed

Files Changed

  • sregym/conductor/problems/pod_cidr_exhaustion_hotel_reservation.py — new problem class
  • sregym/conductor/problems/registry.py — added registry entry
  • clients/stratus/configs/mitigation_agent_prompts.yaml — cross-namespace investigation hint
  • kind/setup_kind_cluster.sh — setup script to create a Kind cluster with Calico CNI

Infrastructure Notes

  • Requires Calico CNI (replace kindnet in kind configs)
  • Works on /16 clusters — no cluster CIDR changes needed
  • macOS with Docker Desktop: ClusterStateOracle and AlertOracle fail due to
    127.0.0.1 vs host.docker.internal networking. This is a known Docker Desktop
    limitation unrelated to the benchmark problem. Testing on Linux is recommended for
    accurate mitigation oracle results.

Related Work

data-pipeline pods are labeled with priority: low and workload-type: batch — following real-world SRE practice of labeling batch workloads so they are identifiable as safe to scale down during incidents.


/validate-problem pod_cidr_exhaustion_hotel_reservation

@Saadmrp1038 Saadmrp1038 added the problem Adding a new problem to the benchmark label May 24, 2026
@Saadmrp1038
Saadmrp1038 self-requested a review May 24, 2026 04:18

@Saadmrp1038 Saadmrp1038 left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@MohammadElsharqawy Thanks for the PR! This fault is quite interesting and the code looks great. One thing, we'll need to automate the config-related changes.

The "Cluster Requirements" section lists manual steps that need to happen outside the PR:

# kind-config.yaml additions
networking:
  disableDefaultCNI: true
  podSubnet: "10.244.0.0/24"
kubectl apply -f https://raw.githubusercontent.com/projectcalico/calico/v3.27.0/manifests/calico.yaml
kubectl patch ipamconfig default --type=merge -p '{"spec":{"strictAffinity":true}}'
rm ~/cache_dir/cluster_baseline_state.json

These are not added in the PR. The kind cluster currently uses kindnet (no Calico), and the real cluster has Calico but with a /16 pool. So this problem requires manual cluster reconfiguration to work on either environment. We want things to be as push-button as possible.

A few suggestions/concerns I had:

  1. The kind config (kind/kind-config-x86.yaml and kind/kind-config-arm.yaml) needs to be updated with disableDefaultCNI: true and podSubnet: "10.244.0.0/24". Additionally, we'd need a post-creation setup script to run kubectl apply -f calico.yaml. We can reference this script in the README so that someone using kind also sees that running an additional script is needed. Worth thinking about whether replacing kindnet with Calico causes issues for other problems.

  2. The strictAffinity patch is runtime-configurable, so it can go in inject_fault/recover_fault.

  3. The harder issue is the /24 podSubnet. I think the Calico IPPool CIDR is immutable after cluster creation, so it can't be changed at runtime. We currently use /16. To apply this to the real cluster we will need to edit scripts/ansible/setup_cluster.yml too to change from /16 to /24. A change from /16 to /24 feels like quite a big change.

One alternative approach that could potentially avoid the /24 requirement entirely is instead of shrinking the cluster CIDR, we could disable the default /16 pool at runtime and create a tiny pool that's easy to exhaust. This way both kind and real cluster can stay on /16. This is just something I am hypothesizing based on the docs I skimmed. So it could be totally unviable. Could you please check it out if it's suitable or not?

Refs:

Disabling an IP pool only prevents new IP address allocations; it does not affect the networking of existing pods.

So, technically pods from other namespaces should be fine.


Feel free to ping us if you have any questions or confusions!

- IPPool + blockSize:32 approach (no /24 CIDR change needed)
- strictAffinity and IPPool managed at runtime by inject_fault/recover_fault
- Cross-namespace investigation hint added to mitigation prompt
- kind configs updated with disableDefaultCNI: true
- Added kind/setup_calico_cluster.sh setup script
@MohammadElsharqawy

Copy link
Copy Markdown
Contributor Author

Thanks for the review @Saadmrp1038! The runtime IPPool approach worked! I've updated the PR with the new implementation. No cluster CIDR changes needed — everything is handled at runtime. Details are in the updated PR description. Let me know if you have any questions or need any changes!

@Saadmrp1038

Saadmrp1038 commented May 26, 2026 •

Copy link
Copy Markdown
Collaborator

@MohammadElsharqawy Is there a reason to use a separate kind config in the script?

My idea was to create the cluster with the existing config (with the slight modification of disabling default CNI already baked in), and the script would just install Calico as the CNI. From this point forward, when someone is going to create a kind cluster they just run the script. There's no point in making the config in the script different from the checked-in config.

Some nits:

  1. Let's name the script setup_kind_cluster.sh since this will be the new creation script for kind clusters.
  2. Let's not use a custom name for the kind cluster. Some of our problems depend on name matching to detect kind docker containers, so renaming would break those faults.
  3. In the README we can replace the current kind creation command with this new script.

My reasoning is Calico covers everything kindnet does and adds extra features. So all new kind clusters should be using Calico. So, in future if we want to implement a new networking fault that requires these features it would be easier.


When you address these ping me. I will start reviewing your PR.

@MohammadElsharqawy

Copy link
Copy Markdown
Contributor Author

@Saadmrp1038 Adding disableDefaultCNI: true to the kind configs breaks the CI smoke test since the CI creates the cluster but doesn't install Calico, so I reverted it and created it inline in the script to pass smoke test.

Now I've renamed the script to setup_kind_cluster.sh, simplified it to use the kind config files directly, and added disableDefaultCNI: true back to both kind configs as you suggested. I've also updated the main README to replace the old kind create cluster commands with the new script.

The CI smoke test will break again with this change — would you like me to add a Calico install step to .github/workflows/smoke-test.yml as part of this PR?

Cool, Calico seems a superset of kindnet.

@Saadmrp1038

Copy link
Copy Markdown
Collaborator

@MohammadElsharqawy Thanks!

The CI smoke test will break again with this change — would you like me to add a Calico install step to .github/workflows/smoke-test.yml as part of this PR?

That's fine. I will fix it in a separate PR.

Comment thread kind/setup_kind_cluster.sh

@Saadmrp1038 Saadmrp1038 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@MohammadElsharqawy The injection and recovery working well. I left some inline comments below. Also some other things:

  1. The error message I'm seeing is:
 Failed to create pod sandbox: rpc error: code = Unknown desc = failed to set up sandbox container "5503345a71 │
│ 4a2a0bacb1b240bbf97f4368bab93725fb04376b2b3a09ac57b65c" network for pod "mongodb-geo-8cb7d9f5b-mmqlx": networkPlugin cni failed to set up pod "mongodb-geo-8cb7d9f5b-mmqlx_hotel-reservation" network: plugin type="calico │
│ " failed (add): cannot allocate new block due to per host block limit

This is a per-host block limit error, not global IP/CIDR exhaustion. The PR description mentions failed to request IPv4 addresses. The fault seems to hit the block limit before actual IP exhaustion. This might cause the agent to score lower in diagnosis since the root cause description mentions CIDR exhaustion but the actual signal says block limit. Is the block size manipulation really needed? After removing the blockSize: 32 part fault injection still works correctly and now expected error message is showing:

 rpc error: code = Unknown desc = failed to set up sandbox container "2a0a38347dc569b140745f351635e3a3d7135c9ba73d │
│ ca5e126c41a8e5ecdff1" network for pod "wrk2-job-25vp4": networkPlugin cni failed to set up pod "wrk2-job-25vp4_hotel-reservation" network: plugin type="calico" failed (add): failed to request IPv4 addresses: Assigned 0 │
│  out of 1 requested IPv4 addresses; No more free affine blocks and strict affinity enabled

Regarding the concern about IP reclamation without blockSize: 32; This only matter when the agent scales down batch-worker and this is expected behaviour. Ideal mitigation path (fixing the ippool) actually results in the stuck pods becoming ready almost instantly.

  1. We should add a global cleanup for the IPPool/strictAffinity manipulation. If someone exits mid-run, the app cleanup won't fix the IPPool state. The entire cluster becomes unusable (no pods can get IPs). I actually ran into this issue while testing. We have something similar for kubelet crash. This needs to run on conductor startup. Check that for reference.

Besides these the fault look great! After you address these I will merge the PR.

Comment thread clients/stratus/configs/mitigation_agent_prompts.yaml Outdated
Comment thread kind/setup_kind_cluster.sh Outdated
Comment thread sregym/conductor/problems/pod_cidr_exhaustion_hotel_reservation.py Outdated
Comment thread sregym/conductor/problems/pod_cidr_exhaustion_hotel_reservation.py Outdated
Comment thread sregym/conductor/problems/pod_cidr_exhaustion_hotel_reservation.py
… Calico cleanup, add polling, revert cross-namespace hint
@MohammadElsharqawy

MohammadElsharqawy commented May 27, 2026 •

Copy link
Copy Markdown
Contributor Author

@Saadmrp1038 Thanks for the detailed feedback!

  1. Makes sense! I removed blockSize: 32.

One thing to flag on the ideal mitigation path — the agent currently doesn't have RBAC to patch IPPools. I checked:

kubectl auth can-i patch ippools.crd.projectcalico.org --as=system:serviceaccount:sregym:mcp-server
no
  1. Done. I also cleaned up the runtime-created namespace

Please take a look and let me know if anything else needs to be addressed!

@Saadmrp1038

Copy link
Copy Markdown
Collaborator

kubectl auth can-i patch ippools.crd.projectcalico.org --as=system:serviceaccount:sregym:mcp-server
no

I see! I will add it to the RBAC. I tested with claude code and since it directly executes the commands it managed to mitigate the fault just fine.

@Saadmrp1038

Copy link
Copy Markdown
Collaborator

@MohammadElsharqawy Thanks for addressing all my reviews! The fault is working great now.
Congrats on your first PR merged here 🎉

@Saadmrp1038
Saadmrp1038 merged commit cea3da3 into SREGym:main May 27, 2026
3 of 5 checks passed
ermias19 pushed a commit to ermias19/SREGym that referenced this pull request Jun 5, 2026
* Add pod_cidr_exhaustion_hotel_reservation benchmark problem

* Add pod_cidr_exhaustion_hotel_reservation benchmark problem

- IPPool + blockSize:32 approach (no /24 CIDR change needed)
- strictAffinity and IPPool managed at runtime by inject_fault/recover_fault
- Cross-namespace investigation hint added to mitigation prompt
- kind configs updated with disableDefaultCNI: true
- Added kind/setup_calico_cluster.sh setup script

* Add block affinity cleanup to recover_fault for clean subsequent runs

* Revert block affinity cleanup from recover_fault - not needed with blockSize:32

* Replace fixed sleep with polling in inject_fault to fix race condition

* Revert polling to simple sleep in inject_fault

* Fix CI: revert disableDefaultCNI from kind configs, use inline config in setup script

* Add setup_kind_cluster.sh, update kind configs with disableDefaultCNI, fix incident URL

* Update README: replace kind create command with setup_kind_cluster.sh

* Fix default arch to x86 in setup_kind_cluster.sh

* remove blockSize:32, rename namespaces, update root_cause, add global Calico cleanup, add polling, revert cross-namespace hint

---------

Co-authored-by: Saad Mohammad Rafid Pial <[email protected]>
ermias19 pushed a commit to ermias19/SREGym that referenced this pull request Jun 5, 2026
* Add pod_cidr_exhaustion_hotel_reservation benchmark problem

* Add pod_cidr_exhaustion_hotel_reservation benchmark problem

- IPPool + blockSize:32 approach (no /24 CIDR change needed)
- strictAffinity and IPPool managed at runtime by inject_fault/recover_fault
- Cross-namespace investigation hint added to mitigation prompt
- kind configs updated with disableDefaultCNI: true
- Added kind/setup_calico_cluster.sh setup script

* Add block affinity cleanup to recover_fault for clean subsequent runs

* Revert block affinity cleanup from recover_fault - not needed with blockSize:32

* Replace fixed sleep with polling in inject_fault to fix race condition

* Revert polling to simple sleep in inject_fault

* Fix CI: revert disableDefaultCNI from kind configs, use inline config in setup script

* Add setup_kind_cluster.sh, update kind configs with disableDefaultCNI, fix incident URL

* Update README: replace kind create command with setup_kind_cluster.sh

* Fix default arch to x86 in setup_kind_cluster.sh

* remove blockSize:32, rename namespaces, update root_cause, add global Calico cleanup, add polling, revert cross-namespace hint

---------

Co-authored-by: Saad Mohammad Rafid Pial <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

problem Adding a new problem to the benchmark

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants