Skip to content

Add service wrong pod selection benchmark for Hotel Reservation - #757

Merged
Saadmrp1038 merged 2 commits into
SREGym:mainfrom
jackj6441:jack-service-wrong-pod-selection
May 19, 2026
Merged

Saadmrp1038 merged 2 commits into
SREGym:mainfrom
jackj6441:jack-service-wrong-pod-selection

Conversation

@jackj6441

@jackj6441 jackj6441 commented May 18, 2026 •

Copy link
Copy Markdown
Contributor

The Real-World Failure Story

This problem is inspired by a Kubernetes Service selector misconfiguration where a Service selector is too broad or labels are accidentally shared across unrelated workloads.

In this kind of failure, the Service still has endpoints, but the endpoint list is polluted with pods that do not actually serve the expected traffic. This is harder to diagnose than a simple empty-endpoint failure because the pods may all be Running, the Service may look healthy at first glance, and the endpoint list is not empty.
Reference:
https://medium.com/emburse/intermittent-connection-refused-errors-after-new-service-deployment-in-kubernetes-22dd5feae1ad
However, some requests can still be routed to unintended pods, which may reject connections, time out, or return invalid responses.

How We Simulate the Failure on SREGym

This PR adds a new SREGym benchmark problem:

service_wrong_pod_selection_hotel_reservation

The problem uses the Hotel Reservation application and targets the frontend Service.

Under normal behavior, the frontend Service selects frontend pods using:

io.kompose.service: frontend

The injected fault changes the Service selection behavior by:

  1. Adding sregym.io/frontend-route=true to the intended frontend deployment.
  2. Adding sregym.io/frontend-route=true to the unintended search deployment.
  3. Replacing the frontend Service selector with:
sregym.io/frontend-route: "true"

After the fault is injected, the frontend Service selects both the real frontend pod and the unintended search pod.

frontend Service
├── frontend pod  correct endpoint
└── search pod    wrong endpoint

The frontend Service routes traffic to targetPort 5000, while the search container listens on port 8082. Therefore, traffic routed to the search pod through the frontend Service can fail.

The Problem Runtime Behavior

After fault injection:

  • The Hotel Reservation application deploys successfully.
  • Application pods can remain Running and Ready.
  • The frontend Service still has endpoints.
  • The frontend EndpointSlice is non-empty but contains both frontend and search pods.
  • Requests through frontend:5000 may fail because some traffic can be routed to the search pod, which does not serve frontend traffic on port 5000.
  • The failure is caused by Service endpoint pollution, not an empty endpoint set or pod crash.

The deterministic failure signal is EndpointSlice pollution. Curl failures are useful supporting evidence, but the mitigation oracle does not depend on the exact curl failure rate.

Useful debugging commands:

kubectl get svc frontend -n hotel-reservation -o yaml

kubectl get endpointslices -n hotel-reservation \
  -l kubernetes.io/service-name=frontend \
  -o yaml

kubectl get pods -n hotel-reservation --show-labels

kubectl describe svc frontend -n hotel-reservation

Expected broken-state evidence:

frontend Service selector:
{"sregym.io/frontend-route":"true"}

frontend EndpointSlice contains:
pod=frontend-...
pod=search-...

The Agent Behavior

A successful agent should diagnose that the frontend Service is selecting unintended pods because its selector was broadened to:

sregym.io/frontend-route: "true"

The agent should identify that both the intended frontend pod and the unintended search pod match this selector.

The root cause is the frontend Service selector misconfiguration. It is not a pod crash, not an empty endpoint set, and not a failure inside the search service itself.

Expected diagnosis:

The frontend Service selector is too broad. It selects both frontend pods and the search pod, causing the frontend Service endpoints to be polluted. The search pod does not serve frontend traffic on targetPort 5000, so traffic through the frontend Service can fail.

Expected mitigation is to restore the frontend Service selector to its original value:

io.kompose.service: frontend

The mitigation should result in the frontend EndpointSlice containing only frontend pods.

The mitigation oracle checks that:

  • The frontend Service selector is exactly restored to {"io.kompose.service": "frontend"}.
  • Frontend endpoints are non-empty.
  • All selected endpoint pods are actual frontend pods.
  • No unintended search endpoint remains.

In manual validation, the diagnosis submission passed the LLM-as-a-judge evaluation with 100/100, and the mitigation oracle passed after restoring the frontend Service selector and removing the polluted endpoint.

Validation

Local compile check passed:

PYTHONPATH=. uv run python -m py_compile \
  sregym/conductor/problems/service_wrong_pod_selection_hotel_reservation.py \
  sregym/conductor/oracles/wrong_pod_selection_mitigation.py \
  sregym/conductor/problems/registry.py

Registry check passed:

True

End-to-end SREGym validation passed:

  • The problem started successfully.
  • Fault injection succeeded.
  • The frontend Service selector changed to {"sregym.io/frontend-route":"true"}.
  • The frontend EndpointSlice contained both frontend and search pods.
  • Curl to frontend:5000 showed connection failures as supporting runtime evidence.
  • Diagnosis passed with 100/100.
  • Mitigation restored the frontend Service selector.
  • The frontend EndpointSlice contained only frontend pods.
  • The mitigation oracle passed.

Files Changed

Problem List.md
sregym/conductor/problems/registry.py
sregym/conductor/tasklist.yml.example
sregym/conductor/problems/service_wrong_pod_selection_hotel_reservation.py
sregym/conductor/oracles/wrong_pod_selection_mitigation.py

@tianyin
tianyin requested a review from Saadmrp1038 May 18, 2026 05:16
@tianyin

tianyin commented May 18, 2026

Copy link
Copy Markdown
Contributor

@jackj6441 This is fast. I let @Saadmrp1038 to review it.

The Real-World Failure Story

Can you put the link of the real-world failures there as the evidence support that it is "real world"?

@Saadmrp1038 Saadmrp1038 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the effort! I left some inline comments below, please ping me when you've addressed them.

I also had a couple of queries:

  1. This looks quite similar to wrong_service_selector. Can you explain what's different here?

  2. If possible can you run this problem end-to-end and drop the agent trace here? For example to run on the stratus agent you can:

uv run python main.py --problem service_wrong_pod_selection_hotel_reservation --agent stratus

Please feel free to ping us if you have any questions or confusions!

Comment thread sregym/conductor/problems/service_wrong_pod_selection_hotel_reservation.py Outdated
Comment thread sregym/conductor/problems/service_wrong_pod_selection_hotel_reservation.py Outdated
Comment thread sregym/conductor/problems/service_wrong_pod_selection_hotel_reservation.py Outdated
Comment thread sregym/conductor/oracles/wrong_pod_selection_mitigation.py Outdated
@Saadmrp1038 Saadmrp1038 added the problem Adding a new problem to the benchmark label May 18, 2026
@jackj6441

Copy link
Copy Markdown
Contributor Author

@jackj6441 This is fast. I let @Saadmrp1038 to review it.

The Real-World Failure Story

Can you put the link of the real-world failures there as the evidence support that it is "real world"?

Thanks! I added the real-world failure link as evidence in the PR description. https://medium.com/emburse/intermittent-connection-refused-errors-after-new-service-deployment-in-kubernetes-22dd5feae1ad

@jackj6441

Copy link
Copy Markdown
Contributor Author

Thanks for the effort! I left some inline comments below, please ping me when you've addressed them.

I also had a couple of queries:

  1. This looks quite similar to wrong_service_selector. Can you explain what's different here?
  2. If possible can you run this problem end-to-end and drop the agent trace here? For example to run on the stratus agent you can:
uv run python main.py --problem service_wrong_pod_selection_hotel_reservation --agent stratus

Please feel free to ping us if you have any questions or confusions!

  1. I agree that they are similar, the key difference is the endpoint state:
  • wrong_service_selector: the Service selector does not match the intended pods, so the Service has no ready endpoints.
  • service_wrong_pod_selection_hotel_reservation: the Service still has ready endpoints, but the EndpointSlice is polluted with both the intended frontend pod and the unintended search pod.

So this problem models a broader-selector failure where the Service can look partially healthy because endpoints are non-empty, but some traffic is routed to the wrong pod.

  1. I also ran this locally with Stratus:

The problem deployed successfully and reached the intended injected state. The trace confirmed that the frontend Service selector was changed to service-route=frontend, causing the Service to select both the intended frontend pod and the unintended search pod.

Stratus did not solve the intended fault in this run. It focused on unrelated startup noise around rate / mongodb-rate, patched the rate Deployment with an initContainer, and submitted that as mitigation. The benchmark oracle correctly rejected the attempt because the polluted frontend Service endpoints were not fixed. The run keep retrying in worng direction so I stopped the loop with Ctrl-C.
deployment.apps/rate patched

@Saadmrp1038

Saadmrp1038 commented May 19, 2026 •

Copy link
Copy Markdown
Collaborator

@jackj6441 the changes look great!

@tianyin

tianyin commented May 19, 2026

Copy link
Copy Markdown
Contributor

@Saadmrp1038 We should try to merge it and make it an official problem of the benchmark.

Certainly, please keep giving feedback to help @jackj6441 improve the PR till the quality you want.

@Saadmrp1038

Copy link
Copy Markdown
Collaborator

The code looks good. I also ran the problem and everything seems to be working correctly. I am merging the PR.
Congrats on the first PR merged here @jackj6441 🎉

@Saadmrp1038
Saadmrp1038 merged commit f432ae0 into SREGym:main May 19, 2026
2 checks passed
Rahuldrabit pushed a commit to Rahuldrabit/SREGym that referenced this pull request Sep 28, 2026
…ery (SREGym#565)

* fix: Return HTTP response immediately without waiting for fault recovery

This change prevents the submission API from blocking while performing slow operations like fault injection/recovery and app teardown. These operations can take minutes when running multiple problems concurrently.

Changes:

- Add async cleanup method that performs fault recovery, undeploy, and cluster reconciliation in the background

- Update _finish_problem to set state to tearing_down immediately and schedule async cleanup

- Modify submit() to capture results snapshot and return immediately

- Prevent noise restart during teardown state

- Add tearing_down to possible stages in API documentation

Fixes SREGym#244

Co-authored-by: Jackson Clark <[email protected]>

* free event loop from cleanup

* Use force exit to shutdown api server

* suppress irrelevant error

* Return without evaluating

* Close SREGym#757

* Mount .aws directory

* taking over fleetcast

* update fleetcast owner

---------

Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: Jackson Clark <[email protected]>
Co-authored-by: Yiming Su <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

problem Adding a new problem to the benchmark

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants