Repository navigation
Add service wrong pod selection benchmark for Hotel Reservation - #757
Conversation
|
@jackj6441 This is fast. I let @Saadmrp1038 to review it.
Can you put the link of the real-world failures there as the evidence support that it is "real world"? |
Saadmrp1038
left a comment
There was a problem hiding this comment.
Thanks for the effort! I left some inline comments below, please ping me when you've addressed them.
I also had a couple of queries:
-
This looks quite similar to
wrong_service_selector. Can you explain what's different here? -
If possible can you run this problem end-to-end and drop the agent trace here? For example to run on the stratus agent you can:
uv run python main.py --problem service_wrong_pod_selection_hotel_reservation --agent stratus
Please feel free to ping us if you have any questions or confusions!
Thanks! I added the real-world failure link as evidence in the PR description. https://medium.com/emburse/intermittent-connection-refused-errors-after-new-service-deployment-in-kubernetes-22dd5feae1ad |
So this problem models a broader-selector failure where the Service can look partially healthy because endpoints are non-empty, but some traffic is routed to the wrong pod.
The problem deployed successfully and reached the intended injected state. The trace confirmed that the Stratus did not solve the intended fault in this run. It focused on unrelated startup noise around |
|
@jackj6441 the changes look great! |
|
@Saadmrp1038 We should try to merge it and make it an official problem of the benchmark. Certainly, please keep giving feedback to help @jackj6441 improve the PR till the quality you want. |
|
The code looks good. I also ran the problem and everything seems to be working correctly. I am merging the PR. |
…ery (SREGym#565) * fix: Return HTTP response immediately without waiting for fault recovery This change prevents the submission API from blocking while performing slow operations like fault injection/recovery and app teardown. These operations can take minutes when running multiple problems concurrently. Changes: - Add async cleanup method that performs fault recovery, undeploy, and cluster reconciliation in the background - Update _finish_problem to set state to tearing_down immediately and schedule async cleanup - Modify submit() to capture results snapshot and return immediately - Prevent noise restart during teardown state - Add tearing_down to possible stages in API documentation Fixes SREGym#244 Co-authored-by: Jackson Clark <[email protected]> * free event loop from cleanup * Use force exit to shutdown api server * suppress irrelevant error * Return without evaluating * Close SREGym#757 * Mount .aws directory * taking over fleetcast * update fleetcast owner --------- Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: Jackson Clark <[email protected]> Co-authored-by: Yiming Su <[email protected]>
The Real-World Failure Story
This problem is inspired by a Kubernetes Service selector misconfiguration where a Service selector is too broad or labels are accidentally shared across unrelated workloads.
In this kind of failure, the Service still has endpoints, but the endpoint list is polluted with pods that do not actually serve the expected traffic. This is harder to diagnose than a simple empty-endpoint failure because the pods may all be Running, the Service may look healthy at first glance, and the endpoint list is not empty.
Reference:
https://medium.com/emburse/intermittent-connection-refused-errors-after-new-service-deployment-in-kubernetes-22dd5feae1ad
However, some requests can still be routed to unintended pods, which may reject connections, time out, or return invalid responses.
How We Simulate the Failure on SREGym
This PR adds a new SREGym benchmark problem:
service_wrong_pod_selection_hotel_reservationThe problem uses the Hotel Reservation application and targets the
frontendService.Under normal behavior, the
frontendService selects frontend pods using:The injected fault changes the Service selection behavior by:
sregym.io/frontend-route=trueto the intendedfrontenddeployment.sregym.io/frontend-route=trueto the unintendedsearchdeployment.frontendService selector with:After the fault is injected, the
frontendService selects both the realfrontendpod and the unintendedsearchpod.The
frontendService routes traffic to targetPort5000, while thesearchcontainer listens on port8082. Therefore, traffic routed to thesearchpod through thefrontendService can fail.The Problem Runtime Behavior
After fault injection:
frontendService still has endpoints.frontendEndpointSlice is non-empty but contains bothfrontendandsearchpods.frontend:5000may fail because some traffic can be routed to thesearchpod, which does not serve frontend traffic on port5000.The deterministic failure signal is EndpointSlice pollution. Curl failures are useful supporting evidence, but the mitigation oracle does not depend on the exact curl failure rate.
Useful debugging commands:
Expected broken-state evidence:
The Agent Behavior
A successful agent should diagnose that the
frontendService is selecting unintended pods because its selector was broadened to:The agent should identify that both the intended
frontendpod and the unintendedsearchpod match this selector.The root cause is the
frontendService selector misconfiguration. It is not a pod crash, not an empty endpoint set, and not a failure inside thesearchservice itself.Expected diagnosis:
Expected mitigation is to restore the
frontendService selector to its original value:The mitigation should result in the
frontendEndpointSlice containing only frontend pods.The mitigation oracle checks that:
frontendService selector is exactly restored to{"io.kompose.service": "frontend"}.searchendpoint remains.In manual validation, the diagnosis submission passed the LLM-as-a-judge evaluation with 100/100, and the mitigation oracle passed after restoring the frontend Service selector and removing the polluted endpoint.
Validation
Local compile check passed:
Registry check passed:
End-to-end SREGym validation passed:
frontendService selector changed to{"sregym.io/frontend-route":"true"}.frontendEndpointSlice contained bothfrontendandsearchpods.frontend:5000showed connection failures as supporting runtime evidence.frontendService selector.frontendEndpointSlice contained only frontend pods.Files Changed