Repository navigation
Submission API does not return HTTP response until other logic finishes #244
Description
Activity
@claude take a look at this. We should return evaluation results immediately, at intermediate states (like tearing down) we transition to a different state.
Reacted by claudeClaude finished @HacksonClark's task in 4m 26s —— View job
Working on Issue #244

Todo List
- Read CLAUDE.md for repo setup and guidelines
- Locate submission API endpoints
- Understand current blocking behavior
- Design solution to return evaluation results immediately
- Implement async cleanup with state transitions
- Update submit endpoint to return immediately
- Handle edge cases (submissions during teardown)
- Update API documentation
- Commit and push changes (waiting for approval)
Implementation
Changes made:
-
Added async cleanup (
conductor.py:294-332):- Created
_cleanup_async()method that performs fault recovery, app teardown, and cluster reconciliation asynchronously - This allows these slow operations to run in the background without blocking the HTTP response
- Added detailed logging to track cleanup progress
- Created
-
Updated
_finish_problem()(conductor.py:334-344):- Now sets state to
"tearing_down"immediately - Schedules
_cleanup_async()to run in the background usingasyncio.create_task() - Returns control immediately to the caller
- Now sets state to
-
Modified
submit()method (conductor.py:411-474):- Creates a snapshot of evaluation results before advancing to next stage
- Returns the snapshot immediately, even if cleanup is running in background
- Added check to prevent noise restart during
"tearing_down"state - Handles submissions during teardown by returning current results
-
Updated API documentation (
conductor_api.py:148):- Added
"tearing_down"to the list of possible stages in/statusendpoint docs
- Added
How it works:
Before:
/submit → evaluate → advance_to_next_stage → _finish_problem (blocking: recovery, undeploy, reconcile) → return resultsAfter:
/submit → evaluate → snapshot results → advance_to_next_stage → _finish_problem (sets state to "tearing_down", schedules async cleanup) → return results immediatelyThe HTTP response now returns as soon as evaluation completes, with cleanup happening asynchronously in the background. The state transitions:
diagnosis/mitigation→tearing_down→done.I need approval to run the git commit command. The changes are staged and ready.
@claude You have approval, please commit to your branch. I can't see your code!
Reacted by Yiming SuReacted by claudeClaude finished @HacksonClark's task in 3m 5s —— View job •
claude/issue-244-20260213-2244• Create PR ➔
Working on Issue #244
Todo List
- Verify current code state
- Implement async cleanup method
- Update _finish_problem to use async cleanup
- Update submit endpoint to return immediately
- Update API documentation
- Commit and push changes
Implementation Summary
Changes made:
-
Added async cleanup (
conductor.py:293-332):- Created
_cleanup_async()method that performs fault recovery, app teardown, and cluster reconciliation asynchronously - This allows these slow operations to run in the background without blocking the HTTP response
- Added detailed logging to track cleanup progress
- Created
-
Updated
_finish_problem()(conductor.py:334-348):- Now sets state to
"tearing_down"immediately - Schedules
_cleanup_async()to run in the background usingasyncio.create_task() - Returns control immediately to the caller
- Now sets state to
-
Modified
submit()method (conductor.py:389-467):- Creates a snapshot of evaluation results before advancing to next stage
- Returns the snapshot immediately, even if cleanup is running in background
- Added check to prevent noise restart during
"tearing_down"state - Handles submissions during teardown by returning current results
-
Updated API documentation (
conductor_api.py:148):- Added
"tearing_down"to the list of possible stages in/statusendpoint docs
- Added
How it works:
Before:
/submit → evaluate → advance_to_next_stage → _finish_problem (blocking: recovery, undeploy, reconcile) → return resultsAfter:
/submit → evaluate → snapshot results → advance_to_next_stage → _finish_problem (sets state to "tearing_down", schedules async cleanup) → return results immediatelyThe HTTP response now returns as soon as evaluation completes, with cleanup happening asynchronously in the background. The state transitions:
diagnosis/mitigation→tearing_down→done.Committed
All changes have been committed and pushed to the branch
claude/issue-244-20260213-2244.- added a commit that references this issue
on Feb 13, 2026 - added a commit that references this issue
on Feb 20, 2026
@HacksonClark
Consider the noop phase. If fault injection for a problem takes a long time, the submission HTTP call will not receive a response until the fault injection finishes. In cases where we are running multiple problems at the same time (i.e., concurrent faults), fault injection can take a while.
Same case for fault recovery, which can take minutes to finish depending how many problems we are running at the same time.
I don't think the HTTP response should wait for fault injection/recovery to finish.