Repository navigation
Add automatic retry logic for transient Gemini API errors (503, 429) - #385
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
beaf0dc to
5152918
Compare
65139a1 to
5152918
Compare
|
Your branch is 1 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
1 similar comment
|
Your branch is 1 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
|
Your branch is 6 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
|
Your branch is 8 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
|
Your branch is 11 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
|
Your branch is 15 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
Fixes google#240 This change implements exponential backoff retry for transient errors in the Gemini provider, preventing entire document processing failures when a single chunk encounters temporary service overload (503) or rate limiting (429) errors. Changes: - Add retry configuration parameters (max_retries, retry_delay, max_retry_delay) - Implement _is_retryable_error() to distinguish temporary vs permanent errors - Add exponential backoff retry logic in _process_single_prompt() - Each chunk retries independently without affecting other chunks - Add comprehensive test coverage (30 test cases) Benefits: - Prevents API quota waste from re-processing entire documents - Reduces 429 errors from excessive retries - Improves reliability for large batch processing
Maintainer polish on top of google#385: - Typed google.genai APIError.code classification for 408/429/5xx; fall through to httpx transient subclasses (TimeoutException, NetworkError, RemoteProtocolError, ProxyError) and stdlib ConnectionError/TimeoutError. LocalProtocolError and UnsupportedProtocol intentionally excluded since they indicate client/config bugs. - Narrow regex fallback for exceptions that only carry the status in a message; avoids matching permanent failures like bare "quota" or "unavailable". - Multiplicative jitter (uniform(0.5, 1.5)) with post-jitter cap at max_retry_delay so the named maximum bounds the real sleep. - Keyword-only retry knobs (max_retries, retry_delay, max_retry_delay) with init-time validation. - Guard against stacking with google-genai HttpOptions.retry_options. Matches the SDK's own normalization: retry_options with attempts in {0, 1} is treated as "retries disabled" and allowed; None or >1 raises InferenceConfigError. Covers HttpOptions object, HttpOptionsDict snake-case, and HttpOptionsDict camelCase shapes. - Rename tests/test_gemini_retry.py -> tests/gemini_retry_test.py to match the repo's pytest discovery pattern (python_files = "*_test.py"). - 65 parametrized tests (absl.testing.parameterized) covering the classifier, retry loop, parallel-chunk scenarios, init validation, and the SDK stacking guard matrix.
55e2fdd to
eb4dda3
Compare
Maintainer polish on top of google#385: - Typed google.genai APIError.code classification for 408/429/5xx; fall through to httpx transient subclasses (TimeoutException, NetworkError, RemoteProtocolError, ProxyError) and stdlib ConnectionError/TimeoutError. LocalProtocolError and UnsupportedProtocol intentionally excluded since they indicate client/config bugs. - Narrow regex fallback for exceptions that only carry the status in a message; avoids matching permanent failures like bare "quota" or "unavailable". - Multiplicative jitter (uniform(0.5, 1.5)) with post-jitter cap at max_retry_delay so the named maximum bounds the real sleep. - Keyword-only retry knobs (max_retries, retry_delay, max_retry_delay) with init-time validation. - Guard against stacking with google-genai HttpOptions.retry_options. Matches the SDK's own normalization: retry_options with attempts in {0, 1} is treated as "retries disabled" and allowed; None or >1 raises InferenceConfigError. Covers HttpOptions object, HttpOptionsDict snake-case, and HttpOptionsDict camelCase shapes. - Rename tests/test_gemini_retry.py -> tests/gemini_retry_test.py to match the repo's pytest discovery pattern (python_files = "*_test.py"). - 65 parametrized tests (absl.testing.parameterized) covering the classifier, retry loop, parallel-chunk scenarios, init validation, and the SDK stacking guard matrix.
eb4dda3 to
7e5f7ba
Compare
Description
Fixes #240
This change implements exponential backoff retry for transient errors in the Gemini provider, preventing entire document processing failures when a single chunk encounters temporary service overload (503) or rate limiting (429) errors.
Changes:
max_retries,retry_delay,max_retry_delay)_is_retryable_error()to distinguish temporary vs permanent errors_process_single_prompt()How Has This Been Tested?
Test suite in
tests/test_gemini_retry.pywith 30 test cases covering error classification, retry logic, parallel processing, and configuration.Checklist