Repository navigation
Relax Testcontainers image pull timeouts - #12777
gh-worker-dd-mergequeue-cf854d[bot] merged 7 commits into
Conversation
🟡 Java Benchmark SLOs — Performance SLO warning (near threshold)
PR vs. master results
Commit: Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c8fc514cbe
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
More details
The longer image-pull defaults preserve inherited environment overrides. Recovery from the reported CI image-pull stall remains unproven.
🤖 Bits Code Review · Commit c8fc514
There was a problem hiding this comment.
More details
The timeout variables are correctly configured in GitLab’s shared test-job template, targeting CI image pulls. Recovery from the reported image-pull stall remains unproven.
🤖 Bits Code Review · Commit 88d725d
|
/merge -f --reason "No need to run MQ, since it is CI change only, validated in PR" |
|
View all feedbacks in Devflow UI.
The expected merge time in
Warning This change was merged without running any pre merge CI checks Reason: No need to run MQ, since it is CI change only, validated in PR |
What Does This Do
Relax Testcontainers image-pull defaults in the shared test convention:
TESTCONTAINERS_PULL_TIMEOUTTESTCONTAINERS_PULL_PAUSE_TIMEOUTThe defaults apply to all
Testtasks using this convention, including forked and latest-dependency suites.putIfAbsentpreserves explicit environment overrides. Container startup/readiness timeouts are unchanged.This is an unproven mitigation attempt, not a root-cause fix. The intent is to give temporarily stalled image pulls more time to recover; we hope this reduces failures, but the observed CI stall has not been reproduced locally.
Motivation
CouchbaseClient31V0Testfailed during image fetching, before container creation or Couchbase initialization. The test report showed repeatedDocker image pull has not made progress in 30s - aborting pullmessages, followed by exhaustion of the image-pull retry window. A SQL Server test in the same run failed with the same image-fetch timeout chain, supporting a shared runner/registry issue rather than a Couchbase-specific startup regression.Precedent: #12723 retries an
IOExceptionduring Couchbase node renaming after container startup begins. This failure occurs earlier, so that retry cannot address it. That PR changed neither image-pull timeout; this investigation found no evidence that it introduced the reported failure. The evidence also does not establish whether the original node-renaming race is fully resolved.Also I noticed that
OracleContainervery close to default120stimeout, so better to have more time.Additional Notes
Concrete failure evidence, with the registry reference omitted:
The pause timeout is an inactivity watchdog: Docker download/extraction progress updates reset it. When it expires, Testcontainers interrupts that pull and retries internally while its overall retry window allows, before the test fails and Develocity retries the test. The total timeout is a retry window, not a strict wall-clock cap: an active attempt can run beyond it. These settings do not guarantee a fixed number of retries.
The test passed on a Develocity retry. That is consistent with a transient stall, but does not prove that relaxing either timeout fixes it. The registry may have recovered, or Docker may have reused layers downloaded during the failed attempt. Docker daemon/pull-progress evidence is needed to distinguish slow progress from a permanently stuck pull. Longer timeouts also delay reporting genuine pull failures.
Validation: the four shared Couchbase retry unit tests and the Couchbase 3.1/3.2 V0 latest-dependency suites (10 and 12 tests) passed locally with the shared convention. This exercises normal container startup and tracing; it does not reproduce or prove recovery from the CI image-pull stall.
Contributor Checklist
type:and (comp:orinst:) labels in addition to any other useful labelsclose,fix, or any linking keywords when referencing an issueUse
solvesinstead, and assign the PR milestone to the issue