Skip to content

Add rollout diagnostics for cloud-native deployment wait failures - #1445

Open
wakqasahmed wants to merge 2 commits into
hyperledger:mainfrom
wakqasahmed:fix/fsat-rollout-diagnostics-1438
Open

wakqasahmed wants to merge 2 commits into
hyperledger:mainfrom
wakqasahmed:fix/fsat-rollout-diagnostics-1438

Conversation

@wakqasahmed

Copy link
Copy Markdown

Refs #1438 - improves diagnostics, does not close it (the underlying CI failure is still unresolved).

The cloud/console CI jobs have been failing since early July with only one line of output: error: deployment fabric-operator exceeded its progress deadline. I traced why the log is that terse rather than confirming the actual root cause of the deployment not becoming ready - I couldn't reproduce the full kind/k8s environment locally to pin that down further.

Two things compound to make it a dead end right now:

  • wait_for_deployment() only runs kubectl rollout status, which on a ProgressDeadlineExceeded failure prints nothing beyond that one summary line - no pod status, no events, no container logs.
  • LOG_ERROR_LINES defaults to 1, so even the debug log's own tail-on-exit only ever shows the single last line, no matter how much a failing command actually printed.

This adds a diagnostics dump (describe deploy, get pods, describe pods, logs --previous and logs, scoped to the deployment's own selector via kubectl get deploy -o json | jq) when wait_for_deployment fails, and bumps LOG_ERROR_LINES to 200 so it actually surfaces in the log instead of being clipped to one line.

Doesn't fix the underlying regression - just turns the next occurrence (including a rerun of this same job) into something a maintainer can actually diagnose from the CI log.

@wakqasahmed
wakqasahmed requested a review from a team as a code owner September 3, 2026 12:22
Refs hyperledger#1438. The cloud/console CI jobs in the Full Stack Asset Transfer
Guide have been failing since early July with only a single terse line
of output: 'error: deployment fabric-operator exceeded its progress
deadline'. Traced this to two compounding issues in
infrastructure/sample-network:

- wait_for_deployment() only ran 'kubectl rollout status', which prints
  nothing beyond that one summary line on a ProgressDeadlineExceeded
  failure - no pod status, no events, no container logs.
- LOG_ERROR_LINES defaults to 1, so even the debug log's own tail on
  exit only ever surfaces the single last line, regardless of how much
  useful output a command actually produced.

This doesn't fix the underlying cause of the operator deployment not
becoming ready (I could not reproduce the full kind/k8s environment
locally to pin that down), but it should turn the next occurrence -
including a rerun of this exact CI job - into something diagnosable
from the log instead of a dead end.

Signed-off-by: wakqasahmed <wakqasahmed@gmail.com>
@wakqasahmed
wakqasahmed force-pushed the fix/fsat-rollout-diagnostics-1438 branch from 45a0b36 to 92d94f9 Compare September 3, 2026 12:41
@wakqasahmed

Copy link
Copy Markdown
Author

Hi @bestbeforetoday — noticed this PR doesn't have a reviewer assigned yet — it's been sitting for a while, CI is green and it's mergeable. Would you (or whoever's best placed) be able to take a look, or point me to who should? Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant