Add rollout diagnostics for cloud-native deployment wait failures - #1445
Open
wakqasahmed wants to merge 2 commits into
Open
wakqasahmed wants to merge 2 commits into
wakqasahmed wants to merge 2 commits into
Conversation
Refs hyperledger#1438. The cloud/console CI jobs in the Full Stack Asset Transfer Guide have been failing since early July with only a single terse line of output: 'error: deployment fabric-operator exceeded its progress deadline'. Traced this to two compounding issues in infrastructure/sample-network: - wait_for_deployment() only ran 'kubectl rollout status', which prints nothing beyond that one summary line on a ProgressDeadlineExceeded failure - no pod status, no events, no container logs. - LOG_ERROR_LINES defaults to 1, so even the debug log's own tail on exit only ever surfaces the single last line, regardless of how much useful output a command actually produced. This doesn't fix the underlying cause of the operator deployment not becoming ready (I could not reproduce the full kind/k8s environment locally to pin that down), but it should turn the next occurrence - including a rerun of this exact CI job - into something diagnosable from the log instead of a dead end. Signed-off-by: wakqasahmed <wakqasahmed@gmail.com>
wakqasahmed
force-pushed
the
fix/fsat-rollout-diagnostics-1438
branch
from
September 3, 2026 12:41
45a0b36 to
92d94f9
Compare
Author
|
Hi @bestbeforetoday — noticed this PR doesn't have a reviewer assigned yet — it's been sitting for a while, CI is green and it's mergeable. Would you (or whoever's best placed) be able to take a look, or point me to who should? Thanks! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #1438 - improves diagnostics, does not close it (the underlying CI failure is still unresolved).
The cloud/console CI jobs have been failing since early July with only one line of output:
error: deployment fabric-operator exceeded its progress deadline. I traced why the log is that terse rather than confirming the actual root cause of the deployment not becoming ready - I couldn't reproduce the full kind/k8s environment locally to pin that down further.Two things compound to make it a dead end right now:
wait_for_deployment()only runskubectl rollout status, which on aProgressDeadlineExceededfailure prints nothing beyond that one summary line - no pod status, no events, no container logs.LOG_ERROR_LINESdefaults to1, so even the debug log's own tail-on-exit only ever shows the single last line, no matter how much a failing command actually printed.This adds a diagnostics dump (
describe deploy,get pods,describe pods,logs --previousandlogs, scoped to the deployment's own selector viakubectl get deploy -o json | jq) whenwait_for_deploymentfails, and bumpsLOG_ERROR_LINESto 200 so it actually surfaces in the log instead of being clipped to one line.Doesn't fix the underlying regression - just turns the next occurrence (including a rerun of this same job) into something a maintainer can actually diagnose from the CI log.