Skip to content

Flaky test: Test_FirstApplicationSample exhausts HTTP retries before application starts listening #12935

Description

@willdavsmith

Bug information

Steps to reproduce

Run the scheduled noncloud functional tests, including Test_FirstApplicationSample in test/functional-portable/samples/noncloud/tutorial_test.go. The failure is intermittent and occurs when the tutorial application takes longer to start listening than the post-deployment verification allows.

Observed in the September 6, 2026 workflow run 34003702190, samples-noncloud job, originally reported by #12913.

Observed behavior

Deployment and Kubernetes validation succeed, but the post-deployment HTTP verification exhausts three immediate retries before the application starts listening on port 3000:

[rad] Deployment Complete
Pod is ready tutorial:demo-tutorial-58f5675c5f-9fxcw
running post-deploy verification
01:35:39.530092 ... localhost:3000 ... connection refused
01:35:39.566572 ... localhost:3000 ... connection refused
01:35:39.594423 ... localhost:3000 ... connection refused
Get "http://localhost:34267/": EOF
tests failed after 3 retries

The container-log artifact records:

2026-09-06T01:35:40.846732302Z ... Server is running at http://localhost:3000

All three requests fail within approximately 64 milliseconds; the application starts listening approximately 1.25 seconds after the final attempt. Kubernetes pod readiness here does not establish that the application's HTTP endpoint is accepting connections.

Desired behavior

Post-deployment verification should allow a bounded application-startup interval, using context-aware readiness polling or retry backoff rather than exhausting all attempts immediately. Failed port-forward sessions should be recreated as needed. A genuinely unavailable application must still fail within a clear timeout and retain actionable diagnostics.

Consider an application HTTP readiness probe as well, rather than treating pod readiness alone as sufficient for the HTTP check.

Workaround

Rerun the failed job. The next scheduled September 6, 13:07 UTC run passed the samples job at the same Radius commit. This demonstrates intermittency, not a code fix.

System information

rad Version

CLI built from Radius commit c8ad9211a25699c377c45268890e4f67070aa114 in the failing workflow; released version output was not captured in this triage.

Operating system

Linux GitHub Actions runner with a local Kubernetes test cluster; see the linked job for its runner configuration.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or not working as expectedon-callIssue for on-call to research/resolve.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions