Conversation
Ziinc
approved these changes
Sep 17, 2026
Ziinc
deleted the
adammokan/o11y-2427-handle-db_connection-connection_error-telemetry-for
branch
September 17, 2026 12:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
DBConnection sheds read pool checkouts under saturation (
:queue_timeout) or when a queued request's deadline expires. Nothing counts them today, and the shed wait lands inread_pool.checkout.pool_timeas if a connection had been obtained.DBConnection already hands every failed checkout to the
:logcallback we pass onCh.query/4, withconnection_timenil. This matches that shape and emits a counter from the callback, wherebackend_idand the read cluster label are already in scope.Key Changes
handle_read_pool_log/3head matches aConnectionErrorresult withconnection_time: niland emits[:logflare, :clickhouse, :read_pool, :checkout_error]taggedbackend_id,read_cluster,reason. A successful checkout always records a:prepareevent first, so post-checkout errors never matchpool_timelogflare.clickhouse.read_pool.checkout_errordeclared as asum.reasonmirrorsDBConnection.ConnectionError.reasoncheckout_errorandquery_errordescriptions state the overlap: aqueue_timeoutappears in bothRisks / Considerations
checkout_erroris per attempt and the pool-health signal;query_erroris per query and the user-facing failure ratepool_timep99 drops under saturation Shed checkouts were the largest samples. Re-baseline any alerts on it