You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A check that completes successfully but slower than our window is recorded identically to one that never worked. Deal checks have a single hard cutoff (dealJobTimeoutSeconds, default 360s in config/constants.ts). When the window expires, classifyFailureStatus emits failure.timedout and the deal row is persisted with status = failed (deal.service.ts#L664). Nothing revisits that row when the addPieces message lands on chain later, so an eventually-successful deal stays a permanent failure in the DB and in every downstream metric.
The 2026-07-20/21 mainnet congestion made this concrete (internal discussion: Slack thread): elevated base fee pushed addPieces confirmations well past the deal window while the data itself stored and retrieved fine, and dealbot recorded deal failures. Those false negatives are indistinguishable from an SP that lost the data.
What this tracks
An outcome class between success and failure: "worked, but outside our time requirement" (e.g. a pending_confirm-style state; naming open). Metrics and dashboards need to separate "SP broken" from "SP or chain slow".
Problem
A check that completes successfully but slower than our window is recorded identically to one that never worked. Deal checks have a single hard cutoff (
dealJobTimeoutSeconds, default 360s inconfig/constants.ts). When the window expires,classifyFailureStatusemitsfailure.timedoutand the deal row is persisted withstatus = failed(deal.service.ts#L664). Nothing revisits that row when the addPieces message lands on chain later, so an eventually-successful deal stays a permanent failure in the DB and in every downstream metric.The 2026-07-20/21 mainnet congestion made this concrete (internal discussion: Slack thread): elevated base fee pushed addPieces confirmations well past the deal window while the data itself stored and retrieved fine, and dealbot recorded deal failures. Those false negatives are indistinguishable from an SP that lost the data.
What this tracks
pending_confirm-style state; naming open). Metrics and dashboards need to separate "SP broken" from "SP or chain slow".failedrows when the deal later confirms on chain (overlaps database synchronization with chainstate #465).Related
failure.timedout, binary-cutoff artifact), RPC timeout mismatch: dealbot (10s) gives up before eRPC (30s) can fail over #603 (slow RPC recorded as hard deal failures), lower timeouts for deals/retrievals #267 (timeout tuning), database synchronization with chainstate #465 (DB vs chainstate sync)