Skip to content

Fix tryClaim busy-looping on a slot held through a thread-local cache - #310

Open
soerenreichardt wants to merge 1 commit into
chrisvest:mainfrom
soerenreichardt:fix-tryclaim-spin-on-tlr-claimed-slot
Open

soerenreichardt wants to merge 1 commit into
chrisvest:mainfrom
soerenreichardt:fix-tryclaim-spin-on-tlr-claimed-slot

Conversation

@soerenreichardt

@soerenreichardt soerenreichardt commented Oct 1, 2026 •

Copy link
Copy Markdown

Problem

tryClaim() can fail to return while another thread holds an object it claimed through its thread-local cache. While it spins, it allocates without bound. We hit this in production: a pool of 4 objects shared by 8 threads drove a 30 GB JVM into back-to-back full GCs within seconds.

Root cause

BlazePool.slowClaim(BSlotCache) only polls the live-queue when its local slot is null:

slot = newAllocations.pop();
for (;;) {
  if (slot == null) {
    slot = live.poll();
  }
  if (slot == null) {
    disregardPile.refill();
    return null;
  } else if (slot.live2claim()) {
    ...
  } else {
    disregardPile.push(slot);   // slot is not reset
  }
}

A slot claimed through the thread-local fast path (tlrClaim, live2claimTlr) is not removed from the live-queue. It sits there in the TLR_CLAIMED state. Each pass of the loop then does the following:

  1. live.poll() returns that slot.
  2. live2claim() fails, because the state is TLR_CLAIMED, not LIVING.
  3. The slot is pushed onto the disregard pile. That allocates a RefillSlot, which stays reachable from the pile.
  4. slot is still non-null, so the next pass skips the poll and retries the same slot.

The loop only exits once the holder calls release() and the CAS finally succeeds. Until then, tryClaim() behaves like an unbounded blocking claim that also leaks memory. The next refill() then calls live.offer once per pushed node. The live-queue fills with duplicates of the same slot, and later claims have to drain them.

The timed slowClaim(Timeout, BSlotCache) doesn't have this problem. It reassigns slot from newAllocations.pop() or live.poll(...) at the top of every pass, and it is bounded by the timeout. So claim(new Timeout(0, ...)) returns null immediately in the same situation, while tryClaim() hangs.

Fix

Set slot = null at the end of each pass, so every pass polls the live-queue, matching the timed variant.

🤖 Generated with Claude Code

The non-blocking `slowClaim(cache)` used by `tryClaim()` only polls the
live-queue when its local `slot` variable is null, but never resets it
after a failed `live2claim()`. A slot that another thread has claimed
through its thread-local cache stays in the live-queue in the
TLR_CLAIMED state. When `tryClaim()` polls such a slot, the CAS fails,
the slot is pushed onto the disregard pile, and the loop retries the
same slot without polling again. `tryClaim()` therefore does not return
until the other thread releases the object, and every iteration
allocates a `RefillSlot` that stays reachable from the disregard pile.
Under contention this exhausts the heap within seconds, and the next
`refill()` floods the live-queue with duplicates of the same slot.

Reset `slot` at the end of each iteration, so every attempt polls the
live-queue, like the timed `slowClaim(timeout, cache)` already does.
The loop then terminates once the live-queue is drained.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@chrisvest

Copy link
Copy Markdown
Owner

Created https://codeberg.org/chrisvest/stormpot/issues/21 to track

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants