Skip to content

Synchronize concurrent include hashing in Parser - #16

Open
CharlesXu-HQ wants to merge 1 commit into
deepseek-ai:mainfrom
CharlesXu-HQ:fix/parser-concurrent-hashing
Open

CharlesXu-HQ wants to merge 1 commit into
deepseek-ai:mainfrom
CharlesXu-HQ:fix/parser-concurrent-hashing

Conversation

@CharlesXu-HQ

Copy link
Copy Markdown

Parser shares its include-hash cache and cycle-detection set across calls. Native threads parsing the same tracked header can race on those containers and report a circular include for an acyclic header. This happens before Runtime reaches its in-memory cache.

Guard both public hashing entry points with one recursive mutex, since recursive include traversal calls between them. The lock covers parsing only; NVCC compilation is outside it. Add an eight-thread regression that starts concurrent parses of the same tracked header and checks every digest.

Validation:

  • On the base revision, an eight-thread reproducer fails with circular include may occur: shared/header.cuh for an acyclic header. With this change, it passes 10 consecutive rounds.
  • The full python tests/test_cuda.py suite passes on an RTX 5090 with CUDA 13.0 and PyTorch 2.9.1+cu130, including the new concurrent parser test, header self-containment checks, and eight-process cache publication tests.
  • git diff --check passes.

This change synchronizes calls through Parser's hashing methods. Direct concurrent mutation of its public containers remains unsupported; the separate in-memory cache synchronization is addressed by #12.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant