You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
ThreadSanitizer intermittently reports a data race between lsl::factory::new_sample(double, bool) on the inlet's receiving thread and lsl::data_receiver::pull_sample_typed<int>() on the application thread.
This was observed while validating #291 at commit 4a3e8865e776cf9fa8d93b7533edfe433e12ad24. Filing separately for later investigation; this is not the outlet queue producer race addressed by that PR.
Evidence
The focused [outlet],[open],[reopen],[sync] suite passed 1,710 assertions across 20 test cases, with no TSan warnings in that run.
Repeating have_consumers becomes false after disconnect during push100 times produced this report in 15 runs, one warning per affected run.
All 100 runs passed their Catch2 assertions (5 assertions per run). The 15 affected processes nevertheless exited unsuccessfully with SIGABRT after TSan reported the warning. There were no timeouts.
No reports of the outlet queue producer race were observed in these runs.
Debug build; library, test executable, and fetched Catch2 built with TSan
TSAN_OPTIONS="history_size=7 halt_on_error=0"; no suppressions
Before testing liblsl, a minimal empty program ran successfully under TSan and a deliberately racy program produced the expected data-race warning. These checks also passed with Homebrew LLVM 23.1.1; the liblsl results above used Apple Clang.
Run repeatedly, retaining each report and bounding each run to 120 seconds:
importosimportsubprocessfrompathlibimportPathlogs=Path("tsan-race-logs")
logs.mkdir(exist_ok=True)
env=dict(os.environ, TSAN_OPTIONS="history_size=7 halt_on_error=0")
foriinrange(100):
with (logs/f"run-{i:03}.log").open("w") aslog:
try:
result=subprocess.run(
["./build-tsan/testing/lsl_test_exported",
"have_consumers becomes false after disconnect during push"],
stdout=log, stderr=subprocess.STDOUT, env=env, timeout=120)
print(i, result.returncode)
exceptsubprocess.TimeoutExpired:
print(i, "timeout")
Representative report
The report identifies an 8-byte write in new_sample() and an 8-byte read in pull_sample_typed<int>() to the same address in a factory-allocated heap block. Full representative warning below (source line symbolization was unavailable in this run):
The source suggests the conflicting field is the sample timestamp: new_sample() resets timestamp_ when reusing an object, while pull_sample_typed() reads the timestamp to return it to the caller. The pull path holds a sample_p; recycling should happen only after the final reference is released.
The root cause is not established. This issue records a reproducible sanitizer report, not a proven premature-reuse bug. Investigation should trace sample ownership, reference-count release/fence synchronization, and publication through the factory freelist, and determine whether this is a real synchronization/lifetime defect or a synchronization pattern TSan does not recognize.
No incorrect timestamps or sample corruption were demonstrated by these tests. Do not infer a fix (or add a suppression) solely from the two reported function names.
Observed behavior
ThreadSanitizer intermittently reports a data race between
lsl::factory::new_sample(double, bool)on the inlet's receiving thread andlsl::data_receiver::pull_sample_typed<int>()on the application thread.This was observed while validating #291 at commit
4a3e8865e776cf9fa8d93b7533edfe433e12ad24. Filing separately for later investigation; this is not the outlet queue producer race addressed by that PR.Evidence
[outlet],[open],[reopen],[sync]suite passed 1,710 assertions across 20 test cases, with no TSan warnings in that run.have_consumers becomes false after disconnect during push100 times produced this report in 15 runs, one warning per affected run.Environment
TSAN_OPTIONS="history_size=7 halt_on_error=0"; no suppressionsBefore testing liblsl, a minimal empty program ran successfully under TSan and a deliberately racy program produced the expected data-race warning. These checks also passed with Homebrew LLVM 23.1.1; the liblsl results above used Apple Clang.
Reproduction
Starting from the PR commit above:
Run repeatedly, retaining each report and bounding each run to 120 seconds:
Representative report
The report identifies an 8-byte write in
new_sample()and an 8-byte read inpull_sample_typed<int>()to the same address in a factory-allocated heap block. Full representative warning below (source line symbolization was unavailable in this run):Interpretation and investigation scope
The source suggests the conflicting field is the sample timestamp:
new_sample()resetstimestamp_when reusing an object, whilepull_sample_typed()reads the timestamp to return it to the caller. The pull path holds asample_p; recycling should happen only after the final reference is released.The root cause is not established. This issue records a reproducible sanitizer report, not a proven premature-reuse bug. Investigation should trace sample ownership, reference-count release/fence synchronization, and publication through the factory freelist, and determine whether this is a real synchronization/lifetime defect or a synchronization pattern TSan does not recognize.
No incorrect timestamps or sample corruption were demonstrated by these tests. Do not infer a fix (or add a suppression) solely from the two reported function names.