Skip to content

Add execution-context generation token to DI snapshots - #6233

Draft
p-datadog wants to merge 5 commits into
masterfrom
di-snapshot-generation-token
Draft

p-datadog wants to merge 5 commits into
masterfrom
di-snapshot-generation-token

Conversation

@p-datadog

@p-datadog p-datadog commented Aug 24, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

Adds a per-thread execution-context generation token to DI snapshots. Each snapshot now includes a generation value alongside the existing thread id, so snapshots from different execution contexts that share a recycled thread id can be distinguished.

Motivation:

A thread's id can be reused after the thread exits, so two snapshots sharing a thread id are not guaranteed to come from the same execution context. Pairing the thread id with a generation token makes the ambiguity detectable: same id with different generation means different threads.

Change log entry

Yes. Dynamic Instrumentation: snapshots now carry an execution-context generation token to disambiguate reused thread ids.

Additional Notes:

N/A

How to test the change?

  • Unit tests added
  • Integration tests added
  • System test: Test_Debugger_Snapshot_Correlation_Fields::test_generation_token (system-tests#7425)

The snapshot logger object now carries a generation token alongside
thread_id. A thread's ident can be reused after the thread exits, so two
snapshots sharing a thread_id are not guaranteed to come from the same
execution context; pairing the id with a per-thread generation counter
makes the ambiguity detectable.

The token is a lazily assigned counter keyed by the Thread object (not
its ident), held in a process-wide ObjectSpace::WeakMap so finalized
threads do not leak. Tokens are unique only within a runtime id, which
is emitted alongside them in the snapshot envelope.

Closes the generation-token gap flagged in the Casual Correlation RFC
and matches the system-tests correlation gate
(Test_Debugger_Snapshot_Correlation_Fields::test_generation_token), which
reads logger.generation.
@p-datadog p-datadog added the AI Generated Largely based on code generated by an AI or LLM. This label is the same across all dd-trace-* repos label Aug 24, 2026
@dd-octo-sts dd-octo-sts Bot added the debugger Live Debugger (+Dynamic Instrumentation, +Symbol Database) label Aug 24, 2026
@dd-octo-sts

dd-octo-sts Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Thank you for updating Change log entry section 👏

Visited at: 2026-08-26 17:20:08 UTC

@dd-octo-sts

dd-octo-sts Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Typing analysis

Note: Ignored files are excluded from the next sections.

Untyped methods

This PR introduces 5 partially typed methods, and clears 30 partially typed methods. It increases the percentage of typed methods from 70.06% to 71.16% (+1.1%).

Partially typed methods (+5-30) ❌ Introduced:
sig/datadog/open_feature/evaluation_engine.rbs:20
└── def fetch_value: (
        ::String flag_key,
        default_value: untyped,
        expected_type: ::Symbol,
        ?evaluation_context: ::OpenFeature::SDK::EvaluationContext?
      ) -> ResolutionDetails
sig/datadog/open_feature/native_evaluator.rbs:14
└── def get_assignment: (
        ::String flag_key,
        default_value: untyped,
        context: ::OpenFeature::SDK::EvaluationContext::fields_t,
        expected_type: ::Symbol
      ) -> ResolutionDetails
sig/datadog/open_feature/native_evaluator.rbs:25
└── def build_resolution_details: (
        Core::FeatureFlags::ResolutionDetails result,
        untyped default_value
      ) -> ResolutionDetails
sig/datadog/open_feature/provider.rbs:74
└── def sdk_success_details: (
        ResolutionDetails result,
        ::Hash[::String, untyped] flag_meta
      ) -> ::OpenFeature::SDK::Provider::ResolutionDetails
sig/datadog/open_feature/provider.rbs:87
└── def build_flag_metadata: (
        ResolutionDetails result,
        ::Integer eval_time_ms
      ) -> ::Hash[::String, untyped]
✅ Cleared:
sig/datadog/open_feature/evaluation_engine.rbs:20
└── def fetch_value: (
        ::String flag_key,
        default_value: untyped,
        expected_type: ::Symbol,
        ?evaluation_context: ::OpenFeature::SDK::EvaluationContext?
      ) -> (ResolutionDetails | Core::FeatureFlags::ResolutionDetails)
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:69
└── def record: (
          flag_key: ::String,
          variant: ::String?,
          allocation_key: ::String?,
          targeting_key: ::String?,
          eval_time_ms: ::Integer,
          attrs: ::Hash[::String, untyped]?,
          ?error_message: ::String?,
          ?runtime_default: bool?
        ) -> void
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:80
└── def flush_and_reset: () -> ::Hash[::Symbol, untyped]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:82
└── def prune_context: (::Hash[::String, untyped]? attrs) -> ::Hash[::String, untyped]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:84
└── def self.prune_context: (::Hash[::String, untyped]? attrs) -> ::Hash[::String, untyped]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:86
└── def self.flatten_context: (::Hash[::String, untyped]? attrs) -> ::Hash[::String, untyped]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:88
└── def self.flatten_value: (
          ::String prefix,
          untyped value,
          ::Hash[::String, untyped] output,
          ::Hash[::Integer, bool] seen,
          ::Integer depth
        ) -> void
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:96
└── def canonical_context_key: (::Hash[::String, untyped]? attrs) -> ::String
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:100
└── def context_value_bytes: (untyped value) -> ::String
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:104
└── def new_entry: (
          ::Integer evaluation_time_ms,
          runtime_default: bool,
          ?error_message: ::String?,
          ?targeting_key: ::String?,
          ?context_attrs: ::Hash[::String, untyped]?
        ) -> ::Hash[::Symbol, untyped]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:112
└── def observe: (::Hash[::Symbol, untyped] entry, ::Integer evaluation_time_ms) -> void
sig/datadog/open_feature/flag_evaluation/writer.rbs:65
└── def enqueue: (**untyped event) -> void
sig/datadog/open_feature/flag_evaluation/writer.rbs:75
└── def snapshot_context_value: (
          untyped value,
          ::Hash[::Integer, bool] seen,
          ::Integer depth
        ) -> untyped
sig/datadog/open_feature/flag_evaluation/writer.rbs:103
└── def build_events: (::Hash[::Symbol, untyped] snapshot) -> ::Array[::Hash[::String, untyped]]
sig/datadog/open_feature/flag_evaluation/writer.rbs:105
└── def build_event: (
          flag_key: ::String,
          variant: ::String?,
          allocation_key: ::String?,
          targeting_key: ::String?,
          entry: ::Hash[::Symbol, untyped],
          flush_time_ms: ::Integer,
          tier: ::Symbol
        ) -> ::Hash[::String, untyped]
sig/datadog/open_feature/flag_evaluation/writer.rbs:115
└── def send_payload_batches: (::Array[::Hash[::String, untyped]] events) -> void
sig/datadog/open_feature/flag_evaluation/writer.rbs:117
└── def send_payload_batch: (::Array[::Hash[::String, untyped]] events) -> untyped
sig/datadog/open_feature/flag_evaluation/writer.rbs:119
└── def encoded_event_for_payload: (
          ::Hash[::String, untyped] event,
          ::Integer base_payload_size
        ) -> [::Hash[::String, untyped], ::Integer, bool]?
sig/datadog/open_feature/flag_evaluation/writer.rbs:124
└── def encoded_event: (::Hash[::String, untyped] event) -> [::Hash[::String, untyped], ::Integer]
sig/datadog/open_feature/flag_evaluation/writer.rbs:128
└── def degrade_event_for_payload_limit: (
          ::Hash[::String, untyped] event
        ) -> ::Hash[::String, untyped]?
sig/datadog/open_feature/flag_evaluation/writer.rbs:134
└── def event_count: (::Hash[::String, untyped] event) -> ::Integer
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:15
└── def finally: (
          hook_context: untyped,
          evaluation_details: untyped,
          **untyped _opts
        ) -> void
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:23
└── def extract_targeting_key: (untyped evaluation_context) -> ::String?
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:25
└── def extract_attributes: (untyped evaluation_context) -> ::Hash[::String, untyped]
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:27
└── def extract_allocation_key: (untyped evaluation_details) -> ::String?
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:29
└── def extract_error_message: (untyped evaluation_details) -> ::String?
sig/datadog/open_feature/hooks/flag_eval_evp_hook.rbs:31
└── def runtime_default?: (untyped evaluation_details) -> bool
sig/datadog/open_feature/native_evaluator.rbs:10
└── def get_assignment: (
        ::String flag_key,
        default_value: untyped,
        context: ::OpenFeature::SDK::EvaluationContext::fields_t,
        expected_type: ::Symbol
      ) -> (Core::FeatureFlags::ResolutionDetails | ResolutionDetails)
sig/datadog/open_feature/provider.rbs:74
└── def sdk_success_details: (
        (ResolutionDetails | Core::FeatureFlags::ResolutionDetails) result,
        ::Hash[::String, untyped] flag_meta
      ) -> ::OpenFeature::SDK::Provider::ResolutionDetails
sig/datadog/open_feature/provider.rbs:87
└── def build_flag_metadata: (
        (ResolutionDetails | Core::FeatureFlags::ResolutionDetails) result,
        ::Integer eval_time_ms
      ) -> ::Hash[::String, untyped]

Untyped other declarations

This PR clears 2 partially typed other declarations. It increases the percentage of typed other declarations from 85.33% to 85.69% (+0.36%).

Partially typed other declarations (+0-2) ✅ Cleared:
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:51
└── @full: ::Hash[::Array[untyped], ::Hash[::Symbol, untyped]]
sig/datadog/open_feature/flag_evaluation/aggregator.rbs:53
└── @degraded: ::Hash[::Array[untyped], ::Hash[::Symbol, untyped]]

If you believe a method or an attribute is rightfully untyped or partially typed, you can add # untyped:accept on the line before the definition to remove it from the stats.

@pr-commenter

pr-commenter Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-08-26 17:44:27

Comparing candidate commit af216e0 in PR branch di-snapshot-generation-token with baseline commit cf557cf in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 49 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

@datadog-prod-us1-4

datadog-prod-us1-4 Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 80.00%
• Overall Coverage: 90.30% (+0.05%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: af216e0 | Docs | View more details | Give us feedback!

ThreadGeneration stored per-thread generation tokens in an
ObjectSpace::WeakMap keyed by Thread. On Ruby 2.5 and 2.6,
ObjectSpace::WeakMap#[]= raises
"ArgumentError: cannot define finalizer for Integer" for any object key,
because WeakMap's internal finalizer registration is broken on those
versions. Every DI snapshot build therefore crashed in
ProbeNotificationBuilder#build_snapshot_base when it read the generation
token via ThreadGeneration.current, breaking all 114 snapshot-producing
tests on the Ruby 2.6 CI job.

Replace the WeakMap with a thread-local stored on the Thread object via
Thread#thread_variable_get/set, keyed by :datadog_di_thread_generation.
This matches the codebase's existing per-thread state pattern
(lib/datadog/appsec/rate_limiter.rb and
lib/datadog/appsec/api_security/sampler.rb use the same THREAD_KEY +
thread_variable approach). Thread-locals are reclaimed when the thread
is garbage collected, so finalized threads do not leak, and the approach
works across the full supported matrix (Ruby 2.6..4.0). DI requires
Ruby 2.6+ (script_compiled), so 2.6 is the lower bound where this code
executes.

Verified: spec/datadog/di/thread_generation_spec.rb passes (4 examples,
0 failures) on Ruby 3.3.12; the four spec assertions were replicated
standalone on Ruby 2.6.10, 2.7.8, 3.4.10, and 4.0.6 (WeakMap error gone,
per-thread token stable and distinct across threads). steep check and
standardrb clean on the changed file.
The generation token added to the snapshot logger object in
probe_notification_builder.rb#build_snapshot_base was missing from the
expected_snapshot_payload logger hashes in
everything_from_remote_config_spec.rb, so the full-payload match
assertions failed on Ruby 2.7..4.0 (5 tests) because the actual payload
includes logger.generation, which the expected payload omitted.

Add generation: Integer to the four expected logger hashes, matching the
matcher pattern the PR already used in
probe_notification_builder_spec.rb. The Integer matcher accommodates the
token's non-deterministic value (a process-wide monotonic counter), the
same way timestamp and duration use Integer.

Verified: spec/datadog/di/integration/everything_from_remote_config_spec.rb
passes (11 examples, 0 failures) on Ruby 3.3.12 with
TEST_DATADOG_INTEGRATION=1. standardrb clean on the changed file.
thread_generation.rbs used bare Integer, Symbol, Thread, and Thread::Mutex.
rbs.md requires core types and stdlib classes in RBS to carry the :: prefix
so Steep resolves them to the global class rather than a same-named constant
in the current module namespace; existing DI sigs follow this (component.rbs,
instrumenter.rbs).

Verified: rspec spec/datadog/di/thread_generation_spec.rb
spec/datadog/di/probe_notification_builder_spec.rb (0 failures).
standard/steep run in CI (absent from the default Gemfile on Ruby 3.2).
State and State#initialize carried only an @api private tag with no
docstring. documentation.md requires docstrings on every public method
including initialize, and comments.md requires a class docstring stating
what the class is. Add a class docstring for State (the process-wide
per-thread generation ledger) and an initializer docstring.

Verified: rspec spec/datadog/di/thread_generation_spec.rb (0 failures).

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds a per-thread execution-context generation token to Dynamic Instrumentation (Live Debugger) snapshot payloads, emitted as logger.generation, to disambiguate cases where a thread ident may be reused after thread exit.

Changes:

  • Introduces Datadog::DI::ThreadGeneration to lazily assign a monotonically increasing, per-thread generation token via thread-local storage.
  • Adds logger.generation to the snapshot logger envelope produced by ProbeNotificationBuilder.
  • Updates DI unit/integration specs and adds an RBS signature for the new module.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
spec/datadog/di/thread_generation_spec.rb New unit spec for thread generation token behavior.
spec/datadog/di/probe_notification_builder_spec.rb Asserts logger.generation is present and matches the current thread’s token; updates expected payload shapes.
spec/datadog/di/integration/everything_from_remote_config_spec.rb Updates integration expectations to include logger.generation.
sig/datadog/di/thread_generation.rbs Adds RBS typings for Datadog::DI::ThreadGeneration.
lib/datadog/di/thread_generation.rb Implements the per-thread generation token ledger/state.
lib/datadog/di/probe_notification_builder.rb Adds logger.generation to snapshot base payload.
lib/datadog/di/boot.rb Ensures DI boot sequence loads the new thread generation implementation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +380 to 383
# Execution-context generation token, paired with thread_id: a reused
# thread id is distinguishable by a different generation.
generation: DI::ThreadGeneration.current,
version: 2,
Comment on lines +3 to +6
require "datadog/di/thread_generation"

RSpec.describe Datadog::DI::ThreadGeneration do
it "returns the same token for the same thread" do

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AI Generated Largely based on code generated by an AI or LLM. This label is the same across all dd-trace-* repos debugger Live Debugger (+Dynamic Instrumentation, +Symbol Database)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants