Skip to content

feat(llmobs): emit gen_ai.* attributes on APM spans - #20083

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 6 commits into
mainfrom
max.zhang/llmobs-gen-ai-apm-tags
Sep 10, 2026
Merged

gh-worker-dd-mergequeue-cf854d[bot] merged 6 commits into
mainfrom
max.zhang/llmobs-gen-ai-apm-tags

Conversation

@mz1119

@mz1119 mz1119 commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

The APM trace UI merges gen_ai.* keys into spans client-side from the LLMObs track, so the values render but aren't indexed — you can't filter, facet, or monitor on model, provider, or token usage in APM. web-ui code ishere

This change emits a subset from tracer instead, making them real searchable APM tags: gen_ai.operation.name, gen_ai.request.model, gen_ai.provider.name, gen_ai.application.name, and gen_ai.conversation.id, plus gen_ai.usage.{input,output,total,cache_read_input,cache_write_input,reasoning_output}_tokens as metrics.

Message bodies (input, output, tool definitions, retrieval documents) stay off the APM span and continue to come from the LLMObs track.

Testing

gen_ai fields are queryable in APM if sent with changes made in the SDK. Example span

Querying that span through the APM span API shows the tags in the payload (trimmed to the relevant keys):

"custom": {
  "component": "openai",
  "gen_ai": {
    "application": { "name": "max-test" },
    "operation": { "name": "llm" },
    "provider": { "name": "openai" },
    "request": { "model": "gpt-4.1-mini-2025-04-14" },
    "usage": {
      "input_tokens": 13,
      "output_tokens": 10,
      "total_tokens": 23,
      "cache_read_input_tokens": 0,
      "cache_write_input_tokens": 0,
      "reasoning_output_tokens": 0
    }
  }
}

Unit coverage in tests/llmobs/test_llmobs_gen_ai_apm_tags.py.

The APM trace UI reproduces gen_ai.* keys today by querying the LLMObs track
and merging the result into each span client-side. Those merged values are
displayable but not searchable, so you cannot filter, facet, or monitor on
model, provider, or token usage in APM.

Emit the scalar subset from the tracer instead, so they become real APM tags:

  meta:    gen_ai.operation.name, gen_ai.request.model, gen_ai.provider.name,
           gen_ai.application.name, gen_ai.conversation.id
  metrics: gen_ai.usage.{input,output,total,cache_read_input,
           cache_write_input}_tokens

Message bodies (input, output, tool definitions, retrieval documents) stay off
the APM span: they are unbounded and not usefully queryable once serialized.
The UI keeps joining the LLMObs track for those.

Emission is centralized in LLMObs._prepare_llmobs_span_data, before the user
span processor and _normalize_llmobs_meta, so it covers every integration, the
decorators, and manual LLMObs.annotate() from one call site. Running early
keeps the tags on spans whose LLMObs event a processor drops, and reads
model_name and model_provider before normalization pops them for kinds other
than llm/embedding; LLMObsSpan exposes none of those fields, so nothing is
lost by not waiting. BaseLLMIntegration's shadow-tag path covers the
LLMObs-disabled case, where no meta_struct exists, with the same per-
integration coverage as the existing _dd.llmobs.* shadow tags.

gen_ai.operation.name carries the raw LLMObs span kind rather than the OTel
gen_ai enum, matching what the UI already writes into that key so the
tracer-emitted and enriched values agree.

The two helpers live in llmobs/_utils.py rather than a dedicated module. A new
module importing Span would add a fresh product:llmobs -> product:tracing edge
and join the llmobs import tangle, which detect_layering_violations and
detect_circular_imports both reject. _utils already owns the other llmobs
span/meta_struct accessors and already carries that edge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Tests

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 599c4fe | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-09-11 15:54:14

Comparing candidate commit f9254a2 in PR branch max.zhang/llmobs-gen-ai-apm-tags with baseline commit 8df3fb0 in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 3 performance regressions! Performance is the same for 590 metrics, 10 unstable metrics, 7 known flaky benchmarks, 17 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:iastaspects-join_noaspect

  • 🟥 execution_time [+11.692µs; +14.917µs] or [+7.801%; +9.953%]

scenario:iastaspects-title_aspect

  • 🟥 execution_time [+52.387µs; +57.914µs] or [+18.383%; +20.322%]

scenario:samplingrules-high_match

  • 🟥 execution_time [+13.968µs; +16.196µs] or [+9.061%; +10.507%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-724.867ns; +740.773ns] or [-6.574%; +6.718%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-36.786ns; +28.957ns] or [-5.969%; +4.698%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1694.220ns; +1585.687ns] or [-9.946%; +9.309%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-1410.067ns; +1062.994ns] or [-10.714%; +8.077%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-324.070ns; +323.979ns] or [-8.843%; +8.840%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-237.562ns; +266.592ns] or [-8.155%; +9.151%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-76.576ns; +74.987ns] or [-6.672%; +6.534%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-3744.660ns; +4201.105ns] or [-9.202%; +10.324%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-724.895ns; +876.468ns] or [-8.928%; +10.794%]

scenario:packagesupdateimporteddependencies-import_many_stdlib_cached

  • unstable execution_time [-49.900µs; +62.367µs] or [-8.515%; +10.643%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+1.933µs; +2.059µs] or [+11.476%; +12.226%]

scenario:iastaspects-casefold_noaspect

  • 🟥 execution_time [+30.200µs; +36.180µs] or [+11.936%; +14.299%]

scenario:iastaspectsospath-ospathbasename_aspect

  • 🟥 execution_time [+121.133µs; +128.015µs] or [+29.507%; +31.184%]

scenario:iastaspectssplit-rsplit_aspect

  • 🟥 execution_time [+13.983µs; +19.082µs] or [+9.752%; +13.308%]

scenario:span-start

  • 🟥 execution_time [+1.645ms; +1.820ms] or [+11.671%; +12.916%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+325.991ns; +368.832ns] or [+11.242%; +12.719%]

scenario:tracer-small

  • 🟥 execution_time [+41.862µs; +44.419µs] or [+12.925%; +13.714%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:iastaspects-casefold_aspect
  • scenario:iastaspects-index_aspect
  • scenario:iastaspects-ljust_noaspect
  • scenario:iastaspects-lower_aspect
  • scenario:iastaspects-replace_aspect
  • scenario:iastaspects-rstrip_aspect
  • scenario:iastaspects-swapcase_aspect
  • scenario:iastaspects-title_noaspect
  • scenario:iastaspects-translate_aspect
  • scenario:iastaspects-translate_noaspect
  • scenario:iastaspects-upper_noaspect
  • scenario:packagespackageforrootmodulemapping-cache_off
  • scenario:packagespackageforrootmodulemapping-cache_on
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

@mz1119
mz1119 force-pushed the max.zhang/llmobs-gen-ai-apm-tags branch from 38aecdc to 9c1eb5e Compare September 8, 2026 15:17
@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against main using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

No remaining files require a CODEOWNERS review.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 228 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 228 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=134)
ddtrace.llmobs._integrations.anthropic -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=132)
ddtrace.llmobs._integrations.mcp -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=132)
ddtrace.internal.opentelemetry.trace -×-> ddtrace.trace  (product:opentelemetry -> product:tracing, score=132)
ddtrace.internal.openfeature._span_enrichment -×-> ddtrace.trace  (product:openfeature -> product:tracing, score=132)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@mz1119
mz1119 force-pushed the max.zhang/llmobs-gen-ai-apm-tags branch from 9c1eb5e to 3993b5c Compare September 8, 2026 15:43
@mz1119
mz1119 marked this pull request as ready for review September 8, 2026 17:22
@mz1119
mz1119 requested review from a team as code owners September 8, 2026 17:22
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-08T17:33:37.982750Z 3993b5c Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3993b5cb29

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/llmobs/_integrations/base.py
Comment thread ddtrace/llmobs/_llmobs.py Outdated
Comment thread tests/utils.py
Comment thread ddtrace/llmobs/_utils.py
@mz1119
mz1119 force-pushed the max.zhang/llmobs-gen-ai-apm-tags branch 2 times, most recently from 4a605e8 to 1e46f6a Compare September 8, 2026 20:54
The five mirrored token metrics left out reasoning_output_tokens, which the
openai, vertexai, litellm, and google integrations already collect into the
LLMObs event. It is a per-request count the provider returns, on the same
footing as the other five, and reasoning-token spend is worth monitoring on
directly in APM.

Emitted as gen_ai.usage.reasoning_output_tokens, gated to llm and embedding
kinds like the rest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Max Zhang <max.zhang@datadoghq.com>

@ncybul ncybul left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few questions and suggestions, but overall the logic makes sense to me! What might be good is to query APM's span search API to verify that the tags are on the APM span correctly and attach the link to the APM span with the payload you searched to the PR description so it's easy to reference.

Comment thread releasenotes/notes/gen-ai-apm-span-tags-fc48a9459da0a6f9.yaml Outdated
Comment thread tests/llmobs/test_llmobs_gen_ai_apm_tags.py
Comment thread tests/utils.py Outdated
Comment thread ddtrace/llmobs/_integrations/base.py
Comment thread ddtrace/llmobs/_utils.py
mz1119 and others added 2 commits September 9, 2026 14:55
Co-authored-by: ncybul <124532568+ncybul@users.noreply.github.com>
Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
Co-authored-by: ncybul <124532568+ncybul@users.noreply.github.com>
Signed-off-by: Max Zhang <max.zhang@datadoghq.com>

@emmettbutler emmettbutler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Release note lgtm

@mz1119
mz1119 force-pushed the max.zhang/llmobs-gen-ai-apm-tags branch from cb80890 to 1200608 Compare September 9, 2026 20:38
@mz1119

mz1119 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-10 18:56:47 UTC ℹ️ Start processing command /merge


2026-09-10 18:57:06 UTC ℹ️ MergeQueue: Pull request is not mergeable yet

It will be processed automatically as soon as GitHub reports it as mergeable. View in MergeQueue UI.

  • Run /code blockers to see what is blocking it.
  • Run /remove to cancel it.

2026-09-10 19:11:29 UTC ℹ️ MergeQueue: merge request added to the queue

The expected merge time in main is approximately 55m (p90).


2026-09-10 19:53:48 UTC ❌ MergeQueue: The checks failed on this merge request

Tests failed on this commit 0602375:

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.

@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot merged commit c8f94a2 into main Sep 10, 2026
1443 checks passed
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot deleted the max.zhang/llmobs-gen-ai-apm-tags branch September 10, 2026 20:41
brettlangdon pushed a commit that referenced this pull request Sep 14, 2026
## Description

The APM trace UI merges `gen_ai.*` keys into spans client-side from the LLMObs track, so the values render but aren't indexed — you can't filter, facet, or monitor on model, provider, or token usage in APM. web-ui code is[here](https://github.com/DataDog/web-ui/blob/preprod/packages/apps/apm/toolkit/hooks/use-llm-spans-by-apm-trace-id/enrich-trace-llm.ts)

This change emits a subset from tracer instead, making them real searchable APM tags: `gen_ai.operation.name`, `gen_ai.request.model`, `gen_ai.provider.name`, `gen_ai.application.name`, and `gen_ai.conversation.id`, plus `gen_ai.usage.{input,output,total,cache_read_input,cache_write_input,reasoning_output}_tokens` as metrics.

Message bodies (input, output, tool definitions, retrieval documents) stay off the APM span and continue to come from the LLMObs track.

## Testing

gen_ai fields are queryable in APM if sent with changes made in the SDK. [Example span](https://app.datadoghq.com/apm/trace/6aa1b48e000000002d4bfc9b713c72aa?graphType=json&shouldShowLegend=true&spanID=12232979254673635221&timeHint=1788982415825&trace=AwAAAaCHqVHRRDs8egAAABhBYUNIcVZKTUFBQTFZWDhhbDRaRnlrc3QAAAAkZjFhMDg3YWMtOWEzZS00NDhlLWFmN2ItMWUzMTIwZjE0NjVhAAAMLQ&traceQuery=)

Querying that span through the APM span API shows the tags in the payload (trimmed to the relevant keys):

```json
"custom": {
  "component": "openai",
  "gen_ai": {
    "application": { "name": "max-test" },
    "operation": { "name": "llm" },
    "provider": { "name": "openai" },
    "request": { "model": "gpt-4.1-mini-2025-04-14" },
    "usage": {
      "input_tokens": 13,
      "output_tokens": 10,
      "total_tokens": 23,
      "cache_read_input_tokens": 0,
      "cache_write_input_tokens": 0,
      "reasoning_output_tokens": 0
    }
  }
}
```

Unit coverage in `tests/llmobs/test_llmobs_gen_ai_apm_tags.py`.


Co-authored-by: max.zhang <max.zhang@datadoghq.com>
gh-worker-dd-mergequeue-cf854d Bot pushed a commit that referenced this pull request Sep 16, 2026
## Description

#20083 added `gen_ai.*` attributes to APM spans through two paths: the LLMObs span-finish path, and `_apply_shadow_metrics` when `llmobs_enabled` is `False`. 

It turns out that these gen_ai tags cause apm spans to be switched back to llmobs spans by spanIsRelevantForLLMObs in dd-go/trace/apps/trace-router/processors/llmobs/processor.go

## Testing

llmobs enabled span: https://app.datadoghq.com/llm/traces/trace/6aaaf8fe0000000098c1c60fa44fd043?is_llm_session=false&selectedTab=overview&spanId=8751386663934863565
https://app.datadoghq.com/apm/trace/6aaaf8fe0000000079b891c7157810fd?graphType=json&shouldShowLegend=true&trace__spanID=8751386663934863565&traceQuery=


llmobs disabled span: 
https://app.datadoghq.com/apm/trace/6aaaf8f40000000039da33fef2e16e51?graphType=json&shouldShowLegend=true&trace__spanID=4425947145328519374&traceQuery=

## Risks

Low. Customers who enabled LLMObs see no change. Anyone who started querying `gen_ai.*` in APM without LLMObs enabled loses those tags, which is the intent.

## Additional Notes

Follow-up to #20083.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: max.zhang <max.zhang@datadoghq.com>
github-actions Bot added a commit that referenced this pull request Sep 17, 2026
## Description

#20083 added `gen_ai.*` attributes to APM spans through two paths: the LLMObs span-finish path, and `_apply_shadow_metrics` when `llmobs_enabled` is `False`.

It turns out that these gen_ai tags cause apm spans to be switched back to llmobs spans by spanIsRelevantForLLMObs in dd-go/trace/apps/trace-router/processors/llmobs/processor.go

## Testing

llmobs enabled span: https://app.datadoghq.com/llm/traces/trace/6aaaf8fe0000000098c1c60fa44fd043?is_llm_session=false&selectedTab=overview&spanId=8751386663934863565
https://app.datadoghq.com/apm/trace/6aaaf8fe0000000079b891c7157810fd?graphType=json&shouldShowLegend=true&trace__spanID=8751386663934863565&traceQuery=

llmobs disabled span:
https://app.datadoghq.com/apm/trace/6aaaf8f40000000039da33fef2e16e51?graphType=json&shouldShowLegend=true&trace__spanID=4425947145328519374&traceQuery=

## Risks

Low. Customers who enabled LLMObs see no change. Anyone who started querying `gen_ai.*` in APM without LLMObs enabled loses those tags, which is the intent.

## Additional Notes

Follow-up to #20083.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: max.zhang <max.zhang@datadoghq.com>
(cherry picked from commit a3af7c9)

Co-authored-by: Max Zhang <40515363+mz1119@users.noreply.github.com>
mz1119 added a commit that referenced this pull request Sep 17, 2026
## Description

Prepares the 4.15.1 patch release at the current head of the `4.15`
branch.

- sets the version string in `pyproject.toml` to `4.15.1` (the branch
was left at `4.15.0` after that release shipped);
- refreshes the pinned `DataDog/system-tests` revision (`de534bc7bc` ->
`1fc149bfb1`) so it validates the unreleased `4.15` change.

The only unreleased commit on `4.15` since `v4.15.0` is #20401, the
backport of #20375, which fixes a bug introduced by #20083 in 4.15.0.

## Testing

Covered by the `4.15` branch CI and the system-tests pipelines this PR
repins.

## Risks

Low. Version string and CI pin only, no runtime behavior change.

## Additional Notes

After this PR merges and release checks pass (including
`check-slo-breaches`), its merge commit should be used as the target for
the 4.15.1 GitHub release draft.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
gh-worker-dd-mergequeue-cf854d Bot pushed a commit that referenced this pull request Sep 21, 2026
## Summary
- When `set_gen_ai_apm_tags` writes `gen_ai.*` attributes onto an APM span, it now also sets an internal `_dd.llmobs.artificial_gen_ai_tags` tag.
- This lets our backend processor distinguish gen_ai tags that ddtrace added artificially from tags a user set directly on the span. Uses the existing `_dd.` prefix convention for internal, non-user-searchable tags (see `_dd.llmobs.*` shadow tags in `ddtrace/llmobs/_constants.py`).

Follow-up to #20083 / #20375.

## Testing plan
- Extended `tests/llmobs/test_llmobs_gen_ai_apm_tags.py` to assert the new tag is set alongside the existing `gen_ai.*` tags.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: max.zhang <max.zhang@datadoghq.com>
brettlangdon pushed a commit that referenced this pull request Sep 23, 2026
## Summary
- When `set_gen_ai_apm_tags` writes `gen_ai.*` attributes onto an APM span, it now also sets an internal `_dd.llmobs.artificial_gen_ai_tags` tag.
- This lets our backend processor distinguish gen_ai tags that ddtrace added artificially from tags a user set directly on the span. Uses the existing `_dd.` prefix convention for internal, non-user-searchable tags (see `_dd.llmobs.*` shadow tags in `ddtrace/llmobs/_constants.py`).

Follow-up to #20083 / #20375.

## Testing plan
- Extended `tests/llmobs/test_llmobs_gen_ai_apm_tags.py` to assert the new tag is set alongside the existing `gen_ai.*` tags.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: max.zhang <max.zhang@datadoghq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants