feat(llmobs): emit gen_ai.* attributes on APM spans - #20083
Conversation
The APM trace UI reproduces gen_ai.* keys today by querying the LLMObs track
and merging the result into each span client-side. Those merged values are
displayable but not searchable, so you cannot filter, facet, or monitor on
model, provider, or token usage in APM.
Emit the scalar subset from the tracer instead, so they become real APM tags:
meta: gen_ai.operation.name, gen_ai.request.model, gen_ai.provider.name,
gen_ai.application.name, gen_ai.conversation.id
metrics: gen_ai.usage.{input,output,total,cache_read_input,
cache_write_input}_tokens
Message bodies (input, output, tool definitions, retrieval documents) stay off
the APM span: they are unbounded and not usefully queryable once serialized.
The UI keeps joining the LLMObs track for those.
Emission is centralized in LLMObs._prepare_llmobs_span_data, before the user
span processor and _normalize_llmobs_meta, so it covers every integration, the
decorators, and manual LLMObs.annotate() from one call site. Running early
keeps the tags on spans whose LLMObs event a processor drops, and reads
model_name and model_provider before normalization pops them for kinds other
than llm/embedding; LLMObsSpan exposes none of those fields, so nothing is
lost by not waiting. BaseLLMIntegration's shadow-tag path covers the
LLMObs-disabled case, where no meta_struct exists, with the same per-
integration coverage as the existing _dd.llmobs.* shadow tags.
gen_ai.operation.name carries the raw LLMObs span kind rather than the OTel
gen_ai enum, matching what the UI already writes into that key so the
tracer-emitted and enriched values agree.
The two helpers live in llmobs/_utils.py rather than a dedicated module. A new
module importing Span would add a fresh product:llmobs -> product:tracing edge
and join the llmobs import tangle, which detect_layering_violations and
detect_circular_imports both reject. _utils already owns the other llmobs
span/meta_struct accessors and already carries that edge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
🎉 All green!🧪 All tests passed 🔗 Commit SHA: 599c4fe | Docs | View more details | Give us feedback! |
BenchmarksBenchmark execution time: 2026-09-11 15:54:14 Comparing candidate commit f9254a2 in PR branch Found 0 performance improvements and 3 performance regressions! Performance is the same for 590 metrics, 10 unstable metrics, 7 known flaky benchmarks, 17 flaky benchmarks without significant changes.
|
38aecdc to
9c1eb5e
Compare
Codeowners resolved asResolved from the full PR diff against No remaining files require a CODEOWNERS review. |
Circular import analysis
|
Dependency direction analysis
|
9c1eb5e to
3993b5c
Compare
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3993b5cb29
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
4a605e8 to
1e46f6a
Compare
The five mirrored token metrics left out reasoning_output_tokens, which the openai, vertexai, litellm, and google integrations already collect into the LLMObs event. It is a per-request count the provider returns, on the same footing as the other five, and reasoning-token spend is worth monitoring on directly in APM. Emitted as gen_ai.usage.reasoning_output_tokens, gated to llm and embedding kinds like the rest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
ncybul
left a comment
There was a problem hiding this comment.
A few questions and suggestions, but overall the logic makes sense to me! What might be good is to query APM's span search API to verify that the tags are on the APM span correctly and attach the link to the APM span with the payload you searched to the PR description so it's easy to reference.
Co-authored-by: ncybul <124532568+ncybul@users.noreply.github.com> Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
Co-authored-by: ncybul <124532568+ncybul@users.noreply.github.com> Signed-off-by: Max Zhang <max.zhang@datadoghq.com>
cb80890 to
1200608
Compare
|
/merge |
|
View all feedbacks in Devflow UI.
It will be processed automatically as soon as GitHub reports it as mergeable. View in MergeQueue UI.
The expected merge time in
Tests failed on this commit 0602375: What to do next?
|
## Description The APM trace UI merges `gen_ai.*` keys into spans client-side from the LLMObs track, so the values render but aren't indexed — you can't filter, facet, or monitor on model, provider, or token usage in APM. web-ui code is[here](https://github.com/DataDog/web-ui/blob/preprod/packages/apps/apm/toolkit/hooks/use-llm-spans-by-apm-trace-id/enrich-trace-llm.ts) This change emits a subset from tracer instead, making them real searchable APM tags: `gen_ai.operation.name`, `gen_ai.request.model`, `gen_ai.provider.name`, `gen_ai.application.name`, and `gen_ai.conversation.id`, plus `gen_ai.usage.{input,output,total,cache_read_input,cache_write_input,reasoning_output}_tokens` as metrics. Message bodies (input, output, tool definitions, retrieval documents) stay off the APM span and continue to come from the LLMObs track. ## Testing gen_ai fields are queryable in APM if sent with changes made in the SDK. [Example span](https://app.datadoghq.com/apm/trace/6aa1b48e000000002d4bfc9b713c72aa?graphType=json&shouldShowLegend=true&spanID=12232979254673635221&timeHint=1788982415825&trace=AwAAAaCHqVHRRDs8egAAABhBYUNIcVZKTUFBQTFZWDhhbDRaRnlrc3QAAAAkZjFhMDg3YWMtOWEzZS00NDhlLWFmN2ItMWUzMTIwZjE0NjVhAAAMLQ&traceQuery=) Querying that span through the APM span API shows the tags in the payload (trimmed to the relevant keys): ```json "custom": { "component": "openai", "gen_ai": { "application": { "name": "max-test" }, "operation": { "name": "llm" }, "provider": { "name": "openai" }, "request": { "model": "gpt-4.1-mini-2025-04-14" }, "usage": { "input_tokens": 13, "output_tokens": 10, "total_tokens": 23, "cache_read_input_tokens": 0, "cache_write_input_tokens": 0, "reasoning_output_tokens": 0 } } } ``` Unit coverage in `tests/llmobs/test_llmobs_gen_ai_apm_tags.py`. Co-authored-by: max.zhang <max.zhang@datadoghq.com>
## Description #20083 added `gen_ai.*` attributes to APM spans through two paths: the LLMObs span-finish path, and `_apply_shadow_metrics` when `llmobs_enabled` is `False`. It turns out that these gen_ai tags cause apm spans to be switched back to llmobs spans by spanIsRelevantForLLMObs in dd-go/trace/apps/trace-router/processors/llmobs/processor.go ## Testing llmobs enabled span: https://app.datadoghq.com/llm/traces/trace/6aaaf8fe0000000098c1c60fa44fd043?is_llm_session=false&selectedTab=overview&spanId=8751386663934863565 https://app.datadoghq.com/apm/trace/6aaaf8fe0000000079b891c7157810fd?graphType=json&shouldShowLegend=true&trace__spanID=8751386663934863565&traceQuery= llmobs disabled span: https://app.datadoghq.com/apm/trace/6aaaf8f40000000039da33fef2e16e51?graphType=json&shouldShowLegend=true&trace__spanID=4425947145328519374&traceQuery= ## Risks Low. Customers who enabled LLMObs see no change. Anyone who started querying `gen_ai.*` in APM without LLMObs enabled loses those tags, which is the intent. ## Additional Notes Follow-up to #20083. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: max.zhang <max.zhang@datadoghq.com>
## Description #20083 added `gen_ai.*` attributes to APM spans through two paths: the LLMObs span-finish path, and `_apply_shadow_metrics` when `llmobs_enabled` is `False`. It turns out that these gen_ai tags cause apm spans to be switched back to llmobs spans by spanIsRelevantForLLMObs in dd-go/trace/apps/trace-router/processors/llmobs/processor.go ## Testing llmobs enabled span: https://app.datadoghq.com/llm/traces/trace/6aaaf8fe0000000098c1c60fa44fd043?is_llm_session=false&selectedTab=overview&spanId=8751386663934863565 https://app.datadoghq.com/apm/trace/6aaaf8fe0000000079b891c7157810fd?graphType=json&shouldShowLegend=true&trace__spanID=8751386663934863565&traceQuery= llmobs disabled span: https://app.datadoghq.com/apm/trace/6aaaf8f40000000039da33fef2e16e51?graphType=json&shouldShowLegend=true&trace__spanID=4425947145328519374&traceQuery= ## Risks Low. Customers who enabled LLMObs see no change. Anyone who started querying `gen_ai.*` in APM without LLMObs enabled loses those tags, which is the intent. ## Additional Notes Follow-up to #20083. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: max.zhang <max.zhang@datadoghq.com> (cherry picked from commit a3af7c9) Co-authored-by: Max Zhang <40515363+mz1119@users.noreply.github.com>
## Description Prepares the 4.15.1 patch release at the current head of the `4.15` branch. - sets the version string in `pyproject.toml` to `4.15.1` (the branch was left at `4.15.0` after that release shipped); - refreshes the pinned `DataDog/system-tests` revision (`de534bc7bc` -> `1fc149bfb1`) so it validates the unreleased `4.15` change. The only unreleased commit on `4.15` since `v4.15.0` is #20401, the backport of #20375, which fixes a bug introduced by #20083 in 4.15.0. ## Testing Covered by the `4.15` branch CI and the system-tests pipelines this PR repins. ## Risks Low. Version string and CI pin only, no runtime behavior change. ## Additional Notes After this PR merges and release checks pass (including `check-slo-breaches`), its merge commit should be used as the target for the 4.15.1 GitHub release draft. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Summary - When `set_gen_ai_apm_tags` writes `gen_ai.*` attributes onto an APM span, it now also sets an internal `_dd.llmobs.artificial_gen_ai_tags` tag. - This lets our backend processor distinguish gen_ai tags that ddtrace added artificially from tags a user set directly on the span. Uses the existing `_dd.` prefix convention for internal, non-user-searchable tags (see `_dd.llmobs.*` shadow tags in `ddtrace/llmobs/_constants.py`). Follow-up to #20083 / #20375. ## Testing plan - Extended `tests/llmobs/test_llmobs_gen_ai_apm_tags.py` to assert the new tag is set alongside the existing `gen_ai.*` tags. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: max.zhang <max.zhang@datadoghq.com>
## Summary - When `set_gen_ai_apm_tags` writes `gen_ai.*` attributes onto an APM span, it now also sets an internal `_dd.llmobs.artificial_gen_ai_tags` tag. - This lets our backend processor distinguish gen_ai tags that ddtrace added artificially from tags a user set directly on the span. Uses the existing `_dd.` prefix convention for internal, non-user-searchable tags (see `_dd.llmobs.*` shadow tags in `ddtrace/llmobs/_constants.py`). Follow-up to #20083 / #20375. ## Testing plan - Extended `tests/llmobs/test_llmobs_gen_ai_apm_tags.py` to assert the new tag is set alongside the existing `gen_ai.*` tags. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: max.zhang <max.zhang@datadoghq.com>
Description
The APM trace UI merges
gen_ai.*keys into spans client-side from the LLMObs track, so the values render but aren't indexed — you can't filter, facet, or monitor on model, provider, or token usage in APM. web-ui code ishereThis change emits a subset from tracer instead, making them real searchable APM tags:
gen_ai.operation.name,gen_ai.request.model,gen_ai.provider.name,gen_ai.application.name, andgen_ai.conversation.id, plusgen_ai.usage.{input,output,total,cache_read_input,cache_write_input,reasoning_output}_tokensas metrics.Message bodies (input, output, tool definitions, retrieval documents) stay off the APM span and continue to come from the LLMObs track.
Testing
gen_ai fields are queryable in APM if sent with changes made in the SDK. Example span
Querying that span through the APM span API shows the tags in the payload (trimmed to the relevant keys):
Unit coverage in
tests/llmobs/test_llmobs_gen_ai_apm_tags.py.