Skip to content

chore(agents): add dd-trace-java overrides for dd-apm-sdk-review - #12460

Open
robertomonteromiguel wants to merge 13 commits into
robertomonteromiguel/dd-apm-sdk-review-core-copyfrom
robertomonteromiguel/dd-apm-sdk-review-java-overrides
Open

chore(agents): add dd-trace-java overrides for dd-apm-sdk-review#12460
robertomonteromiguel wants to merge 13 commits into
robertomonteromiguel/dd-apm-sdk-review-core-copyfrom
robertomonteromiguel/dd-apm-sdk-review-java-overrides

Conversation

@robertomonteromiguel

Copy link
Copy Markdown
Contributor

What does this PR do?

Adds the dd-trace-java layer on top of the shared skill copy.

  • .agents/dd-apm-sdk-review-overrides/ — Java reviewer facts (performance, conventions, design, security, maintainability, repo context)
  • AGENTS.md — short hook: local agents run the skill; Codex reads review-without-harness.md
  • .llm-validation/ — cases that exercise those rules, plus the GitLab "llm validation" include
  • Replaces the retired /perf-review skill (deleted here; its Java-specific rubric lives in the performance override)

Stacked on the core-copy PR. Shared skill source: dd-apm-sdk-review-core#1.

Motivation

This is the half of #12364 that Java reviewers should actually read.

Additional Notes

How to review

  • Start here: .agents/dd-apm-sdk-review-overrides/
  • Then: the AGENTS.md Review Guidelines block
  • Then: .llm-validation/ if you care about the gate
  • Skip .agents/skills/dd-apm-sdk-review/ — that is the parent PR / core repo

Merge with the parent skill-copy PR, not as a standalone.

Made with Cursor

Repo-specific reviewer facts, AGENTS.md hook, llm-validation suite,
and replace the retired perf-review skill. Stacked on the verbatim
skill copy.
The previous commit copied .gitlab-ci.yml from the old combined
branch and dropped unrelated master changes.
@robertomonteromiguel

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 2ab24b4dc9

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

🎯 Code Coverage (details)
Patch Coverage: 100.00%
Overall Coverage: 59.13% (+0.00%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 705a16a | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.82 s 14.70 s [-0.0%; +1.6%] (no difference)
startup:insecure-bank:tracing:Agent 13.70 s 13.78 s [-1.5%; +0.2%] (no difference)
startup:petclinic:appsec:Agent 17.58 s 17.42 s [+0.1%; +1.7%] (maybe worse)
startup:petclinic:iast:Agent 17.38 s 17.60 s [-2.2%; -0.3%] (maybe better)
startup:petclinic:profiling:Agent 16.67 s 17.19 s [-7.1%; +1.1%] (no difference)
startup:petclinic:sca:Agent 17.62 s 17.27 s [+1.1%; +3.0%] (significantly worse)
startup:petclinic:tracing:Agent 16.57 s 16.68 s [-1.7%; +0.3%] (no difference)

Commit: 705a16a0 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

Keep the stacked overrides branch on the latest verbatim skill copy.
Use the current platform SHA for the GitLab include and the job clone.
@robertomonteromiguel

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f8f15a0d73

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .llm-validation/suites/dd-apm-sdk-review.yaml
Comment thread .llm-validation/suites/dd-apm-sdk-review.yaml Outdated
Comment thread .llm-validation/config.yaml
Keep the stacked overrides branch on the latest verbatim skill copy.
Evaluate crash-handler ordering as a security-lens report, not the
orchestrator's redacted output. Require firstSpanWritten to reset on
the payload-flush hook, not at every map().
@robertomonteromiguel

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7a96d17e1f

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .agents/dd-apm-sdk-review-overrides/reviewers/design.md Outdated
Comment thread .llm-validation/suites/dd-apm-sdk-review.yaml
Comment thread .agents/dd-apm-sdk-review-overrides/reviewers/performance.md Outdated
Don't flag feature-flagging modules against an aspirational -lib split;
point J1 at EscapeShapeBenchmark; make the HashMap fixture lookup-only.
Also merge the skill-copy secret-scan sync.
…robertomonteromiguel/dd-apm-sdk-review-java-overrides
@robertomonteromiguel

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 84bb97ad43

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .agents/dd-apm-sdk-review-overrides/reviewers/design.md Outdated
Comment thread .llm-validation/suites/dd-apm-sdk-review.yaml
@pr-commenter

pr-commenter Bot commented Sep 11, 2026

Copy link
Copy Markdown

LLM Validation

LLM Validation Gate — dd-apm-sdk-review

✅ PASS

  • Overall quality improved by 5.2 points with no blocking-case regressions.

Analysis

Changed instruction file(s): .agents/skills/dd-apm-sdk-review/SKILL.md, .agents/skills/dd-apm-sdk-review/reviewers/_common.md, .agents/skills/dd-apm-sdk-review/reviewers/correctness.md, .agents/skills/dd-apm-sdk-review/reviewers/performance.md, .agents/skills/dd-apm-sdk-review/reviewers/report-template.md, .agents/skills/dd-apm-sdk-review/review-without-harness.md, .agents/skills/dd-apm-sdk-review/reviewers/coherence.md, .agents/skills/dd-apm-sdk-review/reviewers/security.md, .agents/skills/dd-apm-sdk-review/reviewers/design.md, .agents/skills/dd-apm-sdk-review/reviewers/maintainability.md, .agents/skills/dd-apm-sdk-review/reviewers/conventions.md, .agents/skills/dd-apm-sdk-review/reviewers/cross-sdk.md, .agents/dd-apm-sdk-review-overrides/repo-context.md, .agents/dd-apm-sdk-review-overrides/reviewers/security.md, .agents/dd-apm-sdk-review-overrides/reviewers/performance.md, .agents/dd-apm-sdk-review-overrides/reviewers/design.md, .agents/dd-apm-sdk-review-overrides/reviewers/conventions.md, .agents/dd-apm-sdk-review-overrides/reviewers/maintainability.md.

No safety or blocking-case regressions across 8 case(s). Overall pairwise win-rate 66% [58%–73%], quality +5.2 — see the verdict above for whether that clears the noise band.

Results

  • Pairwise win-rate: 66% [58%–73%] — candidate's share of blind comparisons (90% CI; spanning 50% = no clear difference)
  • Overall quality: 81.1 → 86.4 (/100, +5.2)
  • Bad signals introduced (advisory): 0
  • Candidate criteria coverage (advisory): 29/31 (94%) — expected_criteria the candidate met; does not affect the gate
  • Blocking-case regressions: 0

Cases

Case Mode Quality Δ Win-rate (90% CI) Safety
java-perf-lens-wrong-collection-001 block 0.0 50% [50%–50%] ok
java-perf-pipeline-full-review-002 block +15.0 75% [51%–99%] ok
java-security-crash-handler-before-trust block -0.8 50% [50%–50%] ok
java-correctness-capture-before-send block +16.7 100% [100%–100%] ok
java-correctness-sqs-queue-name-incomplete block +14.2 100% [100%–100%] ok
java-maintainability-resource-leak-streams block -5.8 38% [17%–58%] ok
java-correctness-span-events-list-only block +1.3 62% [42%–83%] ok
java-correctness-mapper-state-leak block +1.3 50% [50%–50%] ok

Per-dimension scores, token usage, latency, and estimated cost are in the CI job logs.

Auto-instrumentations still require InstrumenterModule; product
transformers do not. The full-review fixture now states the J10/J12
preconditions the performance override already requires.
@robertomonteromiguel

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: 8cbdbb070e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@robertomonteromiguel
robertomonteromiguel marked this pull request as ready for review September 11, 2026 14:44
@robertomonteromiguel
robertomonteromiguel requested review from bric3 and erikayasuda and removed request for a team September 11, 2026 14:44

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8cbdbb070e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread AGENTS.md

@datadog-datadog-prod-us1 datadog-datadog-prod-us1 Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: FAIL

The gate does not test the review contract without a skill harness. The security case also accepts P1 although the new rule requires P0.

Open Bits AI session

🤖 Datadog Autotest · Commit 8cbdbb0 · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

Comment thread .llm-validation/config.yaml
Comment thread .llm-validation/suites/dd-apm-sdk-review.yaml Outdated
@dougqh

dougqh commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

The dd-gitlab/llm validation and dd-gitlab/default-pipeline checks are currently failing. Since this PR introduces the .llm-validation/ suite itself, worth confirming whether that's a real eval regression or first-run CI wiring before merging.

Also no labels yet — per this repo's conventions, please add:

  • a type: label
  • a comp:/inst: label
  • tag: ai generated

@robertomonteromiguel robertomonteromiguel added comp: tooling Build & Tooling tag: no release notes Changes to exclude from release notes tag: ai generated Largely based on code generated by an AI or LLM type: feature Enhancements and improvements labels Sep 14, 2026
…robertomonteromiguel/dd-apm-sdk-review-java-overrides
…robertomonteromiguel/dd-apm-sdk-review-java-overrides
The full-skill run was losing the gate to an uninstructed baseline.
Evaluate firstSpanWritten the same way as the JS correctness cases.
The security override already calls this P0; the eval case should
not accept P1.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: tooling Build & Tooling tag: ai generated Largely based on code generated by an AI or LLM tag: no release notes Changes to exclude from release notes type: feature Enhancements and improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants