Skip to content

feat(aiguard): evaluate OpenAI hosted MCP tool calls - #20850

Draft
avara1986 wants to merge 3 commits into
mainfrom
studio/aiguard-mcp-delivery-one
Draft

avara1986 wants to merge 3 commits into
mainfrom
studio/aiguard-mcp-delivery-one

Conversation

@avara1986

@avara1986 avara1986 commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Description

Jira: APPSEC-70369

Adds MCP provenance to AI Guard tool-call evaluations for the remote MCP tools the OpenAI Responses API runs on the application's behalf (tools=[{"type": "mcp", ...}], see the OpenAI MCP guide).

Contract. ToolCall gains an optional mcp object (new public MCP TypedDict):

Field Required Value
transport yes unknown for OpenAI hosted MCP (OpenAI does not report streamable HTTP vs SSE)
tool_name yes MCP tool name
name no the tool's server_label
url no the tool's server_url, reduced to scheme, host, port and path; absent for connectors and tunnels

function.name and the rest of the payload are unchanged, so ordinary tool calls keep wire compatibility.

Behaviour, gated by the new DD_AI_GUARD_COLLECT_MCP_ENABLED (default false until the backend validates the optional object):

  • mcp_call items, already evaluated today, now carry the mcp object. OpenAI has already run these calls, so a block only keeps the result from the application.
  • mcp_approval_request items (tools with require_approval) are now evaluated in the after-listener, grouped into one assistant turn placed last so they are the evaluation target. A blocking verdict raises the OpenAI-compatible AIGuardAbortError before the application receives the response, so the call is never approved and never runs. The tool-call ID is OpenAI's approval request ID.
  • Replayed mcp_approval_request input items become history; mcp_approval_response items carry no content and are skipped.
  • The AI Guard span gets ai_guard.mcp.{tool_name,name,url,transport} tags (provisional names pending span alignment).

URL sanitization lives in a new shared helper, ddtrace.internal.utils.http.canonicalize_url. The authorization token and headers of the tool definition are never read.

With the flag off, behaviour is identical to today.

Testing

  • tests/aiguard/openai/test_responses.py: URL map and gating, mcp_call enrichment with and without a configured URL, approval-request ordering, input replay, after-listener payload, and SDK-level calls through a mocked transport: approval blocked (DENY/ABORT, no secrets in the payload), allowed (span tags), and not evaluated when the flag is off.
  • tests/tracer/test_utils.py: canonicalize_url cases (userinfo/query/fragment removal, default ports, IPv6, invalid input).
  • Locally: ai_guard_openai (latest openai and 1.102.0) and ai_guard_api suites pass.

Risks

  • Off by default. When enabled, a response with approval requests triggers its usual single after-evaluation, now targeting the approval requests; there is no extra request.
  • Several approval requests in one response share one evaluation, so a block stops all of them.

Additional Notes

  • Only the Responses API is covered; Chat Completions has no MCP tool type in the Python SDK.
  • DD_AI_GUARD_COLLECT_MCP_ENABLED still needs adding to the feature-parity registry.

🤖 Generated with Claude Code

Add an optional mcp object to the AI Guard tool-call contract (transport,
original tool name, server label and sanitized server URL) and populate it
for the remote MCP tools the OpenAI Responses API runs on the
application's behalf.

When DD_AI_GUARD_COLLECT_MCP_ENABLED is set (false by default):
- mcp_call items are evaluated with their MCP metadata.
- mcp_approval_request items are evaluated before the response reaches
  the application, so a blocking verdict prevents the call from ever
  being approved and run.

Server URLs are reduced to scheme, host, port and path; credentials,
query, fragment, authorization tokens and headers are never sent.

APPSEC-70369

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.llmobs._integrations.bedrock -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.aiguard._api_client -×-> ddtrace.trace  (product:aiguard -> product:tracing, score=130)
ddtrace.internal.ci_visibility.filters -×-> ddtrace.trace  (product:ci_visibility -> product:tracing, score=130)
ddtrace.internal.opentelemetry.context -×-> ddtrace.trace  (product:opentelemetry -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Pipelines  Tests

✨ Unblock PR with BitsAI

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 1 Pipeline job failed

DataDog/apm-reliability/dd-trace-py | validate_supported_configurations_v2_local_file — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: db17dae | Docs | View more details | Give us feedback!

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against main using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

.cursor/rules/ai-guard.mdc                                              @DataDog/asm-python
ddtrace/aiguard/__init__.py                                             @DataDog/asm-python
ddtrace/aiguard/_api_client.py                                          @DataDog/asm-python
ddtrace/aiguard/_constants.py                                           @DataDog/asm-python
ddtrace/aiguard/_listener.py                                            @DataDog/asm-python
ddtrace/aiguard/_streaming.py                                           @DataDog/asm-python
ddtrace/aiguard/_types.py                                               @DataDog/asm-python
ddtrace/aiguard/integrations/_mcp.py                                    @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_responses.py                       @DataDog/asm-python
ddtrace/appsec/ai_guard/__init__.py                                     @DataDog/asm-python
ddtrace/internal/settings/_supported_configurations.py                  @DataDog/apm-python
ddtrace/internal/settings/aiguard.py                                    @DataDog/asm-python
ddtrace/internal/utils/http.py                                          @DataDog/apm-core-python
docs/configuration.rst                                                  @DataDog/python-guild
releasenotes/notes/aiguard-mcp-tool-calls-599955302184ee90.yaml         @DataDog/apm-python
supported-configurations.json                                           @DataDog/apm-python
tests/aiguard/api/test_compat_imports.py                                @DataDog/asm-python
tests/aiguard/openai/conftest.py                                        @DataDog/asm-python
tests/aiguard/openai/test_responses.py                                  @DataDog/asm-python
tests/tracer/test_utils.py                                              @DataDog/apm-sdk-capabilities-python

@pr-commenter

pr-commenter Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-10-07 15:04:01

Comparing candidate commit db17dae in PR branch studio/aiguard-mcp-delivery-one with baseline commit 049682f in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 8 performance regressions! Performance is the same for 540 metrics, 10 unstable metrics, 7 known flaky benchmarks, 17 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationextract-full_t_id_datadog_headers

  • 🟥 execution_time [+2.069µs; +2.195µs] or [+14.496%; +15.377%]

scenario:iastaspects-format_map_noaspect

  • 🟥 execution_time [+71.906µs; +77.083µs] or [+20.126%; +21.575%]

scenario:iastaspects-title_aspect

  • 🟥 execution_time [+89.439µs; +93.672µs] or [+33.819%; +35.419%]

scenario:iastaspectsremodule-re_expand_aspect

  • 🟥 execution_time [+115.239µs; +122.828µs] or [+18.217%; +19.416%]

scenario:msgpackencoderscenario-simple_one_span

  • 🟥 execution_time [+533.587ns; +575.441ns] or [+12.978%; +13.996%]

scenario:otelspan-start

  • 🟥 execution_time [+1.884ms; +2.690ms] or [+7.713%; +11.010%]

scenario:recursivecomputation-shallow

  • 🟥 execution_time [+54.692µs; +57.194µs] or [+7.800%; +8.157%]

scenario:span-start-finish-telemetry

  • 🟥 execution_time [+4.731ms; +4.954ms] or [+11.416%; +11.955%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-730.052ns; +736.325ns] or [-6.921%; +6.981%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-40.016ns; +38.980ns] or [-6.014%; +5.858%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1919.021ns; +1943.930ns] or [-9.532%; +9.655%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-2083.338ns; +1653.695ns] or [-10.637%; +8.444%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-393.066ns; +378.292ns] or [-9.084%; +8.743%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-228.348ns; +234.230ns] or [-8.441%; +8.659%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-95.651ns; +88.141ns] or [-7.088%; +6.531%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-4517.478ns; +4689.243ns] or [-9.472%; +9.832%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-879.721ns; +993.093ns] or [-8.700%; +9.821%]

scenario:packagesupdateimporteddependencies-import_many_stdlib_cached

  • unstable execution_time [-48.706µs; +59.398µs] or [-8.720%; +10.635%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+2.894µs; +3.045µs] or [+20.116%; +21.171%]

scenario:iastaspects-rstrip_aspect

  • 🟥 execution_time [+116.428µs; +122.823µs] or [+34.670%; +36.575%]

scenario:iastaspectsospath-ospathbasename_aspect

  • 🟥 execution_time [+141.917µs; +148.607µs] or [+36.561%; +38.285%]

scenario:iastaspectssplit-rsplit_aspect

  • 🟥 execution_time [+25.439µs; +30.158µs] or [+16.490%; +19.549%]

scenario:span-start

  • 🟥 execution_time [+1.081ms; +1.527ms] or [+8.670%; +12.246%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+217.746ns; +245.838ns] or [+11.166%; +12.606%]

scenario:tracer-small

  • 🟥 execution_time [+43.718µs; +45.291µs] or [+17.568%; +18.200%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:iastaspects-casefold_aspect
  • scenario:iastaspects-casefold_noaspect
  • scenario:iastaspects-index_aspect
  • scenario:iastaspects-ljust_noaspect
  • scenario:iastaspects-lower_aspect
  • scenario:iastaspects-replace_aspect
  • scenario:iastaspects-swapcase_aspect
  • scenario:iastaspects-title_noaspect
  • scenario:iastaspects-translate_aspect
  • scenario:iastaspects-translate_noaspect
  • scenario:iastaspects-upper_noaspect
  • scenario:packagespackageforrootmodulemapping-cache_off
  • scenario:packagespackageforrootmodulemapping-cache_on
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

Remember the AI Guard decision taken on each OpenAI hosted MCP approval
request (bounded, in process) and check every mcp_approval_response with
approve: true before the continuation lets OpenAI run the call:

- a decision taken when the request was returned is reused, so one
  operation is evaluated once and a blocked request stays blocked;
- an unknown request replayed in the input is evaluated;
- an approval with only previous_response_id and no known request is a
  logged coverage gap.

Approval checks also run for streaming requests, since they authorize a
tool call rather than inspect the response. Adds streaming, async,
monitor-mode and blocking-disabled hosted MCP coverage, and documents that
only require_approval prevents hosted MCP execution.

APPSEC-70369

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@avara1986

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T14:04:54.712186Z 403deb3 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 403deb398f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/integrations/_openai_responses.py
Comment thread ddtrace/aiguard/_api_client.py Outdated
"Evaluation",
"Function",
"ImageURL",
"MCP",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Re-export MCP from the compatibility package

Adding MCP to the public ddtrace.aiguard.__all__ without adding it to ddtrace.appsec.ai_guard._PUBLIC breaks the deprecated package's documented invariant that it forwards all top-level public symbols until removal. Consequently, from ddtrace.appsec.ai_guard import MCP raises ImportError even though the new type is advertised as public, unlike every other public AI Guard type.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in db17dae: MCP is now in ddtrace.appsec.ai_guard._PUBLIC, and test_compat_imports covers it.

Comment thread ddtrace/internal/utils/http.py Outdated
Comment thread ddtrace/aiguard/integrations/_mcp.py Outdated
Address review findings on OpenAI hosted MCP evaluation:

- Buffer streamed Responses calls whose hosted MCP tools can request
  approval when MCP collection is on and stream analysis is off, and
  evaluate their approval requests before any event reaches the
  application. Every approval request is now enforced when returned, so
  the approval decision cache only avoids duplicate evaluations.
- Report only the origin (scheme, host, port) of MCP server URLs, since
  credentials can sit in the path.
- Resolve MCP span tags of a tool result from the call it answers.
- Re-export MCP from the deprecated ddtrace.appsec.ai_guard package.
- Use a fork-safe lock for the approval decision cache.

APPSEC-70369

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant