Skip to content

feat(aiguard): evaluate openai-agents MCP server tool calls - #20852

Draft
avara1986 wants to merge 1 commit into
studio/aiguard-mcp-delivery-onefrom
studio/aiguard-mcp-openai-agents
Draft

avara1986 wants to merge 1 commit into
studio/aiguard-mcp-delivery-onefrom
studio/aiguard-mcp-openai-agents

Conversation

@avara1986

@avara1986 avara1986 commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Description

Jira: APPSEC-70369

Stacked on #20850 (OpenAI hosted MCP): this PR targets that branch and contains only the client-managed openai-agents part.

Evaluates the tool calls of MCP servers configured in the OpenAI Agents SDK (MCPServerStdio, MCPServerSse, MCPServerStreamableHttp) before the tools/call request is sent, behind DD_AI_GUARD_COLLECT_MCP_ENABLED (default false).

Detection on the model SDK side. The openai_agents contrib dispatches two core events (no AI Guard import in contrib):

Event Wraps Covers
openai_agents.mcp.invoke_tool.before MCPUtil.invoke_mcp_tool Tools the agent runs after the model selects them; knows the model call ID and model-visible name
openai_agents.mcp.call_tool.before call_tool of every concrete SDK server class (0.23 overrides it without super()) Direct server.call_tool() calls

An agent-driven call is evaluated once, at the adapter; the lower check skips it through a one-shot marker, set only when the call may proceed (a blocked call cannot excuse a later direct call). A re-entrancy guard keeps an override calling super() to one dispatch.

Payload. function.name is the model-visible name, mcp.tool_name the original MCP tool name. mcp.name is the configured server name only: the SDK-generated names (stdio: <command>, sse: <url>) embed the command or raw URL and are dropped. mcp.url is sanitized; transport comes from the server class. Headers, auth, stdio command, arguments and environment are never sent.

Correlation. The OpenAI Chat and Responses after-listeners record the model's tool calls (per context, keyed by call ID). An agent-driven call reuses its real call ID and the conversation that produced it; sibling calls are evaluated when they run. Direct calls get a local dd_mcp_<uuid> ID and no fabricated history.

Blocking follows existing semantics (AIGuardAbortError, a BaseException, propagates out of Runner.run); monitor mode and evaluation errors preserve execution. New openai_agents value for the integration telemetry tag.

Testing

New ai_guard_openai_agents suite (Python 3.10-3.13; openai-agents 0.0.x with openai<1.100, and latest):

  • adapter evaluation with MCP identity, DENY/ABORT blocks with no tools/call, monitor and fail-open
  • direct calls (local ID), direct call after an agent call, blocked agent call does not excuse a direct call
  • conversation reuse by call ID, unknown call ID, same tool on two servers
  • stdio and generated-name servers leak no command, arguments, environment or credentials
  • listener gating
  • end-to-end Runner.run with a mocked model returning a function call and a stubbed MCP session (allow and block)

Locally: 16/16 on openai-agents 0.23.1 (py3.12, py3.10) and 0.0.19 (py3.12); ai_guard_openai (222) and contrib openai_agents (agents 0.14 and 0.0) still pass.

Risks

  • Off by default. When enabled, every openai-agents MCP tool call adds one synchronous AI Guard evaluation.
  • Tools converted before ddtrace patched openai-agents keep the unwrapped invoke_mcp_tool; their calls are still evaluated through the lower call_tool check, as direct calls.
  • Custom MCPServer subclasses outside the SDK are covered for agent-driven calls only.

Additional Notes

Merge order: merge this PR into the #20850 branch before #20850 merges, so both land in main together. If #20850 merges first, this PR is retargeted to main and must be merged on its own.

🤖 Generated with Claude Code

Evaluate the tool calls of MCP servers configured in the OpenAI Agents SDK
before the tools/call request is sent, behind
DD_AI_GUARD_COLLECT_MCP_ENABLED.

The openai_agents contrib dispatches two core events, with no AI Guard
import: openai_agents.mcp.invoke_tool.before around MCPUtil.invoke_mcp_tool
(agent-driven calls, which know the model call ID and model-visible name)
and openai_agents.mcp.call_tool.before around every concrete SDK server
call_tool (direct calls). An agent-driven call is evaluated once, at the
adapter.

Evaluations carry the configured server name (never the SDK-generated
name, which embeds the stdio command or raw URL), the sanitized URL, the
transport and the original tool name. OpenAI Chat and Responses listeners
record the model's tool calls so an agent-driven call reuses its real call
ID and conversation; direct calls get a local ID and no history.

APPSEC-70369

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against studio/aiguard-mcp-delivery-one using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

.cursor/rules/ai-guard.mdc                                              @DataDog/asm-python
.riot/requirements/127d746.txt                                          @DataDog/apm-python
.riot/requirements/189d47d.txt                                          @DataDog/apm-python
.riot/requirements/198d9ea.txt                                          @DataDog/apm-python
.riot/requirements/1bb3879.txt                                          @DataDog/apm-python
.riot/requirements/53d094a.txt                                          @DataDog/apm-python
.riot/requirements/63221c8.txt                                          @DataDog/apm-python
.riot/requirements/b61ec3c.txt                                          @DataDog/apm-python
.riot/requirements/bdbcb83.txt                                          @DataDog/apm-python
ddtrace/aiguard/_constants.py                                           @DataDog/asm-python
ddtrace/aiguard/_listener.py                                            @DataDog/asm-python
ddtrace/aiguard/integrations/_mcp.py                                    @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_agents.py                          @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_chat.py                            @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_responses.py                       @DataDog/asm-python
ddtrace/contrib/internal/openai_agents/patch.py                         @DataDog/ml-observability
docs/configuration.rst                                                  @DataDog/python-guild
releasenotes/notes/aiguard-openai-agents-mcp-tool-calls-88cc75f75c24e907.yaml  @DataDog/apm-python
tests/aiguard/openai/conftest.py                                        @DataDog/asm-python
tests/aiguard/openai_agents/conftest.py                                 @DataDog/asm-python
tests/aiguard/openai_agents/test_mcp.py                                 @DataDog/asm-python
tests/aiguard/suitespec.yml                                             @DataDog/asm-python

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Dependency direction analysis

📈 Existing violations got worse

1 pre-existing violation(s) increased in severity (e.g. their target became more depended-on, or got pulled into an import cycle), though the edge itself isn't new:

ddtrace.internal.settings.aiguard -×-> ddtrace.aiguard._constants  (internal-core -> product:aiguard, score=14, +1 vs base)

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.internal.opentelemetry.span -×-> ddtrace.trace  (product:opentelemetry -> product:tracing, score=130)
ddtrace.profiling.scheduler -×-> ddtrace.trace  (product:profiling -> product:tracing, score=130)
ddtrace.aiguard._api_client -×-> ddtrace.trace  (product:aiguard -> product:tracing, score=130)
ddtrace.llmobs._utils -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@datadog-datadog-us1-prod

datadog-datadog-us1-prod Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

✨ Unblock PR with BitsAI

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 3 Pipeline jobs failed

DataDog/apm-reliability/dd-trace-py | ast-grep rules — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

DataDog/apm-reliability/dd-trace-py | build_docs — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

DataDog/apm-reliability/dd-trace-py | validate_supported_configurations_v2_local_file — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: ddba254 | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Oct 7, 2026

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-10-07 09:55:42

Comparing candidate commit ddba254 in PR branch studio/aiguard-mcp-openai-agents with baseline commit 403deb3 in branch studio/aiguard-mcp-delivery-one.

📊 Benchmarking dashboard

Found 4 performance improvements and 8 performance regressions! Performance is the same for 595 metrics, 10 unstable metrics, 6 known flaky benchmarks, 18 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationextract-empty_headers

  • 🟩 execution_time [-64.726ns; -51.938ns] or [-8.761%; -7.030%]

scenario:httppropagationextract-full_t_id_datadog_headers

  • 🟥 execution_time [+1.714µs; +1.830µs] or [+11.691%; +12.483%]

scenario:httppropagationextract-none_propagation_style

  • 🟩 execution_time [-74.168ns; -59.436ns] or [-9.625%; -7.713%]

scenario:httppropagationextract-wsgi_empty_headers

  • 🟩 execution_time [-73.828ns; -60.506ns] or [-9.968%; -8.170%]

scenario:httppropagationextract-wsgi_invalid_span_id_header

  • 🟩 execution_time [-70.474ns; -59.432ns] or [-9.506%; -8.017%]

scenario:iastaspects-add_aspect

  • 🟥 execution_time [+7.278µs; +8.612µs] or [+8.578%; +10.152%]

scenario:iastaspects-format_map_noaspect

  • 🟥 execution_time [+65.105µs; +69.779µs] or [+17.914%; +19.200%]

scenario:iastaspectsremodule-re_expand_aspect

  • 🟥 execution_time [+99.824µs; +108.690µs] or [+15.538%; +16.918%]

scenario:msgpackencoderscenario-simple_one_span

  • 🟥 execution_time [+572.315ns; +630.781ns] or [+14.004%; +15.434%]

scenario:openfeatureflagevaluation-hook-enqueue-typical

  • 🟥 execution_time [+1.772µs; +1.940µs] or [+8.211%; +8.991%]

scenario:recursivecomputation-shallow

  • 🟥 execution_time [+58.136µs; +62.880µs] or [+8.379%; +9.062%]

scenario:span-start-finish-telemetry

  • 🟥 execution_time [+5.111ms; +5.300ms] or [+12.446%; +12.908%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-559.240ns; +908.260ns] or [-5.386%; +8.748%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-42.247ns; +36.260ns] or [-6.330%; +5.433%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1670.390ns; +2173.734ns] or [-8.399%; +10.930%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-1627.884ns; +2061.980ns] or [-8.519%; +10.791%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-322.701ns; +443.122ns] or [-7.579%; +10.408%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-185.753ns; +273.753ns] or [-6.969%; +10.271%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-125.524ns; +57.392ns] or [-9.118%; +4.169%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-4527.569ns; +4682.586ns] or [-9.482%; +9.807%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-912.897ns; +948.932ns] or [-9.006%; +9.362%]

scenario:packagesupdateimporteddependencies-import_many_stdlib_cached

  • unstable execution_time [-52.360µs; +55.881µs] or [-9.326%; +9.953%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+3.281µs; +3.443µs] or [+23.335%; +24.492%]

scenario:iastaspectsospath-ospathbasename_aspect

  • 🟥 execution_time [+143.219µs; +148.969µs] or [+37.336%; +38.835%]

scenario:iastaspectssplit-rsplit_aspect

  • 🟥 execution_time [+23.561µs; +27.666µs] or [+14.856%; +17.444%]

scenario:span-start

  • 🟥 execution_time [+1.139ms; +1.587ms] or [+9.088%; +12.660%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+168.717ns; +201.584ns] or [+8.494%; +10.149%]

scenario:tracer-small

  • 🟥 execution_time [+49.011µs; +50.353µs] or [+20.038%; +20.587%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:iastaspects-casefold_aspect
  • scenario:iastaspects-casefold_noaspect
  • scenario:iastaspects-index_aspect
  • scenario:iastaspects-ljust_noaspect
  • scenario:iastaspects-lower_aspect
  • scenario:iastaspects-replace_aspect
  • scenario:iastaspects-rstrip_aspect
  • scenario:iastaspects-swapcase_aspect
  • scenario:iastaspects-title_noaspect
  • scenario:iastaspects-translate_aspect
  • scenario:iastaspects-translate_noaspect
  • scenario:iastaspects-upper_noaspect
  • scenario:packagespackageforrootmodulemapping-cache_off
  • scenario:packagespackageforrootmodulemapping-cache_on
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant