Skip to content

fix(telemetry): rebuild worker on identity refresh - #19821

Draft
litianningdatadog wants to merge 2 commits into
tianning.li/3-3-trace-writer-identity-refreshfrom
tianning.li/3-4-telemetry-identity-refresh
Draft

litianningdatadog wants to merge 2 commits into
tianning.li/3-3-trace-writer-identity-refreshfrom
tianning.li/3-4-telemetry-identity-refresh

Conversation

@litianningdatadog

@litianningdatadog litianningdatadog commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Stacked PRs:

Description

Telemetry has the same stale-identity problem as traces. The native telemetry worker is created with the runtime identity available at that time, so an explicit MicroVM identity refresh must replace the worker before later telemetry is sent.

TelemetryWriter registers with the explicit identity-refresh callback registry only when running in a MicroVM (in_aws_lambda_microvm()); elsewhere the callback is never installed and the writer behaves as before. On refresh, it discards the current worker without flushing its queued telemetry, clears the worker binding, and builds a fresh worker. If the previous worker had reported app-started, startup is reported again under the refreshed identity.

A worker rebuild starts with empty native state. The writer therefore replays accepted configuration events in sequence order and restores the latest integration and product-activation state on the replacement worker. Dependency reporting uses preserved tracker state and forces a full re-report so previously collected dependency metadata and SCA metadata are not lost.

Refresh can race with reporting calls that read or write self._worker (metrics, integrations, endpoints, configuration, logs, lifecycle, and fork handling). Each worker-accessing method now has a lock-free _without_lock implementation behind a public conditional-lock wrapper. MicroVM writers use the existing re-entrant lock; non-MicroVM writers use None and call the helper directly, avoiding nullcontext() overhead on normal hot paths.

Production status / native dependency

The successful end-to-end refresh path is gated on the native TelemetryWorker.drop() API. This branch currently pins libdatadog v43.0.1, whose TelemetryWorker exposes stop() but not drop().

stop() flushes queued telemetry and emits app-closing; its send_app_closing argument is currently ineffective. This PR therefore intentionally does not fall back to stop() for identity refresh. Against the current native pin, the MicroVM refresh callback fails explicitly rather than flushing the old runtime's queue or claiming a successful replacement. The runtime identity coordinator retries the callback, and the feature becomes live once the native discard API is available.

Reference

Testing

Added focused coverage for:

  • rebuilding the worker with refreshed runtime and session IDs
  • discarding the old worker without calling stop() or flushing its queue
  • propagating discard failures without mutating the old worker state
  • retrying after replacement-worker build failure and preserving app-started lifecycle state
  • restoring app-started state when needed
  • replaying configuration, integration, and product-activation state
  • registering the identity-refresh callback only for MicroVM writers
  • serializing metric recording against a concurrent identity refresh
  • re-reporting preserved dependency metadata after refresh(), with and without SCA enabled
  • no-op writer compatibility and explicit callback registration

Validation:

  • scripts/lint fmt: passed
  • scripts/lint style: passed
  • focused identity-refresh tests: 3 passed, 1 skipped because the pinned native worker has no drop() API
  • full telemetry suite: 192 passed, 2 skipped, 4 existing failures in tests/telemetry/test_telemetry.py; all changed tests/telemetry/test_writer.py tests passed

Risks

Low outside MicroVM environments: the identity-refresh callback is never registered there, and non-MicroVM worker access remains lock-free. Inside MicroVMs, worker replacement occurs only on explicit runtime identity refresh, and the re-entrant lock serializes refresh with worker access. With the current native pin, refresh fails explicitly and is retried rather than silently flushing stale telemetry. Native/libdatadog changes are not included in this PR.

Files (11)

  • ddtrace/internal/_runtime_id.py [MODIFIED] (+5 -0)
  • ddtrace/internal/runtime/init.py [MODIFIED] (+2 -0)
  • ddtrace/internal/telemetry/dependency.py [MODIFIED] (+4 -0)
  • ddtrace/internal/telemetry/dependency_tracker.py [MODIFIED] (+12 -6)
  • ddtrace/internal/telemetry/noop_writer.py [MODIFIED] (+7 -1)
  • ddtrace/internal/telemetry/writer.py [MODIFIED] (+361 -82)
  • ddtrace/internal/writer/writer.py [MODIFIED] (+19 -1)
  • releasenotes/notes/fix-telemetry-worker-identity-refresh-19fe1775be48b696.yaml [ADDED] (+5 -0)
  • tests/appsec/sca/test_telemetry.py [MODIFIED] (+1 -0)
  • tests/telemetry/test_dependency.py [MODIFIED] (+49 -1)
  • tests/telemetry/test_writer.py [MODIFIED] (+364 -0)

@litianningdatadog litianningdatadog added changelog/no-changelog A changelog entry is not required for this PR. aws-microvm Work related to AWS MicroVM onboarding labels Aug 23, 2026
@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 205 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 205 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.llmobs._integrations.anthropic -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.aiguard._api_client -×-> ddtrace.trace  (product:aiguard -> product:tracing, score=130)
ddtrace.llmobs._integrations.openai_agents -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.claude_agent_sdk -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against tianning.li/3-3-trace-writer-identity-refresh using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

ddtrace/internal/_runtime_id.py                                         @DataDog/apm-core-python
ddtrace/internal/runtime/__init__.py                                    @DataDog/apm-sdk-capabilities-python
ddtrace/internal/telemetry/dependency.py                                @DataDog/apm-python
ddtrace/internal/telemetry/dependency_tracker.py                        @DataDog/apm-python
ddtrace/internal/telemetry/noop_writer.py                               @DataDog/apm-python
ddtrace/internal/telemetry/writer.py                                    @DataDog/apm-python
ddtrace/internal/writer/writer.py                                       @DataDog/apm-core-python
releasenotes/notes/fix-telemetry-worker-identity-refresh-19fe1775be48b696.yaml  @DataDog/apm-python
tests/appsec/sca/test_telemetry.py                                      @DataDog/asm-python
tests/telemetry/test_dependency.py                                      @DataDog/apm-core-python @DataDog/apm-python
tests/telemetry/test_writer.py                                          @DataDog/apm-core-python @DataDog/apm-python

@litianningdatadog litianningdatadog changed the title fix(telemetry): rebuild worker on identity refresh chore(telemetry): rebuild worker on identity refresh Aug 23, 2026
@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Pipelines  Tests

✨ Unblock PR with BitsAI

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 3 Pipeline jobs failed

DataDog/apm-reliability/dd-trace-py | prechecks — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

System Tests | serverless-system-tests / Build end-to-end (function-url)

View more details · View in GitHub Actions

System Tests | system-tests finished

View more details · View in GitHub Actions

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: c44cf5a | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-09-27 19:56:22

Comparing candidate commit c44cf5a in PR branch tianning.li/3-4-telemetry-identity-refresh with baseline commit 3cda934 in branch tianning.li/3-3-trace-writer-identity-refresh.

📊 Benchmarking dashboard

Found 0 performance improvements and 5 performance regressions! Performance is the same for 366 metrics, 11 unstable metrics, 4 known flaky benchmarks, 4 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationextract-b3_single_headers

  • 🟥 execution_time [+1.282µs; +1.352µs] or [+19.088%; +20.143%]

scenario:httppropagationextract-empty_headers

  • 🟥 execution_time [+95.020ns; +112.908ns] or [+13.155%; +15.631%]

scenario:msgpackencoderscenario-simple_one_span

  • 🟥 execution_time [+515.825ns; +580.686ns] or [+12.524%; +14.098%]

scenario:otelspan-start

  • 🟥 execution_time [+1.966ms; +2.823ms] or [+8.060%; +11.577%]

scenario:recursivecomputation-shallow

  • 🟥 execution_time [+53.060µs; +56.420µs] or [+7.550%; +8.028%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-713.634ns; +747.847ns] or [-6.880%; +7.210%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-39.675ns; +39.034ns] or [-5.981%; +5.884%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1939.868ns; +1884.671ns] or [-9.762%; +9.485%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-2295.292ns; +1370.304ns] or [-11.875%; +7.089%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-391.513ns; +369.992ns] or [-9.203%; +8.698%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-235.618ns; +224.916ns] or [-8.820%; +8.420%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-94.235ns; +96.466ns] or [-6.936%; +7.100%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-4134.114ns; +5135.429ns] or [-8.649%; +10.743%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-856.281ns; +1037.781ns] or [-8.424%; +10.210%]

scenario:flasksimple-appsec-get

  • unstable execution_time [-325.409µs; +87.775µs] or [-10.218%; +2.756%]

scenario:flasksqli-appsec-enabled

  • unstable execution_time [-44.036µs; +163.599µs] or [-2.338%; +8.687%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+2.848µs; +2.966µs] or [+19.998%; +20.830%]

scenario:span-start

  • 🟥 execution_time [+1.384ms; +1.816ms] or [+10.954%; +14.375%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+342.046ns; +373.159ns] or [+17.852%; +19.476%]

scenario:tracer-small

  • 🟥 execution_time [+42.816µs; +44.322µs] or [+17.184%; +17.789%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch from d59e112 to 16a5332 Compare August 24, 2026 02:23
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from 2165d36 to a44a4b8 Compare August 24, 2026 02:24
@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch from 16a5332 to 8dd7e8e Compare August 24, 2026 02:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from a44a4b8 to 2b1bf99 Compare August 24, 2026 02:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch 5 times, most recently from cde3045 to a0e3c42 Compare August 24, 2026 23:56
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from 2b1bf99 to b4e8c78 Compare August 25, 2026 13:23
@litianningdatadog
litianningdatadog requested a lite review from Copilot August 25, 2026 13:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensure telemetry emitted after a runtime identity refresh uses the refreshed runtime ID by tearing down the existing native telemetry worker and allowing it to be rebuilt.

Changes:

  • Wire TelemetryWriter to runtime identity changes and rebuild (drop) its native worker on refresh.
  • Stop the live native worker during identity refresh to prevent continued heartbeats with stale identity.
  • Add tests covering worker teardown on identity refresh and wiring through runtime.refresh_identity().

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.

File Description
ddtrace/internal/telemetry/writer.py Subscribes to runtime-id changes and stops/drops the native telemetry worker on identity refresh.
tests/telemetry/test_writer.py Adds identity-refresh tests for worker stop/drop behavior and wiring through runtime.refresh_identity().

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/telemetry/test_writer.py Outdated
Comment thread ddtrace/internal/telemetry/writer.py Outdated
@litianningdatadog
litianningdatadog requested review from a team and ZStriker19 and removed request for a team September 7, 2026 20:24
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from 92a469d to 5a48d93 Compare September 7, 2026 20:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from c8c76f4 to 1e03afd Compare September 7, 2026 20:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch 2 times, most recently from 982f582 to 9c54e0e Compare September 21, 2026 21:21
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 5 times, most recently from 24f368e to fb190e7 Compare September 22, 2026 05:38
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from 7d904ad to f2664ed Compare September 22, 2026 20:21
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from fb190e7 to e96b84d Compare September 22, 2026 20:25
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch 2 times, most recently from 8376c12 to ba3f177 Compare September 23, 2026 19:18
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from e96b84d to b51f682 Compare September 23, 2026 19:27
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from ba3f177 to 805d8d1 Compare September 23, 2026 19:41
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 3 times, most recently from f7876cb to b84f21b Compare September 24, 2026 15:37
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from 805d8d1 to 8dbfa50 Compare September 24, 2026 17:40
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 2 times, most recently from 156be45 to c8b45b6 Compare September 24, 2026 17:57
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from 8dbfa50 to 19be48b Compare September 24, 2026 18:25
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from c8b45b6 to c641a12 Compare September 24, 2026 18:29
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-3-trace-writer-identity-refresh branch from 19be48b to 7f7d597 Compare September 24, 2026 19:13
litianningdatadog and others added 2 commits September 26, 2026 23:40
Rebuild the native telemetry worker when a MicroVM refreshes its runtime
identity, registering the refresh callback only for MicroVM writers.

The rebuild races with reporting-path calls (metrics, integrations,
endpoints, configuration) that could otherwise observe or write to a
stale worker mid-swap. Serialize worker replacement against those paths
with a dedicated lock, and replay previously accepted integration and
product-activation state onto the new worker so a rebuild does not
silently drop it.

Dependency reporting also needs a full re-report against the new worker:
mark the dependency tracker for refresh so already-sent dependencies
(and their SCA metadata) are resent, even with SCA disabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

aws-microvm Work related to AWS MicroVM onboarding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants