fix(telemetry): rebuild worker on identity refresh - #19821
litianningdatadog wants to merge 2 commits into
Conversation
Circular import analysis
|
Dependency direction analysis
|
Codeowners resolved asResolved from the full PR diff against |
❌ ErrorsYour PR has failed checks. Please review the issues below and take necessary action before merging. 🚦 3 Pipeline jobs failed
ℹ️ InfoNo other issues found (see more)🧪 All tests passed Useful? React with 👍 / 👎 This comment will be updated automatically if new data arrives.🔗 Commit SHA: c44cf5a | Docs | View more details | Give us feedback! |
BenchmarksBenchmark execution time: 2026-09-27 19:56:22 Comparing candidate commit c44cf5a in PR branch Found 0 performance improvements and 5 performance regressions! Performance is the same for 366 metrics, 11 unstable metrics, 4 known flaky benchmarks, 4 flaky benchmarks without significant changes.
|
d59e112 to
16a5332
Compare
2165d36 to
a44a4b8
Compare
16a5332 to
8dd7e8e
Compare
a44a4b8 to
2b1bf99
Compare
cde3045 to
a0e3c42
Compare
2b1bf99 to
b4e8c78
Compare
There was a problem hiding this comment.
Pull request overview
Ensure telemetry emitted after a runtime identity refresh uses the refreshed runtime ID by tearing down the existing native telemetry worker and allowing it to be rebuilt.
Changes:
- Wire
TelemetryWriterto runtime identity changes and rebuild (drop) its native worker on refresh. - Stop the live native worker during identity refresh to prevent continued heartbeats with stale identity.
- Add tests covering worker teardown on identity refresh and wiring through
runtime.refresh_identity().
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
ddtrace/internal/telemetry/writer.py |
Subscribes to runtime-id changes and stops/drops the native telemetry worker on identity refresh. |
tests/telemetry/test_writer.py |
Adds identity-refresh tests for worker stop/drop behavior and wiring through runtime.refresh_identity(). |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
92a469d to
5a48d93
Compare
c8c76f4 to
1e03afd
Compare
982f582 to
9c54e0e
Compare
24f368e to
fb190e7
Compare
7d904ad to
f2664ed
Compare
fb190e7 to
e96b84d
Compare
8376c12 to
ba3f177
Compare
e96b84d to
b51f682
Compare
ba3f177 to
805d8d1
Compare
f7876cb to
b84f21b
Compare
805d8d1 to
8dbfa50
Compare
156be45 to
c8b45b6
Compare
8dbfa50 to
19be48b
Compare
c8b45b6 to
c641a12
Compare
19be48b to
7f7d597
Compare
Rebuild the native telemetry worker when a MicroVM refreshes its runtime identity, registering the refresh callback only for MicroVM writers. The rebuild races with reporting-path calls (metrics, integrations, endpoints, configuration) that could otherwise observe or write to a stale worker mid-swap. Serialize worker replacement against those paths with a dedicated lock, and replay previously accepted integration and product-activation state onto the new worker so a rebuild does not silently drop it. Dependency reporting also needs a full re-report against the new worker: mark the dependency tracker for refresh so already-sent dependencies (and their SCA metadata) are resent, even with SCA disabled. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Stacked PRs:
Description
Telemetry has the same stale-identity problem as traces. The native telemetry worker is created with the runtime identity available at that time, so an explicit MicroVM identity refresh must replace the worker before later telemetry is sent.
TelemetryWriterregisters with the explicit identity-refresh callback registry only when running in a MicroVM (in_aws_lambda_microvm()); elsewhere the callback is never installed and the writer behaves as before. On refresh, it discards the current worker without flushing its queued telemetry, clears the worker binding, and builds a fresh worker. If the previous worker had reported app-started, startup is reported again under the refreshed identity.A worker rebuild starts with empty native state. The writer therefore replays accepted configuration events in sequence order and restores the latest integration and product-activation state on the replacement worker. Dependency reporting uses preserved tracker state and forces a full re-report so previously collected dependency metadata and SCA metadata are not lost.
Refresh can race with reporting calls that read or write
self._worker(metrics, integrations, endpoints, configuration, logs, lifecycle, and fork handling). Each worker-accessing method now has a lock-free_without_lockimplementation behind a public conditional-lock wrapper. MicroVM writers use the existing re-entrant lock; non-MicroVM writers useNoneand call the helper directly, avoidingnullcontext()overhead on normal hot paths.Production status / native dependency
The successful end-to-end refresh path is gated on the native
TelemetryWorker.drop()API. This branch currently pins libdatadogv43.0.1, whoseTelemetryWorkerexposesstop()but notdrop().stop()flushes queued telemetry and emitsapp-closing; itssend_app_closingargument is currently ineffective. This PR therefore intentionally does not fall back tostop()for identity refresh. Against the current native pin, the MicroVM refresh callback fails explicitly rather than flushing the old runtime's queue or claiming a successful replacement. The runtime identity coordinator retries the callback, and the feature becomes live once the native discard API is available.Reference
Testing
Added focused coverage for:
stop()or flushing its queuerefresh(), with and without SCA enabledValidation:
scripts/lint fmt: passedscripts/lint style: passeddrop()APItests/telemetry/test_telemetry.py; all changedtests/telemetry/test_writer.pytests passedRisks
Low outside MicroVM environments: the identity-refresh callback is never registered there, and non-MicroVM worker access remains lock-free. Inside MicroVMs, worker replacement occurs only on explicit runtime identity refresh, and the re-entrant lock serializes refresh with worker access. With the current native pin, refresh fails explicitly and is retried rather than silently flushing stale telemetry. Native/libdatadog changes are not included in this PR.
Files (11)