Summary
With workers-runtime-sdk==1.9.2, a Flask stream_with_context response fails when the WSGI adapter continues the body iterator from its asynchronous ReadableStream.pull callback. A normal Flask POST succeeds in the same setup.
This reproduces in standalone, unmodified upstream workerd v1.20260929.1, with the unmodified SDK. The reproduction uses a direct service binding to a Python Worker; it does not involve open-compute's daemon, Worker Loader, or deployment pipeline. I have not tested the hosted Cloudflare edge.
Reproduction
The relevant endpoint and entrypoint from the tested Flask application are:
from workers import WorkerEntrypoint, wsgi
from flask import Flask, request, Response, stream_with_context
app = Flask(__name__)
closed = 0
@app.get("/stream")
def stream():
@stream_with_context
def body():
global closed
try:
yield "flask:"
yield request.args["value"]
finally:
closed += 1
return Response(body(), content_type="text/plain")
class Default(WorkerEntrypoint):
async def fetch(self, request):
return await wsgi.fetch(app, request, self.env, self.ctx)
Load the application with the SDK and Flask dependencies as python_modules/* modules in workerd. The verified setup uses:
- macOS arm64; upstream release
v1.20260929.1 (workerd 2026-09-29).
workers-runtime-sdk==1.9.2, Flask 3.1.3, Werkzeug 3.1.9, Jinja2 3.1.6, MarkupSafe 3.0.4, Click 8.5.0, Blinker 1.9.0, ItsDangerous 2.2.0.
- Python Worker compatibility date
2026-09-08; flags python_workers, enable_python_external_sdk, python_dedicated_snapshot.
- A JavaScript test Worker with a
PY service binding to the Python Worker.
The tested client operation is:
const response = await env.PY.fetch(
"https://worker.test/stream?value=stream-body"
);
assert.equal(response.status, 200);
const body = await response.text();
// Expected: "flask:stream-body". Actual: body consumption rejects.
For the diagnostic run, the client catches that rejection and asserts that it contains ContextVar and flask.app_ctx or flask.request_ctx. It also checks a successful Flask POST in the same run. Thus the diagnostic harness passing means it confirmed the failure, not that the stream succeeded.
Actual behavior
workerd logs Exception while streaming WSGI response body, followed by this chain:
RuntimeError: Working outside of request context.
LookupError: <ContextVar name='flask.request_ctx' ...>
LookupError: <ContextVar name='flask.app_ctx' ...>
The stack enters workers/wsgi.py:269 (pull, next(chunks, _END)), then body_chunks, Werkzeug's response iterator, and Flask's helpers.py:130 (with app_ctx, req_ctx). The response is created with status 200, but consuming its body fails.
Analysis and diagnostic fix
In process_request, the adapter invokes the application and primes the first non-empty body chunk in the initial Python context. Subsequent iteration and cleanup happen inside asynchronous FFI pull/cancel callbacks. Those callbacks do not retain the same Python Context as the initial iteration. Flask's context tokens cannot be used or reset across these contexts; merely copying the values into a different context is insufficient.
The installed wsgi.py is byte-identical to the upstream file at 95ca919b2fa6b3aa5c57c653136e0d01ba85f51b. A fresh read of the current main file has the same SHA-256: 64738dbf226a764a1085befaefa8f69c8aee47a1a7ac22ac33458315eb33e97f.
I tested a diagnostic SDK-only change that creates one owned context per response and uses that same object for application invocation, all iterator advancement, and iterator closing:
context = copy_context()
result = context.run(app, environ, start_response)
raw_result_iter = iter(result)
def contextual_chunks():
while (chunk := context.run(next, raw_result_iter, _END)) is not _END:
yield chunk
result_iter = contextual_chunks()
# In close_all():
context.run(_close_iterable, result)
That candidate changes only wsgi.py. On our pinned local workerd fork it passed two concurrent streams with distinct request values, explicit stream cancellation with exactly-once generator cleanup, and the ordinary POST control. No native runtime or daemon changes were needed. The patched candidate has not yet been tested against stock workerd or the hosted edge, so this is a diagnostic fix rather than a fully qualified upstream patch.
Related reports
Could the WSGI adapter preserve one response-owned Python context across application invocation, iteration, and closing, with a Flask stream_with_context regression test covering completion, concurrent responses, and cancellation?
Summary
With
workers-runtime-sdk==1.9.2, a Flaskstream_with_contextresponse fails when the WSGI adapter continues the body iterator from its asynchronousReadableStream.pullcallback. A normal Flask POST succeeds in the same setup.This reproduces in standalone, unmodified upstream workerd
v1.20260929.1, with the unmodified SDK. The reproduction uses a direct service binding to a Python Worker; it does not involve open-compute's daemon, Worker Loader, or deployment pipeline. I have not tested the hosted Cloudflare edge.Reproduction
The relevant endpoint and entrypoint from the tested Flask application are:
Load the application with the SDK and Flask dependencies as
python_modules/*modules in workerd. The verified setup uses:v1.20260929.1(workerd 2026-09-29).workers-runtime-sdk==1.9.2, Flask3.1.3, Werkzeug3.1.9, Jinja23.1.6, MarkupSafe3.0.4, Click8.5.0, Blinker1.9.0, ItsDangerous2.2.0.2026-09-08; flagspython_workers,enable_python_external_sdk,python_dedicated_snapshot.PYservice binding to the Python Worker.The tested client operation is:
For the diagnostic run, the client catches that rejection and asserts that it contains
ContextVarandflask.app_ctxorflask.request_ctx. It also checks a successful Flask POST in the same run. Thus the diagnostic harness passing means it confirmed the failure, not that the stream succeeded.Actual behavior
workerd logs
Exception while streaming WSGI response body, followed by this chain:The stack enters
workers/wsgi.py:269(pull,next(chunks, _END)), thenbody_chunks, Werkzeug's response iterator, and Flask'shelpers.py:130(with app_ctx, req_ctx). The response is created with status 200, but consuming its body fails.Analysis and diagnostic fix
In
process_request, the adapter invokes the application and primes the first non-empty body chunk in the initial Python context. Subsequent iteration and cleanup happen inside asynchronous FFIpull/cancelcallbacks. Those callbacks do not retain the same PythonContextas the initial iteration. Flask's context tokens cannot be used or reset across these contexts; merely copying the values into a different context is insufficient.The installed
wsgi.pyis byte-identical to the upstream file at 95ca919b2fa6b3aa5c57c653136e0d01ba85f51b. A fresh read of the currentmainfile has the same SHA-256:64738dbf226a764a1085befaefa8f69c8aee47a1a7ac22ac33458315eb33e97f.I tested a diagnostic SDK-only change that creates one owned context per response and uses that same object for application invocation, all iterator advancement, and iterator closing:
That candidate changes only
wsgi.py. On our pinned local workerd fork it passed two concurrent streams with distinct request values, explicit stream cancellation with exactly-once generator cleanup, and the ordinary POST control. No native runtime or daemon changes were needed. The patched candidate has not yet been tested against stock workerd or the hosted edge, so this is a diagnostic fix rather than a fully qualified upstream patch.Related reports
pull/cancelhandlers asynchronous. Its current implementation is present in the SDK tested here; the regression tests in that PR use plain WSGI generators rather than Flaskstream_with_context.Could the WSGI adapter preserve one response-owned Python context across application invocation, iteration, and closing, with a Flask
stream_with_contextregression test covering completion, concurrent responses, and cancellation?