Skip to content

Run /search in the threadpool instead of blocking the event loop - #9

Merged
sls701 merged 1 commit into
mainfrom
fix/unblock-event-loop
Jun 5, 2026
Merged

sls701 merged 1 commit into
mainfrom
fix/unblock-event-loop

Conversation

@Vilin97

@Vilin97 Vilin97 commented Jun 5, 2026

Copy link
Copy Markdown
Collaborator

The core fix for the 502/503/504/429 errors

/search and the /mcp tool handler were async def but did blocking I/O (psycopg2 queries + a ~1s OpenAI/Nebius embedding call). Blocking calls inside an async route run on the event loop, so each container instance processed these requests one at a time.

The API runs on AWS App Runner, which routes up to MaxConcurrency (default 100) concurrent requests to a single instance. With the handlers serialized, that backlog produced exactly the reported errors — all emitted by App Runner/Envoy, never by the app (whose only failure path is HTTP 500):

  • 504 — requests waiting behind the serialized queue exceed App Runner's request timeout
  • 502 — the jammed event loop fails health checks → instance recycled mid-request
  • 503 — no healthy instance during the recycle window
  • 429 — App Runner load-shedding during scale-out

Change

  • search is now a sync def, so FastAPI runs it in its worker threadpool (concurrent requests run in parallel) — matching /graph/embedding, which was already a sync def and does not exhibit the bug.
  • The async /mcp handler offloads via run_in_threadpool(search, …).
  • Guarded the lazy get_openai_client() global init with a threading.Lock (double-checked locking), since it can now be called from several threadpool threads at once on a cold instance.

Evidence (measured against prod)

16 concurrent /search calls: ~17s wall, p50 12s (serialized). The identical-workload sync /graph/embedding stays ~1s flat under the same concurrency.

Review

Reviewed by a multi-agent adversarial pass (FastAPI async semantics, call-site regressions, concurrency) plus a focused second pass on the lock — no outstanding issues.

Companion infra guidance: docs/apprunner-scaling-runbook. Independent of the other PRs; recommend merging this one first.

/search and the /mcp tool handler were `async def` but did blocking I/O
(psycopg2 queries + a ~1s OpenAI embedding call). Blocking calls in an
async route run on the event loop, so each instance processed these
requests one at a time. Under AWS App Runner's per-instance concurrency
(default MaxConcurrency=100) this serialization is what produced the
502/503/504 (health-check failures, recycling, request-timeout) and the
rare 429 (scale-out load-shedding).

Make `search` a plain sync `def` so FastAPI runs it in its worker
threadpool (concurrent requests run in parallel), and have the async
/mcp handler offload to the threadpool via run_in_threadpool.

Measured: 16 concurrent /search calls went from ~17s wall (p50 12s,
serialized) toward threadpool-parallel behavior matching /graph/embedding
(already a sync def), which stays ~1s flat under the same load.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses App Runner 502/503/504/429s caused by blocking I/O inside async request handlers by ensuring /search work runs in FastAPI’s worker threadpool instead of the event loop.

Changes:

  • Converted /search from async def to sync def so blocking psycopg2 + embedding calls execute in FastAPI’s threadpool.
  • Updated the async /mcp handler to offload the search call via run_in_threadpool(...).
  • Added a lock around lazy OpenAI client initialization to avoid concurrent cold-start races.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.

File Description
api/routes/search.py Makes /search synchronous (threadpool execution) and adds thread-safe OpenAI client lazy-init.
api/routes/mcp.py Offloads search execution to the threadpool to keep the MCP handler non-blocking.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread api/routes/search.py
@@ -10,14 +11,21 @@
router = APIRouter()

_openai_client: OpenAI = None
Comment thread api/routes/search.py
Comment on lines 234 to 236
try:
with rds_conn() as conn, conn.cursor() as cur:
cur.execute(
@sls701
sls701 merged commit b6132f8 into main Jun 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants