Skip to content

fix(agents): cap container CPU limit at the Docker host's CPU count (#2318) - #3184

Open
obasilakis wants to merge 1 commit into
devfrom
fix/2318-cpu-clamp-to-host
Open

obasilakis wants to merge 1 commit into
devfrom
fix/2318-cpu-clamp-to-host

Conversation

@obasilakis

Copy link
Copy Markdown
Contributor

Closes #2318.

Problem

Docker refuses to create a container whose CPU limit exceeds the host's CPU count (range of CPUs is from 0.01 to 2.00, as there are only 2 CPUs available). The trinity-system template asks for 4 CPUs, so on any 2-CPU host the system agent is never created, and creation fails again on every backend start. A 2-vCPU t3a.large is the default instance type for the AWS Marketplace listing (#3004). A user agent set above the host's CPUs fails the same way.

Fix

services/docker_utils.containers_run — the single seam every agent create and recreate goes through (system agent, crud.py, lifecycle.py) — caps nano_cpus at docker info's NCPU and logs a warning when it does.

Deliberately not capped: the trinity.cpu label. helpers.py's resource drift check compares that label with the DB setting; a capped label would mismatch a DB value of 4 and recreate the agent on every start. The issue suggested capping the label too; this is why it doesn't.

If the host CPU count can't be read, the request passes through unchanged (Docker decides, as before).

Verification

  • tests/unit/test_2318_cpu_capped_to_host.py: 4→2 on 2 CPUs, unchanged when it fits, odd host count (8→3), no limit untouched, unreadable host count passes through. Failed before the fix on the two capping cases; passes after. Existing test_docker_utils.py passes.
  • Live: on a 2-vCPU EC2 instance (AMI from v1.0.0-aws.1) the backend logged the 400 above. With this file in place and the backend restarted: [#2318] CPU limit 4 exceeds the host's 2 CPUs; capping to 2, then System agent: created; container NanoCpus=2000000000, label trinity.cpu=4.
  • Full unit suite locally: the only failures are identical with and without this change (same 24 failed / 2 errors when those files are run on plain dev), from a local Python 3.11 environment; CI runs 3.13.

🤖 Generated with Claude Code

…2318)

Docker rejects NanoCpus above the host's CPU count with a 400. The
trinity-system template asks for 4 CPUs, so on any 2-CPU host (a t3a.large,
the AWS Marketplace default) the system agent was never created, failing
again on every backend start. A user agent configured above the host's CPUs
failed the same way.

containers_run, the one seam every agent create and recreate goes through,
now caps nano_cpus at docker info's NCPU and logs a warning. The trinity.cpu
label keeps the requested value: the resource drift check compares it with
the DB setting, and a capped label would recreate the agent on every start.
If the host count can't be read, the request passes through unchanged.

Verified on a 2-vCPU EC2 instance: with this file in place, trinity-system
was created and runs with NanoCpus=2e9, label trinity.cpu=4.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant