English | Español
Autonomous Multi-Agent Software Development System governed by Finite State Machines (FSM), Deterministic Verification Gates, Execution Budgets, Git Worktree Isolation, and Cryptographic Evidence Chaining.
Warning
Experimental Research Project & Security Audit Status Notice:
- Published Baseline Commit (
c5fd13c): Represents the initial published baseline containing 202 unit/integration tests, heuristic verification gates, and fast-forward worktree integration. - Current Local Development State (not yet published): Incorporates extensive security hardening from Round 3.6 (immutable verification context, controller-side hidden challenges, HMAC evidence chaining, and atomic Git Compare-and-Swap ref updates) and Round 3.7 Stage 1 test-and-freeze (50 frozen acceptance cases across six test files and two supporting Python files).
- Security Remediation Pending: An independent security audit concluded that Round 3.6 does not receive a security PASS, confirming five security defects (S1–S5) and three functional regressions (F1–F3). Stage 1 established 31 genuine FAIL-before acceptance cases and 19 passing controls. Production code remains unchanged from Round 3.6; production security remediation remains pending.
Most contemporary multi-agent coding systems encounter critical failure modes when operating on real-world software projects:
- Unchecked Hallucinations: Over-reliance on LLM self-evaluations and optimistic "looks good to me" feedback loops.
- Scope Creep & Code Boundary Violations: Agents that arbitrarily modify configuration files, delete existing tests to force pass rates, or edit modules outside their assigned scope.
- Infinite Trial-and-Error Loops: Unbounded token consumption as agents cycle through repetitive debugging attempts without formal termination budgets.
- Working Tree Corruption: Direct modifications made to the active development branch that leave the repository broken when a build or test fails mid-flight.
- Unsafe Git Integrations: Non-atomic merges or silent merge conflicts integrated directly into branches without clean-state validation.
Autonomous Multi-Agent Software Factory decouples generation from governance. LLMs are strictly restricted to proposing code, specifications, and diagnostics, while workflow control, test execution, security analysis, file boundaries, and Git integration are enforced by non-probabilistic deterministic gates, a transactional Finite State Machine (FSM), and cryptographic verification evidence.
flowchart TD
Task["Task: TASK-XXX"] --> Architect["Architect: Gemini 3.8 Flash"]
Architect --> SPEC_GATE{"SPEC_GATE: Contract Validation"}
SPEC_GATE --> Worktree["Isolation: Git Worktree"]
Worktree --> Worker["Worker: DeepSeek Flash"]
Worker --> DIFF_GATE{"DIFF_GATE: Boundary Check"}
DIFF_GATE --> TESTING{"TESTING: Test Verification"}
TESTING --> SAST_SCAN{"SAST_SCAN: Semgrep & Bandit"}
SAST_SCAN --> LOGIC_AUDIT{"LOGIC_AUDIT: LLM Security Audit"}
LOGIC_AUDIT --> EvidenceVerification{"Evidence Verification"}
EvidenceVerification --> CAS{"Git CAS: Ref B to C"}
CAS --> COMPLETED(["COMPLETED: Task Finished"])
Note
Evolution from Published Baseline to Hardened Architecture:
- Essential Pipeline Flow: The diagram illustrates the linear verification sequence. Error paths (diff reversion on boundary violations, triage diagnosis loops on test failures, security replanning on confirmed SAST/audit findings, and circuit-breaker halts) are governed by strict execution budgets in prose and configuration.
- Published Baseline vs. Local Development: In the published baseline commit (
c5fd13c),TESTINGexecuted client-side pytest with coverage reports, and integration was performed bygit merge --ff-onlyon the worktree. In the newer local development state (not yet published),TESTINGevaluates controller-side hidden challenges (restricted to thepassword-v36contract), and integration is executed via an atomic Git ref Compare-and-Swap (git update-refCAS B → C). - Human Review & Recovery: Tasks suspended at
LOGIC_AUDITdue to transient upstream availability failures (HTTP 429/503) enterHUMAN_REVIEWwithout budget consumption, recoverable via--resume-audit. Tasks halted atAUTO_MERGEcan be recovered via--resume-merge. Recovery tokens are intended to be strictly one-use (subject to confirmed defect S5). - Remaining Security Limitations:
- S2 Unresolved: Signed evidence confirms record integrity and sequence, but does not prove concrete gate execution (
_record_successaccepts caller-supplied gate names). - S1 Unresolved: The local freshness witness does not prevent paired state and witness rollback or deletion when directory permissions permit modification.
- F1 Worktree Gap: Ref-only CAS advances the branch tip but does not automatically synchronize the main worktree.
- S2 Unresolved: Signed evidence confirms record integrity and sequence, but does not prove concrete gate execution (
| Role | Default Model | Provider | Responsibility |
|---|---|---|---|
| Architect | gemini-3.8-flash |
Google Gemini | Analyzes requirements and drafts formal tabular specifications (specs/TASK-XXX.md). |
| Worker | deepseek-flash |
DeepSeek | Implements source code and unit tests inside an isolated worktree under RULES.md. |
| Triage | gemini-3.5-flash-lite |
Google Gemini | Analyzes test failures (stdout, stderr, stack traces) and outputs structured diagnosis. |
| Security Filter | gemini-3.5-flash-lite |
Google Gemini | Reviews raw SAST alerts to differentiate true positives from false alarms. |
| Logic Security Auditor | gemini-3.8-flash (Primary)Fallbacks: deepseek-flash, qwen3.8-flash, gemini-3.6-flash (plus glm-5.3 support) |
Google Gemini, DeepSeek, Qwen / DashScope, Zhipu GLM | Verifies that semantic code diffs satisfy all security invariants ([SEC-xx]) under a deterministic, fail-closed multi-provider fallback policy. |
SPEC_GATE(scripts/spec_gate.py): Validates required sections, unambiguous Acceptance Criteria ([AC-xx]), Security Invariants ([SEC-xx]), explicit file boundary whitelists (Allowed files/Forbidden files), and 1:1 traceability in the Test Matrix before code generation.DIFF_GATE(scripts/diff_gate.py): Inspectsgit diff B..Cagainst the frozen specification. Strictly enforces set-theoretic file boundaries:modified_files ⊆ allowed_filesforbidden_files ∩ modified_files = ∅Blocks modifications to root infrastructure files (pyproject.toml,.env*,RULES.md,orchestrator/,scripts/).
TESTING/ Behavioral Verification (scripts/test_runner.py; local research suite:trusted_tests/suite.py):- Published Baseline: Ran
pytestwith ≥ 85% line coverage on task files. - Hardened Architecture: Candidate-reported pytest execution and coverage claims are untrusted. The controller generates 149 hidden behavioral challenges across boundary conditions, Unicode, and invalid types. The controller oracle evaluates candidate outputs against in-memory expected booleans.
- Current Scope: Restricted to the
password-v36contract (require_pytest=false). Pytest completion and general test execution remain general-purpose factory work.
- Published Baseline: Ran
SAST_SCAN(scripts/sast_runner.py; local research rules:scripts/semgrep_rules.yml): Executes static application security testing using local versioned Semgrep rules and Bandit. Scanner invocation is isolated to controller modules; tool errors (exit code 2) fail closed and halt the pipeline before audit or merge.LOGIC_AUDIT(orchestrator.py,orchestrator/logic_audit.pyin local research state,adapters/): An independent LLM auditor verifies that the immutable Git diffB..Cstrictly satisfies all stated security invariants ([SEC-xx]). Protected by a deterministic multi-provider fallback chain (gemini-3.8-flash→deepseek-flash→qwen3.8-flash→gemini-3.6-flash) activating strictly on provider availability failures (HTTP 429/503), strictly rejecting model shopping on semanticFAIL. Simulation mode cannot issue success evidence.MERGE_GATE(scripts/merge_gate.py): Executes an atomic Git ref Compare-and-Swap (CAS) (B → C) usinggit update-ref. Requires that the repository base reference still matchesBand that four valid, authentic HMAC-signed evidence records exist in exact sequence. (See Functional Regression F1 regarding worktree synchronization).
Task lifecycle is managed by orchestrator/state_manager.py, which persists state into orchestrator/state/state_<TASK_ID>.json:
INIT -> SPEC_DESIGN -> SPEC_GATE -> BUILDING -> DIFF_GATE -> TESTING ->
TRIAGING -> SAST_SCAN -> SAST_FILTER -> LOGIC_AUDIT -> AUTO_MERGE
Terminal / suspension states: COMPLETED, HALT_HUMAN, HUMAN_REVIEW.
StateLock: Implements intra-process reentrancy (threading.RLock) combined with inter-process kernel-level file locking (msvcrton Windows,fcntl.flockon POSIX).- Atomic Replacement: State modifications occur exclusively via
_transaction, which reloads authoritative state, validates terminal transitions, increments generation counters, flushes and fsyncs temporary files, and atomically replaces the state file. - Freshness Witness: Schema 36 and a persistent lockfile SHA-256 freshness witness detect single-file snapshot replays when the witness is retained. However, as demonstrated by confirmed defect S1, the witness does not prevent rollback or deletion if both state and witness files are restored or removed together. Direct
save(),_save_unlocked(), and_write_snapshot()APIs have been removed.
- Opaque Capabilities: Recovery APIs (
validate_strict_recovery,can_resume_merge,can_resume_audit,authorize_recovery,execute_merge_recovery_transition,execute_audit_recovery_transition) issue opaque, per-instance, generation-bound objects only after re-verifying current semantic, context, and evidence checks. - Intended One-Use Semantics: Capabilities are intended to be strictly one-use. However, independent security audit confirmed defect S5: if strict recovery validation fails before the mutation transaction, the capability is not consumed, allowing reuse if invalid conditions are reverted.
Configured in orchestrator/config.json:
max_worker_per_epoch: 2 local attempts by Worker per epoch.max_cumulative_worker_runs: 5 total lifetime Worker runs per task.max_logic_replans: 2 replanning attempts on persistent test failures.max_security_replans: 1 replanning attempt on confirmed vulnerabilities.max_spec_syntax_retries: 1 retry for specification syntax formatting.
To ensure maximum availability against upstream API rate limits (HTTP 429) and transport outages (HTTP 502/503/504) without sacrificing security guarantees, LOGIC_AUDIT enforces a deterministic, fail-closed multi-provider fallback policy configured in orchestrator/config.json:
- Primary Model:
gemini(gemini-3.8-flash) - Fallback Candidate 1:
deepseek(deepseek-flash) - Fallback Candidate 2:
qwen(qwen3.8-flash) via DashScope OpenAI-compatible API - Fallback Candidate 3:
gemini(gemini-3.6-flash) (Note: Full adapter support is also implemented for Zhipu GLM viaGLMAdapter).
Strict Fallback Semantics:
- Availability Failures Only: Fallback triggers exclusively on transport-level network errors and provider unavailability (HTTP 429 Too Many Requests, HTTP 502/503/504, connection timeouts, socket aborts) after per-provider retries with exponential backoff (
@retry_with_backoff) are exhausted. - Semantic
FAILNever Triggers Fallback (No Model Shopping): If any candidate model successfully connects and evaluates that a security invariant has been violated (FAIL), the audit verdict is final. No further fallback models are queried. The failure immediately consumes a security replan budget and routes back to Triage and Worker. - Stop at First Verdict: The pipeline halts at the first model that successfully returns
PASS. - Safe Halt on Ambiguity or Exhaustion: If a model returns an ambiguous verdict (
UNCERTAIN) or malformed JSON, execution transitions safely toHUMAN_REVIEWwithout trying subsequent models. If all candidates in the fallback chain fail due to network availability, execution pauses safely inHUMAN_REVIEW. - Zero Budget Consumption on Fallbacks: Switching between fallback models due to transport/availability failures does not consume Worker attempts, logic replans, or security replans. The epoch counter is untouched.
- Structured State Ledger: The model and provider that successfully performed the audit are deterministically recorded under
audit_model_used: {"provider": "<provider>", "model": "<model>"}instate_<TASK_ID>.jsonwithout exposing credentials.
CRASH_REPORT_<TASK_ID>.md: Triggered when any budget is exhausted or an unrecoverable failure occurs. Freezes worktree artifacts for inspection.HUMAN_REVIEW: Triggered when external LLM endpoints return persistent network or rate limit errors (HTTP 429 or 503). Suspends execution in a safe state without spending Worker attempts or replan budgets.--resume-merge: Safe recovery pathway for tasks that cleared every verification gate but halted atAUTO_MERGE(e.g., due to local uncommitted edits ondev). Performs recovery via generation-bound capability without invoking agents or altering budget counters.--resume-audit: Safe recovery pathway for tasks suspended inHUMAN_REVIEWwithblocked_reason.gate == LOGIC_AUDIT(e.g., following upstream HTTP 429 rate limits). Resumes exclusively the pending Logic Security LLM audit without re-running Worker, pytest, SAST, or consuming Worker attempts/replans.
Rather than performing file operations in the main working tree:
- Each task generates an isolated Git worktree at
.worktrees/wt_<TASK_ID>linked to branchtask/<TASK_ID>. - If a Worker introduces broken code, unformatted files, or gate violations, the main branch remains clean and untouched.
- Candidate Extraction: The candidate commit C is extracted once. The immutable manifest derives the target tree using
<C>^{tree}and reads explicit Git objects. - Platform Invariant: Native Windows materialization rejects path traversals and fails closed before creating destinations. (Live POSIX descriptor-relative materialization and Docker containment remain unverified on the Windows development host).
The security model enforces five non-probabilistic invariants:
+-----------------------------------------------------------------------------------+
| IMMUTABLE VERIFICATION CONTEXT |
| B (Base Ref) | C (Candidate Commit) | tree (C^{tree}) | M (Manifest Digest) |
| spec_digest | config_digest | policy_digest |
+-----------------------------------------------------------------------------------+
│
┌─────────────────────┼─────────────────────┐
▼ ▼ ▼
[DIFF_GATE] [TESTING] [SAST_SCAN]
B..C Diff Check 149 Hidden Challenges Local Semgrep + Bandit
│ │ │
└─────────────────────┼─────────────────────┘
▼
[LOGIC_AUDIT]
LLM Diff Security Audit
│
▼
[CRYPTOGRAPHIC EVIDENCE]
HMAC-SHA256 Chained Gate Records
│
▼
[MERGE_GATE]
Atomic Git CAS: Ref B -> C
The candidate commit C is selected once from the task branch. The controller constructs an immutable VerificationContext containing:
Context = (B, C, tree, M, spec_digest, config_digest, policy_digest)
The context is written transactionally before DIFF_GATE. Each subsequent gate re-verifies that current repository specification, configuration, policy bytes, and implementation identities match this frozen context before execution.
To prevent candidate tampering, sandbox escape, or forged test reports:
- The controller generates 149 hidden challenges dynamically in memory.
- Expected boolean outcomes remain exclusively in controller memory; only challenge IDs and input payloads enter the scratch directory.
- Exact boolean evaluations are conducted on the controller side.
- Candidate claims of
PASS, pytest exit codes, or coverage percentages are strictly rejected as authorization channels.
VerificationSessionrecords gate execution outcomes signed with HMAC-SHA256 (key stored under.controller-secrets/evidence.key).- Each record binds:
schema,task_id,gate,context_digest,result,implementation_identities,attempt,sequence,generation,timestamp, and theprevious_record_hash. - Cryptographic Integrity vs. Semantic Provenance: The HMAC signatures guarantee tamper-evident storage and ordered chaining. However, independent audit confirmed defect S2:
_record_success(gate)accepted caller-provided gate names without unforgeable execution proof from concrete gate implementations, demonstrating that cryptographic record integrity does not inherently establish semantic provenance.
- Code integration does not execute generic working-tree merges.
merge_gate.pyusesgit update-refto perform an atomic compare-and-swap (B → C).- CAS succeeds only if:
- The target branch reference still points exactly to base commit B.
- All four mandatory gate evidence records (
DIFF_GATE,TESTING,SAST,LOGIC_AUDIT) are present, authentic, hash-chained, and valid.
- A concurrent base ref update causes CAS to reject, leaving B untouched.
A rigorous independent diagnostic audit was conducted following Round 3.6 (documented locally in remediation/round37/INDEPENDENT_AUDIT.md), followed by the Round 3.7 Stage 1 test-and-freeze checkpoint (documented locally in remediation/round37/STAGE1_CHECKPOINT.md).
While Round 3.6 materially improved candidate binding, CAS integration, and transactional state persistence, the audit confirmed five security defects and three functional regressions.
- S1 — Paired State and Freshness Witness Rollback/Deletion: The authoritative state file and its freshness witness lockfile reside in the same mutable directory (
orchestrator/state). On hosts where access controls are permissive (e.g., Windows inheritedModifyfor Authenticated Users), restoring or deleting both files simultaneously allows a fresh controller process to accept an older state or reinitialize a completed task asINIT/RUNNING. - S2 — Gate Evidence Issued Without Concrete Gate Execution:
VerificationSession._record_success(gate)accepted a caller-provided gate name without requiring unforgeable execution proof from the concrete gate runner. A caller inside the controller process could traverse the FSM, issue all four signed records, and authorize a real CAS without executing the gates. - S3 — Backend Identity Does Not Bind Executing Backend Object: The session captures backend identities during construction but invokes mutable
self.backend. Replacing this object in-process allowed an alternate backend to execute while the signed record claimed the original backend. - S4 — Public Status Mutation Claims Completion Without Integration:
StateManager.set_execution_status('COMPLETED')succeeded fromINIT/RUNNINGwithout verification context, evidence, or CAS execution. - S5 — Failed Recovery Attempt Does Not Consume Capability: Capability tokens were removed only around the transaction following strict recovery validation. If validation failed first, the capability remained active and could be reused once invalid conditions were reverted.
- F1 — Dirty Main Repository Left Inconsistent After Ref-Only CAS: Authentic merge recovery advances the base ref B → C via CAS, but the checked-out worktree and index are not automatically synchronized to the new branch tip.
- F2 — Audit Semantic Failure Strands Task in Active State: If authorized audit recovery encounters a semantic
FAIL, it returnsFalsewhile leaving the task stranded inLOGIC_AUDIT/RUNNINGand consuming a replan budget. - F3 — Missing Preserved Worktree Lacks Explicit Policy: Strict recovery validates a worktree only if it exists, allowing recovery to succeed even after a preserved worktree is removed.
To establish an unassailable baseline before production code remediation, Stage 1 implemented and froze 50 acceptance cases across six test files and two supporting Python files (tests/security_acceptance/ROUND37_FROZEN_SHA256.txt):
- Baseline Reproduction: 16 diagnostic reproduction cases passed in 198.82s.
- Secure Acceptance Execution: 31 genuine FAIL-before assertion failures and 19 passing controls (0 errors, 0 skips, 0 xfails across 50 cases).
- Production Status: Production security remediation remains pending. No production code fixes have been applied; production code remains identical to Round 3.6.
| Invariant / Defect Area | Genuine FAIL-before Cases | Passing Controls | Frozen Scope |
|---|---|---|---|
| S1 Durable State Authority | 4 | 1 | 5 cases in test_s1_durable_authority.py |
| S2 Concrete Gate-Result Authority | 5 | 6 | 11 cases in test_s2_gate_authority.py |
| S3 Executed Implementation Identity | 5 | 5 | 10 cases in test_s3_executed_identity.py & test_s3_implementation_bindings.py |
| S4 Completion Authorization | 13 | 4 | 17 cases in test_s4_completion_authority.py |
| S5 One-Shot Recovery Capability | 4 | 3 | 7 cases in test_s5_recovery_consumption.py |
The test suite reflects three distinct phases of repository development:
- Published Baseline Commit (
c5fd13c): 202 unit and integration tests passing. - Round 3.6 Baseline: 422 collected tests in
tests/. Execution reported 374 passed, 47 failed, and 1 skipped (test_file_policy.py::test_checked_destination_blocks_symlinks_and_reparsedue to platform symlink privileges).- Failure Classification (47 Failures): The independent audit classified these failures into specific categories:
- 33 removed snapshot-write fixture API (
_save_unlocked) - 5 symbolic/placeholder manifest fixtures
- 2 Docker mock fixtures using nonexistent mount paths
- 1 fixture reading UTF-8 source with Windows default encoding (cp1252)
- 1 legacy test reusing a terminal failed task for a review transition
- 1 legacy transition skipping mandatory
SPEC_GATE - 1 obsolete merge signature
- 1 retired candidate identity helper
- 1 tampered snapshot rejected earlier than legacy error regex expects
- 1 unsupported native Windows materialization positive test
- 33 removed snapshot-write fixture API (
- Failure Classification (47 Failures): The independent audit classified these failures into specific categories:
- Current Local Development State (not yet published): Running
pytest tests/ --collect-onlyin the local development repository collects 472 test items (including 131 security acceptance tests and 341 general unit/integration tests). In contrast, the published baseline on GitHub contains 202 passing tests.- Pytest Configuration: In accordance with
pyproject.toml, collection appliesaddopts = "-q --import-mode=importlib --ignore-glob=*e2e*".
- Pytest Configuration: In accordance with
- Restricted Password-v36 Scope: The controller's behavioral verification oracle currently supports only the
password-v36contract. General-purpose factory work, arbitrary task semantics, and authoritative pytest completion remain separate product scope (require_pytest=false). - Unverified Live Docker Daemon: Live Docker daemon execution and isolation remain NOT VERIFIED (
docker_executable=null). The execution backend falls back to local processes with path translation. - Unverified POSIX Materialization: Descriptor-relative, no-follow POSIX materialization is not natively verifiable on Windows hosts.
- Doubled Pipeline Boundaries: In pipeline schedule traces, installed scanner and model boundaries are doubled with instrumented test doubles rather than live external network calls.
Note
The following case studies represent historical production engineering records executed and integrated under the published baseline pipeline. They document the development, triage, and recovery history as originally observed.
-
TASK-001— Token Validator (src/auth/token_validator.py):- Specification:
specs/TASK-001.mddefined token verification using constant-time digest comparison (hmac.compare_digest), rejecting empty or invalid inputs. - Worker: Generated the implementation and unit tests in
tests/test_token_validator.py. - Gates: Passed all deterministic gates and fast-forward merged into
dev.
- Specification:
-
TASK-002— Secure Password Validator (src/auth/password_validator.py):- Specification:
specs/TASK-002.mddefined formal acceptance criteria[AC-01]and[AC-02](minimum length of 8 characters, non-empty, requiring at least one letter and at least one digit) and security invariants[SEC-01]and[SEC-02](passwords must never be written to disk, logged, printed to stdout/stderr, or hardcoded). - Worker: Generated the pure helper function
validate_password(password: str) -> booland 24 unit tests intests/test_password_validator.py. - Gates: Passed
SPEC_GATE,DIFF_GATE,TESTING(100% code coverage),SAST_SCAN, andLOGIC_AUDIT. - Recovery: Successfully integrated into
devusing--resume-mergeafter resolving divergence, verifiable in the Git commit history.
- Specification:
-
TASK-003— In-Memory Rate Limiter (src/security/rate_limiter.py):- Specification:
specs/TASK-003.mddefined a configurable time-window rate limiter classRateLimiter(allow(client_id, now)) with acceptance criteria[AC-01]through[AC-05]and strict security invariants[SEC-01]through[SEC-03](no I/O, no logging of client IDs, no external persistence, encapsulated instance state). - Worker: Generated
RateLimiterand comprehensive unit tests intests/test_rate_limiter.py. - Gates & Integration: Passed all verification gates (
SPEC_GATE,DIFF_GATE,TESTING,SAST_SCAN,LOGIC_AUDIT), initially halted atAUTO_MERGEwithHALT_HUMANbecause of an untracked specification file in the main working tree, and was subsequently integrated intodevafter resolving repository cleanliness.
- Specification:
-
TASK-004— Thread-Safe In-Memory TTL/LRU Cache (src/cache/ttl_cache.py):- Specification:
specs/TASK-004.mddefined a concurrentTTLCachewith configurable capacity (max_size) and default expiration (default_ttl),set,get,delete,clear, LRU eviction for active entries, strict thread safety under concurrent operations ([AC-01]through[AC-10]), and security invariants ([SEC-01]through[SEC-04]). - Worker Attempt 1 & Triage: In Epoch 1, Worker attempt 1 generated initial code and tests. During
TESTING, concurrent execution tests failed due to a race condition. The deterministicTESTINGgate caught the failure and invoked Triage, which inspected the execution traceback and diagnosed the exact concurrency root cause. - Worker Attempt 2: Guided by the structured Triage diagnosis, Worker attempt 2 corrected the concurrency handling and internal locking.
- Gates & Suspension: Attempt 2 passed
DIFF_GATE,TESTING(100% code coverage across 19 unit tests intests/test_ttl_cache.py), andSAST_SCAN(zero findings in Semgrep and Bandit). AtLOGIC_AUDIT, an upstream API rate limit (HTTP 429) was encountered; the pipeline safely transitioned toHUMAN_REVIEWwithout penalizing Worker replan budgets. - Recovery & Integration: Once the provider rate limit cleared, recovery was executed via
--resume-audit TASK-004. The audit passed, the worktree was cleanly detached, and the task completed throughAUTO_MERGE -> COMPLETEDvia Fast-Forward merge intodev.
- Specification:
-
TASK-005— Multi-Tenant Authorization Policy Engine (src/security/authorization_policy.py):- Specification:
specs/TASK-005.mddefined a stateless, deny-by-default authorization policy classAuthorizationPolicywithis_allowed(subject_tenant: str, resource_tenant: str, roles: set[str], action: str) -> bool([AC-01],[AC-02]). Acceptance criteria mandated default deny ([AC-03]), strict tenant isolation where access is permitted only whensubject_tenant == resource_tenantacross all roles includingadmin([AC-04],[AC-05]), granular role permissions (viewer:read;editor:read,write;admin:read,write,delete) ([AC-06]), and strict denial on unknown roles, unknown actions, or empty role sets ([AC-07]through[AC-10]). Security invariants mandated total prohibition of cross-tenant access ([SEC-01]), fail-closed semantics on missing/malformed/unknown inputs ([SEC-02]), zero hidden bypasses or admin tenant overrides ([SEC-03]), non-emission of tenant IDs, roles, and decisions to disk/logs/stdout/stderr ([SEC-04]), immutable stateless execution with no mutable global state ([SEC-05]), and zero dynamic execution (eval,exec, subprocesses, network, or external persistence) ([SEC-06]). - Worker: In Epoch 1 attempt 1, the Worker implemented
AuthorizationPolicyinsrc/security/authorization_policy.pyand 70 comprehensive unit tests intests/test_authorization_policy.py. - Gates & Suspension: Passed
SPEC_GATE,DIFF_GATE,TESTING(100% code coverage across all 70 unit tests), andSAST_SCAN(zero findings in Semgrep and Bandit). AtLOGIC_AUDIT, an upstream API availability failure occurred (persistent HTTP 503 Service Unavailable). The pipeline safely transitioned toHUMAN_REVIEWwithout consuming Worker attempts or replan budgets. - Reconciliation & Integration: After resolving Git divergence against
devwhere independent updates had been integrated, recovery was executed via--resume-audit TASK-005. The successful final audit was performed bygemini:gemini-3.8-flash. The worktree was cleanly detached and fast-forward merged intodev(AUTO_MERGE -> COMPLETED).
- Specification:
Note
The tree below reflects the Current Local Development State (not yet published), illustrating the modules introduced during Round 3.6 and Round 3.7 research alongside the published baseline codebase.
.
├── .env.example # Safe credentials template
├── .gitignore # Git exclusion rules
├── .semgrepignore # SAST exclusion rules
├── CONTRIBUTING.md # Contribution guidelines (English)
├── CONTRIBUTING_ES.md # Guía de contribución (Español)
├── LICENSE # MIT License
├── pyproject.toml # Packaging metadata and pytest configuration
├── README.md # Primary English documentation
├── README_ES.md # Spanish documentation
├── RULES.md # Worker governance directives
│
├── .github/workflows/ # Continuous integration
│ └── ci.yml # GitHub Actions test workflow
│
├── adapters/ # LLM provider adapters
│ ├── contracts.py # Pydantic v2 structured schemas
│ ├── deepseek_adapter.py # Worker / code generation
│ ├── gemini_adapter.py # Architect, Triage, Security Filter, Logic Security
│ ├── glm_adapter.py # Alternative Logic Security provider
│ ├── network_retry.py # HTTP resilience with exponential backoff
│ ├── qwen_adapter.py # DashScope / Qwen Logic Security fallback
│ └── sanitizer.py # Secret redaction and credential hygiene
│
├── orchestrator/ # Core orchestrator and FSM
│ ├── config.json # Budgets, roles, and thresholds
│ ├── env_loader.py # Environment variable loader
│ ├── evidence_store.py # Cryptographic HMAC-SHA256 evidence store
│ ├── execution_backend.py # Sandboxed process & execution management
│ ├── gate_controller.py # VerificationSession & gate coordination
│ ├── logic_audit.py # Logic security audit runner
│ ├── state_manager.py # Transactional FSM controller & StateLock
│ ├── verification_context.py # Immutable verification context tuple
│ ├── verification_manifest.py # Git object manifest deriving C^{tree}
│ └── verification_policy.json # Verification policies and security rules
│
├── remediation/ # Independent audit & remediation records
│ ├── round36/ # Round 3.6 implementation report
│ │ └── REPORT.md
│ └── round37/ # Round 3.7 independent audit & Stage 1 checkpoint
│ ├── INDEPENDENT_AUDIT.md # Independent audit findings (S1–S5, F1–F3)
│ ├── REPRODUCTION.md # Baseline reproduction instructions
│ ├── ROUND37_REMEDIATION_CONTRACT.md # Stage 1-3 remediation contract
│ └── STAGE1_CHECKPOINT.md # Stage 1 test-and-freeze checkpoint report
│
├── scripts/ # Deterministic verification gates
│ ├── diff_gate.py # File boundary enforcement
│ ├── discovery.py # Repository structure analysis
│ ├── file_policy.py # Filesystem & path traversal policy
│ ├── merge_gate.py # Atomic Git Compare-and-Swap (CAS) gate
│ ├── sast_runner.py # Semgrep & Bandit security runner
│ ├── semgrep_rules.yml # Local versioned Semgrep rules
│ ├── spec_gate.py # Specification contract validator
│ ├── test_runner.py # Pytest runner with coverage enforcement
│ └── worktree_manager.py # Git worktree isolation manager
│
├── specs/ # Task specifications
│ ├── TASK-001.md # Token validator specification
│ ├── TASK-002.md # Secure password validator specification
│ ├── TASK-003.md # In-memory rate limiter specification
│ ├── TASK-004.md # Thread-safe TTL/LRU cache specification
│ ├── TASK-005.md # Multi-tenant authorization policy specification
│ └── TEMPLATE.md # Canonical specification template
│
├── src/ # Production code generated and integrated
│ ├── auth/
│ │ ├── password_validator.py
│ │ └── token_validator.py
│ ├── cache/
│ │ └── ttl_cache.py
│ └── security/
│ ├── authorization_policy.py
│ └── rate_limiter.py
│
├── tests/ # Automated test suite (472 collected tests)
│ ├── security_acceptance/ # Hardened security acceptance test suites
│ │ ├── ROUND36_FROZEN_SHA256.txt # Round 3.6 frozen acceptance manifest
│ │ ├── ROUND37_FROZEN_SHA256.txt # Round 3.7 Stage 1 frozen manifest (50 tests)
│ │ ├── round36/ # 8 frozen Round 3.6 acceptance test suites
│ │ └── round37/ # 6 frozen Round 3.7 test files & 2 supporting files (50 cases)
│ ├── test_audit_fallback.py
│ ├── test_authorization_policy.py
│ ├── test_diff_gate.py
│ ├── test_diff_gate_immutable.py
│ ├── test_fail_closed_gates.py
│ ├── test_file_policy.py
│ ├── test_password_validator.py
│ ├── test_rate_limiter.py
│ ├── test_resume_audit.py
│ ├── test_resume_merge.py
│ ├── test_sast_runner.py
│ ├── test_spec_gate.py
│ ├── test_state_concurrency.py
│ ├── test_state_manager.py
│ ├── test_state_transactions.py
│ ├── test_token_validator.py
│ ├── test_ttl_cache.py
│ └── test_worktree_manager.py
│
└── trusted_tests/ # Controller-side hidden challenge suite
└── suite.py # 149 hidden behavioral challenge cases
- Pending Security Remediation: Confirmed vulnerabilities S1–S5 await Stage 2 production remediation.
- Restricted Verification Oracle: Only the
password-v36behavioral contract is currently evaluated by the hidden-challenge controller oracle (require_pytest=false). - Unverified Isolation Infrastructure: Live Docker isolation and native POSIX descriptor-relative materialization remain unverified on the Windows development host.
- Single Task Sequential Execution: Tasks are processed sequentially; concurrent multi-task pipelining is in development.
- Commercial APIs: Requires commercial API keys (Google Gemini, DeepSeek, or Zhipu GLM) unless running in
--simulatemode.
- Stage 2 Remediation: Remediate security defects S1–S5 in production code against the 50 frozen Stage 1 tests.
- Stage 3 Independent Re-Audit: Conduct full independent audit verification of remediated invariants.
- Functional Regressions: Resolve F1 worktree synchronization, F2 audit failure state stranding, and F3 worktree presence policies.
- General-Purpose Verification: Expand the trusted verifier beyond
password-v36to general tasks with authoritative pytest completion. - Local Model Support: Integrate Ollama / vLLM for zero token cost on Worker and Triage roles.
- Polyglot Gate Support: Node.js/TypeScript (Vitest) and Rust (
cargo test).
Note
Scope of Setup and Quick Start Instructions:
The installation and quick start steps below apply directly to the published baseline codebase (c5fd13c) available in this repository. The newer controller-side hidden challenge suite, transactional StateLock, and Stage 1 acceptance test suites described in Section 7 and 8 belong to the local research and remediation state and are not yet part of the public release.
- Python:
>= 3.10 - Git:
>= 2.30 - Semgrep:
>= 1.0.0 - Compatible with Windows (PowerShell) and Linux / macOS.
# 1. Clone the repository
git clone https://github.com/LE0ST/autonomous-multi-agent-software-factory.git
cd autonomous-multi-agent-software-factory
# 2. Create and activate a virtual environment
python -m venv .venv
# On Windows PowerShell:
.venv\Scripts\Activate.ps1
# On Linux / macOS:
# source .venv/bin/activate
# 3. Install dependencies and package in editable mode
pip install -e .Copy .env.example to .env:
cp .env.example .envAdd your API keys to .env:
GEMINI_API_KEY=your-gemini-api-key
DEEPSEEK_API_KEY=your-deepseek-api-key
GLM_API_KEY=your-glm-api-key # Optional for Zhipu GLM
DASHSCOPE_API_KEY=your-dashscope-api-key # Optional for Qwen fallback (qwen3.8-flash)Executing a task safely within the factory requires adherence to clean-state invariants:
-
Always Work from
dev: Switch todevand ensure your local branch is strictly up-to-date with remote upstream:git checkout dev git pull --ff-only origin dev
-
Define the Specification Contract: Create a new specification file (e.g.
specs/TASK-004.md) adhering strictly tospecs/TEMPLATE.md. Define unambiguous Acceptance Criteria ([AC-xx]), Security Invariants ([SEC-xx]), explicit file boundary whitelists (Allowed files/Forbidden files), and the complete Test Matrix. -
Stage and Commit the Specification Before Pipeline Execution:
[!IMPORTANT] Pre-Execution Invariant: You must stage and commit the specification contract before launching
orchestrator.py. Ensure thatgit status --shortreturns a completely clean working tree. The final merge gate (merge_gate.py) verifies working tree cleanliness viagit status --porcelainand will reject integration if untracked or modified files are detected in the repository root.git add specs/TASK-004.md git commit -m "spec: add TASK-004 contract" git status --short # Must be completely empty
-
Execute the Orchestrator:
python orchestrator.py TASK-004
-
Verify the Integration Outcome: After pipeline execution finishes, verify system health:
python -m pytest -v tests/ git status --short git log --oneline --decorate -5
Distributed under the MIT License. See LICENSE for details.