I build applied AI systems around the parts that usually fail after the demo: tool execution, trust boundaries, replay, evals, observability, scheduling and operator workflows.
My work is mostly local-first and inspectable. I prefer deterministic cores, explicit failure modes and evidence you can verify over opaque orchestration.
yashkhou.com · Projects · @yashkhou on X
| Project | What it explores |
|---|---|
| Commander Plus | A local-first agent workstation: MCP tools, reusable skills, durable workspaces, browser/computer control and persistent project context. |
| Glyph | A semantic design system for AI-generated interfaces with constraints, stable IDs, diffs and deterministic rendering targets. |
| Verify | Verification infrastructure for AI-written artifacts, code and reversible actions. |
A set of focused, dependency-light tools for testing and hardening agent infrastructure. Each repository is built around a small deterministic core with tests, CI, examples and tagged releases.
| Project | Reliability boundary |
|---|---|
| Context Firewall | Provenance-aware trust boundaries and fail-closed checks before privileged actions. |
| Tool Contract Fuzzer | Deterministic valid/invalid JSON-Schema cases, boundary mutations and shrinking. |
| MCP Chaos | Deterministic JSON-RPC/MCP fault injection, method-scoped cadence and wire-level chaos testing. |
| Agent Replay | Redacted, hash-chained agent event logs with integrity-aware deterministic replay. |
| Failure Corpus | Normalize, fingerprint and deduplicate failures into reusable regression corpora. |
| Handoff Spec | Canonical, digestible handoffs with authority boundaries and continuation invariants. |
| Toolgraph Profiler | Critical paths, retries, fan-out, lock pressure and idle time in tool-call traces. |
| Agent Policy Compiler | Explainable policy-as-code for deterministic allow/deny decisions. |
| Eval Capsule | Portable eval fixtures, assertions and integrity-checked reproducible archives. |
| Schema Evolution Guard | Detect compatibility-breaking changes in evolving tool and API schemas. |
| Context Budgeter | Token allocation, duplicate detection and policy-collision analysis for prompt context. |
| Agent Scheduler Sim | Deterministic worker/retry/starvation simulation for agent scheduling policies. |
- OpenRetention — self-hosted customer-success software with explainable health scoring, churn risk, renewals and revenue-at-risk prioritisation.
- AI Real Estate CRM — evidence-aware CRM logic for property, owner and lead workflows with deterministic matching and voice-agent handoff.
- AI Product Sourcing Agent — marketplace-agnostic sourcing engine for query planning, normalization, deduplication and evidence-based ranking.
- Opened Stellar-agentic #429 to repair repository-specific README links.
- Reviewed pydantic-ai #8969, a regression fix preventing shared
StructuredDictschema metadata from leaking between output types. - Added current-main implementation analysis to MCP Python SDK #1933 around stdio ownership and process-stream lifecycle.
Agent infrastructure, evals and verification, local-first tooling, reliable computer use, and product systems where AI has to survive contact with real workflows.


