KUMA Python SDK
Knowledge-grounded Universal Measurement for Agents
KUMA is a framework-neutral Python SDK for evaluating AI Agents. It delivers a test Case as step-by-step Inputs, captures bounded Evidence, and produces a Judgment through official or custom Providers. KUMA does not run your Agent or expose private evaluation logic.
- Run repeatable Case-to-Judgment evaluations.
- Use official services or custom Case and Judge Providers, including fully local workflows.
- Capture structured Evidence from results, file changes, logs, and optional OpenTelemetry traces.
- Select a versioned Strategy Group and describe the Agent with an Agent Profile.
Requires Python 3.10 or newer:
python -m pip install --upgrade kuma-defuzexThe package is kuma-defuzex; Python imports and the CLI use kuma.
Run a deterministic local check without an account, API key, Docker, or network:
kuma quickstartSuccessful output starts with:
Local check: PASS
Score: 100/100
Reason: Output exactly matched the published rule.
This checks the bundled local example, not your Agent or the hosted Judge. To evaluate your Agent, follow the SDK guide.
For official Case generation, create_run(..., difficulty="D1") selects challenge
count and intensity: D0 injects zero problems; D1 injects one obvious, low-intensity
problem; D2 injects two subtler or composed problems requiring stronger recognition,
recovery and verification. Necessary inputs and solvability must be preserved.
max_steps remains the same upper bound; D2 does not add a step or change Judge
severity. No measured failure rate is promised; the service builds the challenge.
- Run an official evaluation with the full Case and Judge workflow.
- Observe an Agent locally without a Case, Judge, or automatic upload.
- Try the full-stack example with KUMA and mini-SWE-agent in Docker. It may use service credit and model budget.
SDK guide · Python API reference · Versions and releases · Detailed Judge assessment