Pythia is an end-to-end AI forecasting system for humanitarian questions. It scans countries for emerging hazards, turns triage into forecastable questions, runs an LLM ensemble to produce Subjective Probability Distributions (SPD), and stores forecasts, research, and diagnostics in a DuckDB-backed system of record.
- What is Pythia?
- System at a glance
- Core components
- Regime Change (RC): out-of-pattern detection
- Hazard Tail Packs (HTP): hazard-specific trigger evidence
- Retriever web research (shared evidence packs)
- Data model / DuckDB tables
- Model management
- Quickstart (local)
- Running in GitHub Actions
- Dashboard
- Downloads / exports
- Troubleshooting
- Cross-links
Pythia connects historical hazard data, web research, and LLM reasoning into a repeatable pipeline for early-warning forecasts. It produces triage signals, structured research briefs, and probabilistic forecasts across hazards, then serves those outputs through a DuckDB system of record, a FastAPI service, and a Next.js dashboard.
Resolver facts/base rates → Horizon Scanner per-hazard pipeline:
Per hazard: RC grounding → RC assessment → triage grounding → triage (single pass each by default)
+ Structured data: ReliefWeb, ACAPS, IPC, ACLED political, ENSO, seasonal TC, HDX Signals, GDELT
+ Conflict forecasts: VIEWS, conflictforecast.org, ACLED CAST
+ Adversarial checks (RC L1+) → Calibration advice
→ Track 1: full ensemble SPD + scenarios
→ Track 2: single-model SPD
→ DuckDB → API → Dashboard + downloads
- System of record: DuckDB (
PYTHIA_DB_URL/app.db_url) holds HS triage, questions, structured data, forecasts, and diagnostics. - End-to-end flow: Resolver facts and base rates inform the HS per-hazard pipeline. Each hazard gets its own RC assessment (with dedicated grounding) and triage assessment (with separate grounding), each a single Gemini Flash pass by default (
PYTHIA_HS_RC_PASSES/PYTHIA_HS_TRIAGE_PASSES=2restore the 2-pass merge). Structured data connectors pull evidence from humanitarian, climate, and conflict-forecast APIs including ENSO state, seasonal TC forecasts, HDX Signals, and ACLED CAST. Adversarial checks run for RC L2+ cases. Forecaster routes questions by track — Track 1 (RC-elevated) through the full ensemble, Track 2 (priority, no RC) through a single model — and writes results back to DuckDB.
- Resolver (facts + base rates): Resolver tables in DuckDB (
facts_resolved,facts_deltas,snapshots) provide historical context. Schema is defined inpythia/db/schema.py. - Horizon Scanner (HS):
python -m horizon_scanner.horizon_scannertriages countries and hazards, writeshs_runs/hs_triage, seedsquestions/question_research, and generates country evidence packs. HS runs a fully per-hazard pipeline: each hazard (ACE, DR, FL, HW, TC) gets its own RC grounding call, RC assessment, triage grounding call, and triage assessment. RC and triage each run a single Gemini Flash pass by default (a 2-pass merge is available viaPYTHIA_HS_RC_PASSES/PYTHIA_HS_TRIAGE_PASSES=2, useful only with a distinct pass-2 model). RC grounding uses TRIGGER/DAMPENER/BASELINE signal categories; triage grounding uses SITUATION/RESPONSE/FORECAST/VULNERABILITY categories with longer recency windows. An ACLED low-activity filter skips ACE triage for countries with minimal recent conflict (0 fatalities in 2 months AND <25 in 12 months). - Retriever web research (shared evidence packs): When enabled, the retriever builds evidence packs reused across HS, research, and SPD prompts. The shared retriever defaults to
gemini-3.5-flashwhenPYTHIA_RETRIEVER_ENABLED=1andPYTHIA_RETRIEVER_MODEL_IDis unset; seepythia/web_research/web_research.py. - Structured data connectors: Authoritative humanitarian, climate, and conflict-forecast data pulled from specialist APIs and stored in DuckDB for direct prompt injection. Replaced the former Research LLM stage. See
pythia/acaps.py,pythia/fewsnet_food_security.py,horizon_scanner/reliefweb.py,pythia/acled_political.py,horizon_scanner/enso/,horizon_scanner/seasonal_tc/,horizon_scanner/hdx_signals.py,resolver/connectors/acled_cast.py. - Adversarial evidence checks: Counter-evidence searches for RC Level 1+ cases, stored in
hs_adversarial_checks. Seepythia/adversarial_check.py. - Hazard-specific reasoning: Per-hazard forecasting instructions injected into SPD prompts. See
forecaster/hazard_prompts.py. - Calibration advice: Per-hazard/metric calibration guidance generated from historical performance, injected into forecasting prompts. See
pythia/tools/generate_calibration_advice.py. - Forecaster SPD v2 ensemble:
python -m forecaster.cli --mode pythiaruns SPD v2 prompts across the active ensemble and writes per-model outputs (forecasts_raw) and aggregated results (forecasts_ensemble). Questions are routed by track: Track 1 (RC-elevated, level > 0) uses the full multi-model ensemble; Track 2 (priority, no RC) uses a lightweight single-model path (Gemini Flash). - Sibyl (parallel deep-research track):
python -m sibyl.runindependently re-forecasts 25 affected/fatalities questions per run, chosen floor-then-fill (each hazard's three most volatile first, then the most volatile remaining under a cap of ten per hazard), using agentic open-web research (Claude Opus + Brave Search). K independent belief-updating trials per question are linear-pooled into an SPD written beside the standard track (model_name='sibyl'inforecasts_raw/forecasts_ensemble), so scoring grades the two tracks head-to-head. Takes two inputs from Fred's own data: the Resolver base rate, and one reading of the resolving source per question shown to every trial (a live month-to-date ACLED read for conflict deaths whenSIBYL_LIVE_LOOKUPS_ENABLED=1, the newest Phase 3+ rows for drought, the last six months of resolving rows and GDACS alerts for floods and cyclones;sibyl/resolver_reading.py). A hashed half of questions (SIBYL_PACK_SHARE) is also shown a starting pack of the pipeline's structured feeds, never the ensemble's own research, as a measured experiment (sibyl/pack.py). Hard run budget cap (default $60) and a wall-clock cap on the trial loop (default 180 minutes). A monthly advice loop (python -m sibyl.advice, run by the calibration workflow) measures Sibyl's own resolved record per class and, once a class holds 20 scored questions, shows Sibyl a short YOUR TRACK RECORD note on the next run, to half its questions. Seesibyl/andsibyl/DISCOVERY.md. - Scenarios (Track 1 only): When an ensemble SPD is available, scenarios can be generated for Track 1 questions and written back to DuckDB.
- Batch-API staged pipeline (opt-in): The HS and forecaster LLM families can run through provider Batch APIs at 50% off both input and output tokens.
python -m horizon_scanner.horizon_scanner --stage {hs_submit,hs_rc_collect,hs_finalize}andpython -m forecaster.cli --phase {submit,collect}split each family into a submit wave (prompts built + enqueued inpythia/llm_batch.py, byte-identical bodies to the sync path) and a collect wave (results replayed, per-item sync fallback on any miss). GitHub orchestration:pythia_pipeline_stage.yml(staged runs — the production monthly entry since 2026-07-27, cron on the 13th entering aths_submit) +poll_llm_batches.yml(15-min poller that advances the pipeline when batches finish). Code defaults are OFF (PYTHIA_BATCH_API_ENABLED=0) so the--stage full/--phase fullpaths and local runs are unchanged; the production workflows opt in.run_horizon_scanner.ymlis the manual synchronous fallback. Related: static-first prompt ordering + provider prompt caching (PYTHIA_PROMPT_V3_ORDER/PYTHIA_PROMPT_CACHE_ENABLED, on in the production workflows;scripts/compare_prompt_order_runs.pyis the regression harness). Anthropiccache_controlinside batch bodies is a separate opt-in (PYTHIA_BATCH_PROMPT_CACHE, on inpythia_pipeline_stage.ymlsince 2026-08-03, with a 1h TTL because batches routinely outlive the 5-minute default) — confirm it is paying for itself by checkingcache_read_input_tokensin the batch-economics / post-run diagnostics after each cycle. - Dashboard + API: FastAPI serves
/v1/*endpoints for the dashboard, downloads, and diagnostics. The Next.js UI lives inweb/.
Definition: Regime Change (RC) is the likelihood of a generating-process shift within the next 1–6 months that makes historical base rates less reliable. RC is assessed per-hazard in a dedicated LLM pipeline step that runs before triage. Each hazard type gets its own RC prompt with calibration anchors (default likelihood 0.05, expected ~80% at ≤0.10), dedicated grounding call, and hazard-specific data injection (ENSO context for climate hazards, conflict forecasts for ACE, seasonal TC for TC). RC runs a single Gemini Flash pass by default (2-pass available via PYTHIA_HS_RC_PASSES=2) and writes to hs_triage.
Stored fields in hs_triage:
regime_change_likelihood(probability)regime_change_direction(up,down,mixed,unclear)regime_change_magnituderegime_change_score(computed)regime_change_level(severity)regime_change_windowregime_change_json(full normalized object)
Scoring + levels (defaults; env-overridable via PYTHIA_HS_RC_LEVEL*_LIKELIHOOD):
score = likelihood × magnitude- Level 0:
likelihood < 0.15 - Level 1:
likelihood ≥ 0.15 - Level 2:
likelihood ≥ 0.35 - Level 3:
likelihood ≥ 0.55
Behavioral implications:
- Track routing: Questions with RC level > 0 are assigned Track 1 (full multi-model ensemble); priority questions without RC are Track 2 (lightweight single-model).
need_full_spdoverride: HS forcesneed_full_spd = TRUEwhen RC is elevated (Level ≥2) even if the triage tier is quiet.- Research v2: RC is surfaced in the research prompt, and RC-elevated hazards require at least one
regime_shift_signalsentry (or a rebuttal). - SPD v2: RC guidance is embedded in the SPD prompt, and the model must include a sentence starting with
RC:inhuman_explanationdescribing how RC affected the SPD.
See docs/hs_regime_change.md for full details and env overrides.
Hazard Tail Packs are hazard-scoped evidence bundles generated by HS when RC is elevated. They provide targeted trigger/counter-trigger evidence for downstream Research and SPD.
Trigger rules:
- Generated only for RC Level ≥1 hazards.
- Limited to max 2 hazards per country per HS run (
PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_PER_COUNTRY, default2). - Tail packs are off by default unless
PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1.
Storage table: hs_hazard_tail_packs (one row per hs_run_id × iso3 × hazard_code), including:
rc_level,rc_score,rc_direction,rc_windowquery,report_markdown,sources_json,grounded,grounding_debug_jsonstructural_context,recent_signals_json,created_at
Output format:
recent_signalsbullets are emitted asTRIGGER | <window> | <direction> | <signal>,DAMPENER | ...,BASELINE | ....
Caching behavior:
- Tail packs are reused per
(hs_run_id, iso3, hazard_code); HS skips regeneration if the row already exists.
Phase A / Phase B behavior:
- Phase A (Research): When present, the tail pack is injected into Research v2 evidence (
hs_hazard_tail_pack). - Phase B (SPD v2): For RC-elevated hazards, the tail pack is injected into the SPD prompt. Signals are capped at 12 bullets (
PYTHIA_SPD_TAIL_PACKS_MAX_SIGNALS, default12).
Relevant env vars (defaults):
PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED(0)PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_PER_COUNTRY(2)PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_SOURCES(defaults toPYTHIA_HS_EVIDENCE_MAX_SOURCES)PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_SIGNALS(defaults toPYTHIA_HS_EVIDENCE_MAX_SIGNALS)PYTHIA_SPD_TAIL_PACKS_MAX_SIGNALS(12)
See docs/hs_hazard_tail_packs.md for full details.
The question-level web research pipeline (fetch_evidence_pack / _build_question_evidence_queries) is deprecated. SPD prompts now receive structured data directly via _load_structured_data(), including conflict forecasts, ReliefWeb, HDX Signals, ACAPS, ACLED political, NMME, ENSO, seasonal TC, GDACS event history, and adversarial checks. The question_research table is no longer populated by the pipeline (only placeholder rows). Env vars PYTHIA_RETRIEVER_ENABLED, PYTHIA_WEB_RESEARCH_ENABLED, PYTHIA_SPD_WEB_SEARCH_ENABLED, PYTHIA_FORECASTER_SELF_SEARCH are set to "0" in the workflow.
Grounding for RC and triage uses OpenAI GPT-4.1-mini web search (primary) with Gemini Google Search as fallback. Override with PYTHIA_GROUNDING_PRIMARY_BACKEND=gemini.
Pythia ingests NMME (North American Multi-Model Ensemble) seasonal temperature and precipitation anomaly forecasts from the CPC FTP server. These provide structured climate context for drought, flood, and tropical cyclone assessments.
- Source:
ftp://ftp.cpc.ncep.noaa.gov/NMME/realtime_anom/ENSMEAN/— ensemble mean anomalies at 1° resolution, updated ~9th of each month with 7 lead months. - Probability of a below-normal month (since Oct 2026):
ftp://ftp.cpc.ncep.noaa.gov/NMME/prob/netcdf/prate.YYYYMM.prob.adj.mon.nc, CPC's calibrated tercile probabilities, stored as variableprate_prob_below(0..1). The PA machine's drought gate reads it at 0.5; cells CPC does not forecast are masked, not read as zero.python -m scripts.ci.nmme_prob_crossing 12reports per-country crossing frequency for candidate thresholds (needs FTP access to CPC). - Processing: Country-level area-weighted averages using
xarray+regionmask(Natural Earth admin-0 boundaries). Anomalies expressed in σ (standard deviations from climatology) with derived tercile categories (above/below/near normal). - Storage:
seasonal_forecaststable in Pythia DuckDB (~2,700 rows per monthly update: 195 countries × 2 variables × 7 leads). - Injection: Automatically loaded into HS triage and RC prompts via the existing
climate_dataparameter for DR, FL, TC hazards. Also injected into forecaster research and SPD prompts viaresearch_json. - Run manually:
python -m resolver.tools.ingest_nmme(or--year-month YYYYMMfor a specific month,--dry-runto preview). - Automation: NMME runs as the
nmmesource insideresolver_update.ymlPhase 3 (11th monthly) andingest-structured-data.yml. (The standaloneingest-nmme.ymlworkflow was removed July 2026.)
Pythia ingests external conflict forecasts from three independent quantitative sources and incorporates qualitative expert assessments from a fourth, providing forward-looking signals for the Armed Conflict (ACE) hazard.
- VIEWS (Uppsala/PRIO): ML-based country-month state-based conflict forecasts from the VIEWS Forecasting project. Provides predicted fatalities (
views_predicted_fatalities) and probability of ≥25 battle-related deaths (views_p_gte25_brd) at 1–6 month lead times. Good at trends and baseline levels; weaker at sudden onset. - conflictforecast.org (Mueller/Rauh): News-based armed conflict risk scores from the Conflict Forecast project. Provides armed conflict risk at 3-month and 12-month horizons (
cf_armed_conflict_risk_3m,cf_armed_conflict_risk_12m) and violence intensity outlook (cf_violence_intensity_3m). Better at detecting shifts and escalation signals from media coverage patterns. - ACLED CAST (Conflict Alert System Tool): Event-count forecasts from ACLED's CAST API, disaggregated by event type: total events (
cast_total_events), battles (cast_battles_events), explosions/remote violence (cast_erv_events), and violence against civilians (cast_vac_events) at 6-month lead. CAST forecasts event counts (not fatalities) and the event-type breakdown reveals shifts in the character of violence that aggregate measures would miss. Seeresolver/connectors/acled_cast.py.
- ICG CrisisWatch: International Crisis Group's monthly CrisisWatch bulletin. Per-country directional assessments (Deteriorated/Improved/Unchanged) and forward-looking "On the Horizon" flags (~3 conflict risks + ~1 resolution opportunity per month). Primary data source:
scripts/refresh_crisiswatch.pyruns monthly in CI viarefresh-crisiswatch.yml, fetching the server-rendered page from the Internet Archive's Wayback Machine (--source wayback— the ICG site's Cloudflare blocks direct CI fetches) and parsing with BeautifulSoup; a local Playwright path (--source live) is kept for manual refreshes. Secondary: Gemini grounding calls during HS runs. Data injected into RC and triage evidence for ACE hazards.
- Table:
conflict_forecastsin Pythia DuckDB (keyed on source, iso3, hazard_code, metric, lead_months, forecast_issue_date). - HS integration: Automatically loaded into RC and triage prompts for ACE hazards via
horizon_scanner/conflict_forecasts.py. Includes staleness warnings for data >45 days old. CAST section notes that it measures event counts rather than fatalities. - Forecaster integration: Injected into research prompts for ACE questions as structured quantitative anchors.
- Run manually:
python -m resolver.tools.fetch_conflict_forecasts(or--sources views conflictforecast_org acled_cast,--dry-runto preview).
Pythia pulls structured humanitarian, climate, and conflict-forecast data from authoritative APIs, stores it in DuckDB, and injects it into pipeline prompts (HS triage, RC assessment, and/or forecaster SPD). This replaced the former Research LLM stage with deterministic, reproducible evidence injection.
| Source | Module | What it provides | DuckDB table(s) |
|---|---|---|---|
| ReliefWeb | horizon_scanner/reliefweb.py |
Humanitarian situation reports (45-day window, up to 15/country) | reliefweb_reports |
| ACAPS INFORM Severity | pythia/acaps.py |
Crisis severity scores + trend | acaps_inform_severity, acaps_inform_severity_trend |
| ACAPS Risk Radar | pythia/acaps.py |
Forward-looking risk with triggers | acaps_risk_radar |
| ACAPS Daily Monitoring | pythia/acaps.py |
Analyst-curated daily updates | acaps_daily_monitoring |
| ACAPS Humanitarian Access | pythia/acaps.py |
Access constraint scores (HS triage only) | acaps_humanitarian_access |
| ACLED Political Events | pythia/acled_political.py |
Event-level political data (ACE/DI hazards only) | acled_political_events |
| NMME Seasonal Forecasts | resolver/tools/ingest_nmme.py |
Temp/precip anomalies, 7-month lead (DR/FL/TC) | seasonal_forecasts |
| ENSO State/Forecast | horizon_scanner/enso/ |
ENSO phase + strength COMPUTED from the ONI (NOAA ERDDAP → CPC weekly → CPC ONI table), plus the IRI Quick Look's 9-season outlook, plume and IOD (DR/FL/TC) | enso_state (DB-first; a run that resolves no index carries the last good record forward, labelled stale) |
| Seasonal TC Forecasts | horizon_scanner/seasonal_tc/ |
Basin-level TC activity forecasts from TSR, NOAA CPC, BoM, Météo-France La Réunion (SWI), and an IMD/NIO climatology block across 8 basins (TC only) | seasonal_tc_outlooks, seasonal_tc_context_cache (DB-first) |
| HDX Signals | horizon_scanner/hdx_signals.py |
OCHA automated crisis alerts: conflict, displacement, food insecurity, agricultural stress (all hazards) | hdx_signals (DB-first, falls back to CSV) |
| VIEWS | resolver/connectors/views.py |
ML-based conflict fatality predictions (ACE, 1–6 month leads) | conflict_forecasts |
| conflictforecast.org | resolver/connectors/conflictforecast.py |
News-based conflict risk scores (ACE, 3m/12m) | conflict_forecasts |
| ACLED CAST | resolver/connectors/acled_cast.py |
Event-count forecasts by type: total/battles/ERV/VAC (ACE, 6-month lead) | conflict_forecasts |
| FEWS NET IPC | resolver/connectors/fewsnet_ipc.py |
Phase 3+ population estimates (DR hazard; Current Situation + Most Likely). The stored value is the LOWER bound of FEWS NET's published range (e.g. 1,000,000 for "1.0 - 2.49 million") and value_high its upper bound, which the Phase 3+ prompts print beside it; probe_fewsnet_values.yml re-checks this |
facts_resolved (via Resolver pipeline) |
| IDMC conflict displacement | resolver/ingestion/idmc_conflict.py |
Conflict displacement per country and month, the ACE/PA resolution series: IDMC's recommended figures only (triangulations dropped), records spanning more than 31 days and months above the population held out. A month resolves 90 days after it ends; a missing month is zero only for a country reported in 8 of the 12 months before and bracketed by a later report (base_rate_spd.resolve_conflict_month); probe_idmc_conflict.yml re-measures the lag |
facts_resolved (hazard ACE, metrics new_displacements / new_displacements_held) |
| GDACS | resolver/connectors/gdacs.py |
Disaster population exposure + event occurrence (FL/DR/TC) | facts_resolved (via Resolver pipeline) |
| ICG CrisisWatch | horizon_scanner/crisiswatch.py + scripts/refresh_crisiswatch.py |
Expert conflict arrows + "On the Horizon" flags (ACE RC + triage + SPD). Wayback Machine scraper (monthly, in CI) + Gemini grounding (runtime fallback) | crisiswatch_entries |
| GDELT | pythia/gdelt.py |
Media-derived conflict intensity indicators from GDELT 1.0 daily event exports (CAMEO-tiered, Goldstein, tone; ACE only) | gdelt_conflict_indicators |
For RC Level 1+ cases, pythia/adversarial_check.py runs counter-evidence web searches and synthesizes results into structured output (counter-evidence, historical analogs, stabilizing factors, net assessment). Stored in hs_adversarial_checks.
pythia/tools/generate_calibration_advice.py generates per-hazard/metric calibration guidance from historical scores. Stored in calibration_advice and injected into forecasting prompts.
ACAPS_EMAIL,ACAPS_PASSWORD: ACAPS API credentials (required for ACAPS feeds).PYTHIA_DB_URL: DuckDB path (shared with all connectors).
Key tables (see pythia/db/schema.py and SCHEMAS.md for canonical fields):
- HS:
hs_runs,hs_triage(includes RC columns),hs_country_reports,hs_hazard_tail_packs,hs_adversarial_checks - Seasonal climate:
seasonal_forecasts(NMME country-level temp/precip anomalies) - Conflict forecasts:
conflict_forecasts(VIEWS + conflictforecast.org + ACLED CAST predicted fatalities, risk scores, and event counts) - Structured data:
reliefweb_reports,acled_political_events,acaps_inform_severity,acaps_inform_severity_trend,acaps_risk_radar,acaps_daily_monitoring,acaps_humanitarian_access - Context sources:
enso_state,seasonal_tc_outlooks,seasonal_tc_context_cache,hdx_signals,crisiswatch_entries - Questions + research:
questions,question_research - Forecasts:
forecasts_raw,forecasts_ensemble(both includereasoning_trace_json) - Sibyl:
sibyl_runs(run coverage, cost, budget_capped),sibyl_forecasts(per-question trial traces, pooled quantiles, JS divergences) - Resolutions + scoring:
resolutions,scores,eiv_scores - Calibration:
calibration_weights,calibration_advice,bucket_centroids,bucket_definitions - Diagnostics:
llm_calls,question_run_metrics - Resolver:
facts_resolved,facts_deltas,snapshots,manifests(the last two are written only by the snapshot export CLI, never by the pipeline) - Hazard resolution machine (
resolver/hazard_resolution/):haz_raw_*per-source caches,haz_triggers,haz_impact_candidates,haz_resolutions,haz_revisions,haz_base_rates_occurrence,haz_base_rates_severity— created byhaz-migrate(or any machine entry point); behaviour configured inrulebook.yaml. CLIs:resolve-hazards --hazard {cyclone,flood,drought} --month YYYY-MM [--summary-out PATH],haz-backcast [--time-budget-min N],haz-base-rates,haz-acceptance. Runs in production since Aug 2026 (shadow mode):resolver_update.ymlPhase 2.5 resolves the trailing 3 months each cycle, and the nightlyhaz_backcast.ymlfills history in time-boxed chunks; its answers do not yet feedcompute_resolutions— that flip is gated on the monthly acceptance report in thebackfill-diagnosticsartifact
All model choices are centralized in pythia/config.yaml under llm.models (the model registry: one alias per model family) and llm.profiles.prod (the SPD ensemble list plus a roles block assigning an alias to every other purpose).
llm:
models: # THE single place to swap a model family
gpt: openai:gpt-6-sol
gpt_mini: openai:gpt-6-luna
claude: anthropic:claude-opus-5-5
gemini_pro: google:gemini-3.1-pro-preview
gemini_flash: google:gemini-3.5-flash
gemini_lite: google:gemini-2.5-flash
profile: "prod" # override with PYTHIA_LLM_PROFILE
profiles:
prod:
ensemble: # SPD forecast ensemble — entries reference aliases
- model: gpt
thinking: high
- model: claude
# ... (see config.yaml for the full list and per-model params)
roles: # every non-ensemble purpose -> alias
hs_default: gemini_flash
rc_pass1: gemini_flash
track2_spd: gemini_flash
scenario_writer: gemini_flash
# ... (see config.yaml for all roles)- Upgrade a model family (e.g. GPT-5.6 → GPT-6): change the id on one line in
llm.models, then add a cost entry inpythia/model_costs.json. Every ensemble member and role that references the alias picks up the new model. - Add/remove an ensemble member: add or delete a
- model: <alias>entry (define the alias inllm.modelsfirst). Multiple members from the same provider are supported. - Point a role at a different model: edit the
roles:block. Values are registry aliases or explicitprovider:model_idrefs. - Add a new provider: requires a
call_<provider>()function and dispatch branch inforecaster/providers.py, plus an entry in_PROVIDER_ENV_KEYS.
pythia/tests/test_model_registry.py validates the registry: every alias/role resolves and every reachable model has a cost entry.
Per-model cost rates are stored in pythia/model_costs.json as [input, output] cost per 1,000,000 (1M) tokens in USD — one entry per model id (provider-prefixed lookups are normalized automatically). When switching to a new model id, add its cost entry to this file; models without an entry log $0 cost, and the registry test fails to remind you.
Every role can still be overridden at runtime without touching config: PYTHIA_HS_FALLBACK_MODEL_SPECS, PYTHIA_SCENARIO_MODEL_ID, PYTHIA_TRIAGE_MODEL_PASS1/2, PYTHIA_RC_MODEL_PASS1/2, HS_MODEL_ID, PYTHIA_GROUNDING_MODEL_ID, PYTHIA_WEB_RESEARCH_MODEL_ID. Env vars always win over config roles.
PYTHIA_SPD_ENSEMBLE_SPECS: overrides the entire SPD ensemble at runtime (comma-separatedprovider:model_idpairs).PYTHIA_BLOCK_PROVIDERS: comma-separated provider names to exclude (e.g.google).PYTHIA_SPD_GOOGLE_MODEL_ID: overrides all Google model IDs in the SPD ensemble.
git clone <YOUR_REPO_URL>
cd Pythia
python -m venv .venv
source .venv/bin/activate
pip install -r python_library_requirements.txtexport PYTHIA_DB_URL="duckdb:///data/resolver.duckdb"
export PYTHIA_LLM_PROFILE="test" # test or prodpython - <<'PY'
from pythia.db.schema import ensure_schema
ensure_schema()
print("schema ready")
PYexport PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1 # tail packs are off by default
# Note: The question-level web research pipeline is deprecated.
# SPD prompts receive structured data directly via _load_structured_data().
# Grounding uses OpenAI web search (primary) + Gemini (fallback).- Default list:
horizon_scanner/hs_country_list.txt
PYTHIA_LLM_PROFILE=test python -m horizon_scanner.horizon_scannerPYTHIA_LLM_PROFILE=test python -m forecaster.cli --mode pythia --limit 20 --purpose local_smoke- DuckDB: open the database at
PYTHIA_DB_URL(app.db_urldefault isduckdb:///data/resolver.duckdb). - Debug bundle (optional):
python -m scripts.dump_pythia_debug_bundle \
--db "${PYTHIA_DB_URL}" \
--hs-run-id "<HS_RUN_ID>" \
--forecaster-run-id "<FORECASTER_RUN_ID>"- Bundles default to
debug/pytia_debug_bundle__<run_id>.md(note: filename usespytia_).
- Production pipeline:
run_horizon_scanner.ymlruns HS + forecaster end-to-end. - Resolver Update (backfill):
resolver_update.yml— primary data ingestion (11th monthly, two days ahead of the forecast pipeline on the 13th; until 2026-10-05 the 28th and the 1st). Single-job workflow with 5 sequential phases: (1) Resolver connectors (fatal), (2) Resolution sources, (3) Structured data, (4) Context sources, (5) Verify + export artifact. - Structured data refresh:
ingest-structured-data.yml— mid-cycle refresh for fast-changing sources (weekly Sunday 03:00 UTC): conflict forecasts, GDACS, ReliefWeb, ACLED political. - Post-forecast pipeline:
compute_resolutions.yml,compute_scores.yml,compute_calibration_pythia.yml.compute_resolutionsprints, per hazard and metric, how many horizon-months resolved from a source, defaulted to zero, or stayed open for lag;compute_scoresdrops scores whose resolution was withdrawn. - Canonical DB guards: every workflow that uploads
pythia-resolver-dbruns.github/actions/guard-canonical-uploadfirst and stops red if another run uploaded a newer one after the copy it downloaded; the chained workflows (resolutions, scores, calibration, publish, inspect) skip when their trigger uploaded nothing; and the pipeline gate (scripts/ci/check_pipeline_active.py) answers in flight, idle or unknown, proceeding with a warning on a schedule and stopping on a manual dispatch when it cannot tell. - Maintenance:
compact_resolver_db.ymlcompacts the canonical DB (on the 4th and the 20th, and by hand), and itsconflict_displacement_resetmode exports and deletes every ACE/PA resolution, score and learned row, then starts the resolution chain. - Forecaster CI:
forecaster-ci.ymlcovers SPD unit tests and optional compare artifacts. - ENSO / Seasonal TC refresh: fetched fresh and stored to the DB by
resolver_update.ymlPhase 4 (fetch_and_store_enso()/fetch_and_store_seasonal_tc()). Run manually viapython -m horizon_scanner.enso.enso_module/python -m horizon_scanner.seasonal_tc.seasonal_tc_runner. (The standalonerefresh-enso.yml/refresh-seasonal-tc.ymlworkflows were removed July 2026 — Phase 4 does the real fetch+store.) - CrisisWatch refresh:
refresh-crisiswatch.ymlfetches ICG CrisisWatch data from the Internet Archive's Wayback Machine (Save Page Now capture + CDX snapshot walk — the ICG site's Cloudflare blocks direct CI fetches) on the 3rd/5th/7th/10th of each month, parses with BeautifulSoup, and commits updated JSON tohorizon_scanner/data/crisiswatch_latest.jsononly when a newer edition is available.crisiswatch-ci.ymlruns the refresh-script tests. - SPEI-3 drought feed:
spei3_refresh.ymlproducesresolver/data/spei3_country_means.csvon the 10th of each month — the only drought indicator reaching back before the HDX and NMME ingests began. It is a global raster and zonal statistics do not belong in a resolution run, so the CSV is the product and the rulebook'stabularprovider reads the committed file. Five stages (plan/fetch/reduce/validate/promoteinscripts/build_spei3_country_means.py), plus a dispatch-onlydescribe/probepair that asks the CDS what a request enum accepts and does nothing else. It asks BOTH ERA5-Drought releases, one per month: the consolidated one for the settled end of the record and the intermediate one for the recent months it does not yet hold, which is what lets the feed reach the month the drought backcast needs. Every row records which release it came from, so a provisional intermediate value is re-asked for as soon as its settled version exists rather than waiting on a trailing window; dependencies are pinned in its OWNrequirements-spei3.txtso the pipeline never carriescdsapi. Needs theCDSAPI_KEYsecret. It sits in its own concurrency group, neverpythia-resolver-db, so it cannot contend with the nightly backcast or the monthly ingest — which is why it commits its resume-ledger restale request intoresolver/data/spei3_status.jsonand the nightly backcast applies it. Branch protection refuses a CI push tomain, so the commit needs aSPEI3_COMMIT_TOKENwith bypass — there is no pull-request fallback, because one opened byGITHUB_TOKENnever triggers its own checks and so could never merge. That token carries no expiry, so it is guarded on the REFUSAL: a failed push is classified, the run fails through a step naming the class, and the run issue register reports a refused credential atdegraded. - DuckDB inspection:
inspect_resolver_duckdb.yml— 7 data quality checks on the DB artifact, triggered after Resolver Update. - Release (publish):
publish_latest_data.ymluploads the canonical DB to thepythia-data-latestrelease and then runs the same DB inspection inline. The forecast pipeline releases in the orderHorizon Scanner Triage → Sibyl → publish → inline DB inspect(a single release per cycle): HS Triage completion triggers Sibyl, and Sibyl dispatches publish at the end of its run. The post-scoring chain (Resolver Update → resolutions → scores → calibration) also dispatches publish at the end of calibration; publish has no other trigger. If a Sibyl run breaks, recover the forecast release by manually dispatchingpublish_latest_data.yml.
- Workflow downloads the canonical DB artifact (
pythia-resolver-dbby default), appends new HS + forecaster rows, and uploads the updated DB as a new artifact. - It validates a signature (
scripts/ci/db_signature.py) and writes:diagnostics/db_signature_before.jsondiagnostics/db_signature_after.json
- If signature validation fails, the workflow rejects the candidate DB and tries the next run. A missing/failed signature is a hard-stop; confirm required tables exist or re-run with a reset DB artifact for bootstrap runs.
- Debug bundle:
pythia-debug-bundle—debug/pythia_debug_bundle__<run_id>.zip, written in the Sibyl job (run_sibyl.yml) at the end of the forecast cycle, beside the current-run and forecast attribution bundles (seedocs/bundles.mdfor the four bundles and their join keys). Readanomalies__<run>.jsonfirst: severity, subsystem, one line on what is wrong, and the file holding the evidence.BUNDLE_MANIFEST.jsonnames every other file and what it holds — per-question metrics, SPD tables, every LLM call, the provider batch lifecycle and the raw provider objects, this cycle's Actions logs, the resolved environment and LLM profile, source copies of the files most often implicated in a bad run, connector freshness, prompt-cache and retry reports, per-member forecast completeness, and this run beside the previous two. Build locally with:
python -m scripts.dump_pythia_debug_bundle --db duckdb:///data/resolver.duckdb \
--hs-run-id hs_20260901T040916 --forecaster-run-id fc_1788237725
# PYTHIA_BUNDLE_WORKFLOW_LOGS=0 and PYTHIA_BUNDLE_FETCH_PROVIDER_OBJECTS=0 skip
# the two collectors that need credentials; both degrade cleanly without them.When the zip exceeds 25 MB the workflow logs are split into pythia-debug-bundle-workflow-logs rather than dropped, and the manifest says so.
- Forecast attribution bundle:
pythia-forecast-attribution-bundle—ai_bundle/forecast_attribution__<run_id>.zip, built in the same Sibyl job right after the debug bundle. A signal ledger parsed from every model's reasoning trace (where probability mass moved and what the model said moved it, classed by a versioned regex taxonomy), prior anchoring against the base-rate anchor, RC assessment, trace quality, an input inventory per question, evidence items linked to signals, prompt section fingerprints and token share, model and Sibyl contrasts, run-over-run deltas and a brief per hazard. Its guide opens by saying the ledger is claimed attribution, not measured influence. Build locally with:
python -m scripts.ai_bundle.build_forecast_attribution_bundle --db duckdb:///data/resolver.duckdb \
--out-dir ai_bundle [--run-id fc_1788237725] [--hs-run-id hs_20260901T040916]- Resolver debug bundle:
resolver-debug-bundle-<run_id>— one zip per Resolver Update run, built byif: always()so the runs worth diagnosing are covered. It answers, from the zip alone: which URL each connector called and what came back (http/requests.jsonl,http/envelopes/— including the response fields the connectors themselves discard); what thresholds and windows the run used (config/rulebook.yaml,run/env_effective.json); why any unresolved hazard cell produced no row (hazard/cell_ledger.csv, one reason code per cell); what ceiling a rejected figure was measured against and where that number came from (hazard/figures_ledger.csv); whether the extraction budget was exhausted and by whom (hazard/extraction_budget.csv); which code produced the run (code/,run/git.json); and whether any two sources in the run contradict each other (checks/contradictions.md). Start atREADME.mdinside the zip. Build locally with:
python -m scripts.build_resolver_debug_bundle --db duckdb:///data/resolver.duckdb \
--diagnostics-dir diagnostics --out debug_bundle/resolver-debug-bundle.zip
# Set PYTHIA_RUN_LOG_DIR before the ingest to capture HTTP requests, response
# envelopes and the machine's per-cell / per-figure ledgers; without it the
# bundle reports what the logs say and no more.Hard ceiling of 80 MB compressed. When it binds, code/ is dropped before any evidence is and the manifest names what went; nothing is truncated silently. Every file is scanned for every secret in the build environment before the zip is written, and a hit fails the build.
- Legacy monolithic bundle:
python -m scripts.dump_pythia_debug_bundle --legacystill writes the single markdown file. - SPD compare JSON:
debug/spd_compare_smokeordebug/spd_compare_tests - Diagnostics:
diagnostics/(DB signatures, compare JSON, latency summaries) - HS triage coverage:
diagnostics/hs_triage_coverage__<HS_RUN_ID>.csvanddiagnostics/hs_triage_failures__<HS_RUN_ID>.json - AI analysis bundle:
pythia-ai-analysis-bundle(calibration workflow) — one zip per scoring round joining forecast reasoning to realized outcomes for every scored question; hand it to an AI analyst. Build locally with:
python -m scripts.ai_bundle.build_scored_forecast_bundle --db duckdb:///data/resolver.duckdb --out-dir ai_bundle
# options: --months-back 12, --n-case-studies 10, --include-test,
# --include-sibyl-trials {case-studies|all|none}, --keep-stagingpython scripts/migrate_llm_calls_telemetry.py --db duckdb:///path/to/resolver.duckdb- After the HS run completes, stdout includes
HS_TRIAGE_RERUN_ISO3S=.... To re-run just those countries:
PYTHIA_HS_ONLY_COUNTRIES="<ISO3S>" python -m horizon_scanner.horizon_scannerPYTHIA_DB_URL="duckdb:///data/resolver.duckdb" uvicorn pythia.api.app:app --reload --port 8000Production (Render) runs a single worker — multiple workers each open their own DuckDB connection and caches, multiplying memory on the 2GB instance:
uvicorn pythia.api.app:app --host 0.0.0.0 --port $PORT --workers 1Recommended Render env: PYTHIA_MAX_CONCURRENT_HEAVY=1, MALLOC_ARENA_MAX=2.
The /v1/run endpoint is disabled by default; set PYTHIA_ALLOW_INPROCESS_RUN=1
locally if you want the API to launch pipeline runs in-process.
cd web
npm install
NEXT_PUBLIC_PYTHIA_API_BASE=http://localhost:8000/v1 npm run dev- Forecast Index (Overview): shows the Humanitarian Impact Forecast Index, with RC KPI counts and RC map markers (highest RC level per country).
- Forecasts (
/questions): latest forecasts table includes RC score, triage fields, and track assignment whenlatest_only=true. - HS Triage (
/hs-triage): displays per-run triage with RC likelihood/direction/magnitude/score and tier (quiet/priority). - Question detail (
/questions/[questionId]): RC fields appear alongside research and SPD outputs. - Countries (
/countries): includes highest RC level/score per country from the latest HS run. - Performance (
/performance): forecast evaluation with KPI cards, ensemble selector dropdown (compare ensemble_mean vs. ensemble_bayesmc), median and average scores (Brier, Log, CRPS), and views by Total, Hazard, Run, and Model. - Downloads (
/downloads): forecast/triage exports with RC fields, plus per-question score CSVs, model-level summary CSVs, and rationale exports. - About (
/about): versioned prompt snapshots and system overview history.
- Forecast SPD & EIV export:
/v1/downloads/forecasts.csvand/v1/downloads/forecasts.xlsxinclude RC probability/direction/magnitude/score columns and track assignment per row. - Score exports: per-question CSV for each named ensemble (full 6-month x 5-bin grid, expected impact values, resolutions, and Brier/Log/CRPS scores), model-level summary CSV (avg/median/min/max per model and hazard), and rationale export (human-readable LLM explanations by hazard).
- Countries endpoint:
/v1/countriesincludeshighest_rc_level/highest_rc_score(latest HS run). - Questions endpoint:
/v1/questions?latest_only=trueincludes RC fields (regime_change_*) andtrack.
See PUBLIC_APIS.md for canonical API contracts.
pythia/config.yamlis the authoritative configuration.- There is no
pythia/config.py; all runtime defaults live in the YAML + env vars. app.db_urldefines DuckDB;PYTHIA_DB_URLoverrides at runtime.llm.profileselects the default model bundle; override withPYTHIA_LLM_PROFILE.llm.profiles.<name>.ensembledefines the forecast ensemble; see Model management.pythia/model_costs.jsoncontains per-model cost rates.
- Provider keys:
OPENAI_API_KEY,GEMINI_API_KEY,ANTHROPIC_API_KEY,EXA_API_KEY,PERPLEXITY_API_KEY. - Structured data:
ACAPS_EMAIL,ACAPS_PASSWORD(ACAPS API auth);ACLED_ACCESS_KEY,ACLED_EMAIL(ACLED CAST API auth). - Concurrency:
PYTHIA_LLM_CONCURRENCY(global LLM call cap)HS_MAX_WORKERSFORECASTER_RESEARCH_MAX_WORKERSFORECASTER_SPD_MAX_WORKERS
- Grounding:
PYTHIA_GROUNDING_PRIMARY_BACKEND:openai(default) orgemini— controls which backend is tried first for RC/triage grounding.PYTHIA_ADVERSARIAL_CHECK_ENABLED: Enable adversarial checks for RC L1+ (default1).- Note: The legacy
PYTHIA_WEB_RESEARCH_ENABLED,PYTHIA_RETRIEVER_ENABLED, andPYTHIA_HS_RESEARCH_WEB_SEARCH_ENABLEDare deprecated (set to0in workflow).
- SPD tuning / timeouts (Gemini 3):
PYTHIA_GOOGLE_SPD_THINKING_LEVEL_FLASH,PYTHIA_GOOGLE_SPD_THINKING_LEVEL_PROPYTHIA_GOOGLE_SPD_TIMEOUT_FLASH_SEC,PYTHIA_GOOGLE_SPD_TIMEOUT_PRO_SECPYTHIA_GOOGLE_SPD_RETRIES
- SPD ensemble override:
PYTHIA_SPD_ENSEMBLE_SPECS(e.g.openai:gpt-6-sol,google:gemini-3.5-flash). - HS triage resilience:
PYTHIA_HS_FALLBACK_MODEL_SPECS(defaults tohs_fallbackfrom the active profile, thenopenai:gpt-6-sol; keeps HS triage running when Gemini fails).PYTHIA_HS_ONLY_COUNTRIES(comma-separated ISO3s/names to rerun HS triage for a subset).PYTHIA_PROVIDER_FAILURE_THRESHOLD,PYTHIA_PROVIDER_COOLDOWN_SECONDS,PYTHIA_PROVIDER_RESET_ON_SUCCESSPYTHIA_LLM_RETRY_TIMEOUTS(set0to opt out of timeout retries outside HS triage)PYTHIA_HS_LLM_MAX_ATTEMPTS(defaults to 3 for HS triage retries)PYTHIA_HS_GEMINI_TIMEOUT_SEC(defaults to 120s for HS triage)
| Secret | Used by | Required for |
|---|---|---|
OPENAI_API_KEY |
Forecaster SPD ensemble | OpenAI models in ensemble |
GEMINI_API_KEY |
HS, retriever, forecaster | Gemini models (HS + SPD + retriever) |
ANTHROPIC_API_KEY |
Forecaster SPD ensemble + Sibyl | Anthropic models |
EXA_API_KEY |
Web research | Exa backend (optional) |
PERPLEXITY_API_KEY |
Web research | Perplexity backend (optional) |
ACAPS_EMAIL |
Structured data | ACAPS API auth (INFORM, Risk Radar, etc.) |
ACAPS_PASSWORD |
Structured data | ACAPS API auth |
ACLED_ACCESS_KEY |
Structured data | ACLED CAST API auth |
ACLED_EMAIL |
Structured data | ACLED CAST API auth |
GITHUB_TOKEN |
Actions | Artifact download + summary updates |
Minimum set for full pipeline: GEMINI_API_KEY + at least one of (OPENAI_API_KEY, ANTHROPIC_API_KEY). Missing keys disable their providers, and the run proceeds with a partial ensemble.
- Partial ensembles are expected: providers can timeout; the ensemble uses available members and records
ensemble_meta. - Timeouts cap tail risk: per-provider timeouts are enforced; increase them only if needed.
- Shared retriever: evidence packs are cached and reused to stabilize sources and reduce cost.
- HS triage always writes: HS stores triage outputs even if no questions are produced.
- Prediction market retriever disabled: Metaculus returns 403, Polymarket returns 422. Do not re-enable until upstream APIs are fixed.
- IPC API unavailable: The IPC connector requires
IPC_API_KEYwhich is not available; FEWS NET Phase 3+ data is used instead. - IFRC Montandon sparsity: natural hazard PA data may be sparse for some hazards/countries, reducing base-rate strength.
- DI, CU, HW fully silenced: Displacement Influx, Civil Unrest, and Heatwave hazards are blocked at the hazard catalog level and never enter RC, triage, grounding, or question generation.
- Occasional “200 but ungrounded”: some web search calls can return success without verified sources. OpenAI is the primary grounding backend; Gemini is the fallback.
- No questions generated: HS may mark all tiers quiet. Check
hs_triagein DuckDB and confirm your country list (horizon_scanner/hs_country_list.txt) andhazards_allowedin config. - No hazard tail packs: confirm
PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1and RC Level ≥1 for the hazard. Tail packs are limited to 2 hazards per country. - RC fields missing in API: ensure
hs_triagehasregime_change_*columns (pythia/db/schema.py) and thatlatest_only=trueis set on/v1/questions. - No active models: verify
PYTHIA_LLM_PROFILE,llm.profilesensemble config inpythia/config.yaml, and provider API keys. - Debug bundle too large for step summary: artifacts still exist under
debug/even if GitHub Step Summary truncates. - Slow runs: Gemini tails can dominate latency. Tune
PYTHIA_LLM_CONCURRENCY,FORECASTER_*_MAX_WORKERS, and SPD timeouts. - Interpreting question_run_metrics:
question_run_metrics(if present) records wall-clock vs compute vs queue time per question; seescripts/dump_pythia_debug_bundle.py.
- Non-technical system overview:
docs/fred_overview.md - Config:
pythia/config.yaml - Model costs:
pythia/model_costs.json - Horizon Scanner:
horizon_scanner/horizon_scanner.py - Regime Change docs:
docs/hs_regime_change.md - Hazard Tail Packs docs:
docs/hs_hazard_tail_packs.md - HS country list:
horizon_scanner/hs_country_list.txt - Forecaster CLI:
forecaster/cli.py - Hazard-specific prompts:
forecaster/hazard_prompts.py - Per-hazard RC prompts:
horizon_scanner/rc_prompts.py - Per-hazard triage prompts:
horizon_scanner/hs_triage_prompts.py - RC grounding prompts:
horizon_scanner/rc_grounding_prompts.py - Triage grounding prompts:
horizon_scanner/hs_triage_grounding_prompts.py - ENSO module:
horizon_scanner/enso/ - Seasonal TC module:
horizon_scanner/seasonal_tc/ - HDX Signals:
horizon_scanner/hdx_signals.py - ACLED CAST connector:
resolver/connectors/acled_cast.py - Conflict forecasts loader:
horizon_scanner/conflict_forecasts.py - CrisisWatch:
horizon_scanner/crisiswatch.py,scripts/refresh_crisiswatch.py - Structured data:
pythia/acaps.py,pythia/food_security.py,horizon_scanner/reliefweb.py,pythia/acled_political.py - Adversarial checks:
pythia/adversarial_check.py - Calibration advice:
pythia/tools/generate_calibration_advice.py - Schema:
pythia/db/schema.py - Debug bundle script:
scripts/dump_pythia_debug_bundle.py - Workflows:
run_horizon_scanner.yml,forecaster-ci.yml,refresh-crisiswatch.yml(ENSO/Seasonal TC are fetched+stored byresolver_update.ymlPhase 4) - Public API contracts:
PUBLIC_APIS.md
See CONTRIBUTING.md for coding standards and workflow notes. Archived READMEs live in docs/archive/README_INDEX.md.
This repository follows the licensing terms bundled with the codebase (see LICENSE if present or repository metadata).