Skip to content
kwyjadPublic

About

An experimental humanitarian impact forecasting system, combining parallel structured data injected ensemble forecasts and an agentic deep research forecasting approach. Known as Fred on public interface side

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2,004 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Pythia

Pythia is an end-to-end AI forecasting system for humanitarian questions. It scans countries for emerging hazards, turns triage into forecastable questions, runs an LLM ensemble to produce Subjective Probability Distributions (SPD), and stores forecasts, research, and diagnostics in a DuckDB-backed system of record.

Table of contents

What is Pythia?

Pythia connects historical hazard data, web research, and LLM reasoning into a repeatable pipeline for early-warning forecasts. It produces triage signals, structured research briefs, and probabilistic forecasts across hazards, then serves those outputs through a DuckDB system of record, a FastAPI service, and a Next.js dashboard.

System at a glance

Resolver facts/base rates → Horizon Scanner per-hazard pipeline:
  Per hazard: RC grounding → RC assessment → triage grounding → triage (single pass each by default)
  + Structured data: ReliefWeb, ACAPS, IPC, ACLED political, ENSO, seasonal TC, HDX Signals, GDELT
  + Conflict forecasts: VIEWS, conflictforecast.org, ACLED CAST
  + Adversarial checks (RC L1+) → Calibration advice
  → Track 1: full ensemble SPD + scenarios
  → Track 2: single-model SPD
  → DuckDB → API → Dashboard + downloads
  • System of record: DuckDB (PYTHIA_DB_URL / app.db_url) holds HS triage, questions, structured data, forecasts, and diagnostics.
  • End-to-end flow: Resolver facts and base rates inform the HS per-hazard pipeline. Each hazard gets its own RC assessment (with dedicated grounding) and triage assessment (with separate grounding), each a single Gemini Flash pass by default (PYTHIA_HS_RC_PASSES/PYTHIA_HS_TRIAGE_PASSES=2 restore the 2-pass merge). Structured data connectors pull evidence from humanitarian, climate, and conflict-forecast APIs including ENSO state, seasonal TC forecasts, HDX Signals, and ACLED CAST. Adversarial checks run for RC L2+ cases. Forecaster routes questions by track — Track 1 (RC-elevated) through the full ensemble, Track 2 (priority, no RC) through a single model — and writes results back to DuckDB.

Core components

  • Resolver (facts + base rates): Resolver tables in DuckDB (facts_resolved, facts_deltas, snapshots) provide historical context. Schema is defined in pythia/db/schema.py.
  • Horizon Scanner (HS): python -m horizon_scanner.horizon_scanner triages countries and hazards, writes hs_runs/hs_triage, seeds questions/question_research, and generates country evidence packs. HS runs a fully per-hazard pipeline: each hazard (ACE, DR, FL, HW, TC) gets its own RC grounding call, RC assessment, triage grounding call, and triage assessment. RC and triage each run a single Gemini Flash pass by default (a 2-pass merge is available via PYTHIA_HS_RC_PASSES/PYTHIA_HS_TRIAGE_PASSES=2, useful only with a distinct pass-2 model). RC grounding uses TRIGGER/DAMPENER/BASELINE signal categories; triage grounding uses SITUATION/RESPONSE/FORECAST/VULNERABILITY categories with longer recency windows. An ACLED low-activity filter skips ACE triage for countries with minimal recent conflict (0 fatalities in 2 months AND <25 in 12 months).
  • Retriever web research (shared evidence packs): When enabled, the retriever builds evidence packs reused across HS, research, and SPD prompts. The shared retriever defaults to gemini-3.5-flash when PYTHIA_RETRIEVER_ENABLED=1 and PYTHIA_RETRIEVER_MODEL_ID is unset; see pythia/web_research/web_research.py.
  • Structured data connectors: Authoritative humanitarian, climate, and conflict-forecast data pulled from specialist APIs and stored in DuckDB for direct prompt injection. Replaced the former Research LLM stage. See pythia/acaps.py, pythia/fewsnet_food_security.py, horizon_scanner/reliefweb.py, pythia/acled_political.py, horizon_scanner/enso/, horizon_scanner/seasonal_tc/, horizon_scanner/hdx_signals.py, resolver/connectors/acled_cast.py.
  • Adversarial evidence checks: Counter-evidence searches for RC Level 1+ cases, stored in hs_adversarial_checks. See pythia/adversarial_check.py.
  • Hazard-specific reasoning: Per-hazard forecasting instructions injected into SPD prompts. See forecaster/hazard_prompts.py.
  • Calibration advice: Per-hazard/metric calibration guidance generated from historical performance, injected into forecasting prompts. See pythia/tools/generate_calibration_advice.py.
  • Forecaster SPD v2 ensemble: python -m forecaster.cli --mode pythia runs SPD v2 prompts across the active ensemble and writes per-model outputs (forecasts_raw) and aggregated results (forecasts_ensemble). Questions are routed by track: Track 1 (RC-elevated, level > 0) uses the full multi-model ensemble; Track 2 (priority, no RC) uses a lightweight single-model path (Gemini Flash).
  • Sibyl (parallel deep-research track): python -m sibyl.run independently re-forecasts 25 affected/fatalities questions per run, chosen floor-then-fill (each hazard's three most volatile first, then the most volatile remaining under a cap of ten per hazard), using agentic open-web research (Claude Opus + Brave Search). K independent belief-updating trials per question are linear-pooled into an SPD written beside the standard track (model_name='sibyl' in forecasts_raw/forecasts_ensemble), so scoring grades the two tracks head-to-head. Takes two inputs from Fred's own data: the Resolver base rate, and one reading of the resolving source per question shown to every trial (a live month-to-date ACLED read for conflict deaths when SIBYL_LIVE_LOOKUPS_ENABLED=1, the newest Phase 3+ rows for drought, the last six months of resolving rows and GDACS alerts for floods and cyclones; sibyl/resolver_reading.py). A hashed half of questions (SIBYL_PACK_SHARE) is also shown a starting pack of the pipeline's structured feeds, never the ensemble's own research, as a measured experiment (sibyl/pack.py). Hard run budget cap (default $60) and a wall-clock cap on the trial loop (default 180 minutes). A monthly advice loop (python -m sibyl.advice, run by the calibration workflow) measures Sibyl's own resolved record per class and, once a class holds 20 scored questions, shows Sibyl a short YOUR TRACK RECORD note on the next run, to half its questions. See sibyl/ and sibyl/DISCOVERY.md.
  • Scenarios (Track 1 only): When an ensemble SPD is available, scenarios can be generated for Track 1 questions and written back to DuckDB.
  • Batch-API staged pipeline (opt-in): The HS and forecaster LLM families can run through provider Batch APIs at 50% off both input and output tokens. python -m horizon_scanner.horizon_scanner --stage {hs_submit,hs_rc_collect,hs_finalize} and python -m forecaster.cli --phase {submit,collect} split each family into a submit wave (prompts built + enqueued in pythia/llm_batch.py, byte-identical bodies to the sync path) and a collect wave (results replayed, per-item sync fallback on any miss). GitHub orchestration: pythia_pipeline_stage.yml (staged runs — the production monthly entry since 2026-07-27, cron on the 13th entering at hs_submit) + poll_llm_batches.yml (15-min poller that advances the pipeline when batches finish). Code defaults are OFF (PYTHIA_BATCH_API_ENABLED=0) so the --stage full / --phase full paths and local runs are unchanged; the production workflows opt in. run_horizon_scanner.yml is the manual synchronous fallback. Related: static-first prompt ordering + provider prompt caching (PYTHIA_PROMPT_V3_ORDER / PYTHIA_PROMPT_CACHE_ENABLED, on in the production workflows; scripts/compare_prompt_order_runs.py is the regression harness). Anthropic cache_control inside batch bodies is a separate opt-in (PYTHIA_BATCH_PROMPT_CACHE, on in pythia_pipeline_stage.yml since 2026-08-03, with a 1h TTL because batches routinely outlive the 5-minute default) — confirm it is paying for itself by checking cache_read_input_tokens in the batch-economics / post-run diagnostics after each cycle.
  • Dashboard + API: FastAPI serves /v1/* endpoints for the dashboard, downloads, and diagnostics. The Next.js UI lives in web/.

Regime Change (RC): out-of-pattern detection

Definition: Regime Change (RC) is the likelihood of a generating-process shift within the next 1–6 months that makes historical base rates less reliable. RC is assessed per-hazard in a dedicated LLM pipeline step that runs before triage. Each hazard type gets its own RC prompt with calibration anchors (default likelihood 0.05, expected ~80% at ≤0.10), dedicated grounding call, and hazard-specific data injection (ENSO context for climate hazards, conflict forecasts for ACE, seasonal TC for TC). RC runs a single Gemini Flash pass by default (2-pass available via PYTHIA_HS_RC_PASSES=2) and writes to hs_triage.

Stored fields in hs_triage:

  • regime_change_likelihood (probability)
  • regime_change_direction (up, down, mixed, unclear)
  • regime_change_magnitude
  • regime_change_score (computed)
  • regime_change_level (severity)
  • regime_change_window
  • regime_change_json (full normalized object)

Scoring + levels (defaults; env-overridable via PYTHIA_HS_RC_LEVEL*_LIKELIHOOD):

  • score = likelihood × magnitude
  • Level 0: likelihood < 0.15
  • Level 1: likelihood ≥ 0.15
  • Level 2: likelihood ≥ 0.35
  • Level 3: likelihood ≥ 0.55

Behavioral implications:

  • Track routing: Questions with RC level > 0 are assigned Track 1 (full multi-model ensemble); priority questions without RC are Track 2 (lightweight single-model).
  • need_full_spd override: HS forces need_full_spd = TRUE when RC is elevated (Level ≥2) even if the triage tier is quiet.
  • Research v2: RC is surfaced in the research prompt, and RC-elevated hazards require at least one regime_shift_signals entry (or a rebuttal).
  • SPD v2: RC guidance is embedded in the SPD prompt, and the model must include a sentence starting with RC: in human_explanation describing how RC affected the SPD.

See docs/hs_regime_change.md for full details and env overrides.

Hazard Tail Packs (HTP): hazard-specific trigger evidence

Hazard Tail Packs are hazard-scoped evidence bundles generated by HS when RC is elevated. They provide targeted trigger/counter-trigger evidence for downstream Research and SPD.

Trigger rules:

  • Generated only for RC Level ≥1 hazards.
  • Limited to max 2 hazards per country per HS run (PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_PER_COUNTRY, default 2).
  • Tail packs are off by default unless PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1.

Storage table: hs_hazard_tail_packs (one row per hs_run_id × iso3 × hazard_code), including:

  • rc_level, rc_score, rc_direction, rc_window
  • query, report_markdown, sources_json, grounded, grounding_debug_json
  • structural_context, recent_signals_json, created_at

Output format:

  • recent_signals bullets are emitted as TRIGGER | <window> | <direction> | <signal>, DAMPENER | ..., BASELINE | ....

Caching behavior:

  • Tail packs are reused per (hs_run_id, iso3, hazard_code); HS skips regeneration if the row already exists.

Phase A / Phase B behavior:

  • Phase A (Research): When present, the tail pack is injected into Research v2 evidence (hs_hazard_tail_pack).
  • Phase B (SPD v2): For RC-elevated hazards, the tail pack is injected into the SPD prompt. Signals are capped at 12 bullets (PYTHIA_SPD_TAIL_PACKS_MAX_SIGNALS, default 12).

Relevant env vars (defaults):

  • PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED (0)
  • PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_PER_COUNTRY (2)
  • PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_SOURCES (defaults to PYTHIA_HS_EVIDENCE_MAX_SOURCES)
  • PYTHIA_HS_HAZARD_TAIL_PACKS_MAX_SIGNALS (defaults to PYTHIA_HS_EVIDENCE_MAX_SIGNALS)
  • PYTHIA_SPD_TAIL_PACKS_MAX_SIGNALS (12)

See docs/hs_hazard_tail_packs.md for full details.

Retriever web research (deprecated)

The question-level web research pipeline (fetch_evidence_pack / _build_question_evidence_queries) is deprecated. SPD prompts now receive structured data directly via _load_structured_data(), including conflict forecasts, ReliefWeb, HDX Signals, ACAPS, ACLED political, NMME, ENSO, seasonal TC, GDACS event history, and adversarial checks. The question_research table is no longer populated by the pipeline (only placeholder rows). Env vars PYTHIA_RETRIEVER_ENABLED, PYTHIA_WEB_RESEARCH_ENABLED, PYTHIA_SPD_WEB_SEARCH_ENABLED, PYTHIA_FORECASTER_SELF_SEARCH are set to "0" in the workflow.

Grounding for RC and triage uses OpenAI GPT-4.1-mini web search (primary) with Gemini Google Search as fallback. Override with PYTHIA_GROUNDING_PRIMARY_BACKEND=gemini.

NMME seasonal climate forecasts

Pythia ingests NMME (North American Multi-Model Ensemble) seasonal temperature and precipitation anomaly forecasts from the CPC FTP server. These provide structured climate context for drought, flood, and tropical cyclone assessments.

  • Source: ftp://ftp.cpc.ncep.noaa.gov/NMME/realtime_anom/ENSMEAN/ — ensemble mean anomalies at 1° resolution, updated ~9th of each month with 7 lead months.
  • Probability of a below-normal month (since Oct 2026): ftp://ftp.cpc.ncep.noaa.gov/NMME/prob/netcdf/prate.YYYYMM.prob.adj.mon.nc, CPC's calibrated tercile probabilities, stored as variable prate_prob_below (0..1). The PA machine's drought gate reads it at 0.5; cells CPC does not forecast are masked, not read as zero. python -m scripts.ci.nmme_prob_crossing 12 reports per-country crossing frequency for candidate thresholds (needs FTP access to CPC).
  • Processing: Country-level area-weighted averages using xarray + regionmask (Natural Earth admin-0 boundaries). Anomalies expressed in σ (standard deviations from climatology) with derived tercile categories (above/below/near normal).
  • Storage: seasonal_forecasts table in Pythia DuckDB (~2,700 rows per monthly update: 195 countries × 2 variables × 7 leads).
  • Injection: Automatically loaded into HS triage and RC prompts via the existing climate_data parameter for DR, FL, TC hazards. Also injected into forecaster research and SPD prompts via research_json.
  • Run manually: python -m resolver.tools.ingest_nmme (or --year-month YYYYMM for a specific month, --dry-run to preview).
  • Automation: NMME runs as the nmme source inside resolver_update.yml Phase 3 (11th monthly) and ingest-structured-data.yml. (The standalone ingest-nmme.yml workflow was removed July 2026.)

Conflict forecasts

Pythia ingests external conflict forecasts from three independent quantitative sources and incorporates qualitative expert assessments from a fourth, providing forward-looking signals for the Armed Conflict (ACE) hazard.

Quantitative sources (stored in DuckDB)

  • VIEWS (Uppsala/PRIO): ML-based country-month state-based conflict forecasts from the VIEWS Forecasting project. Provides predicted fatalities (views_predicted_fatalities) and probability of ≥25 battle-related deaths (views_p_gte25_brd) at 1–6 month lead times. Good at trends and baseline levels; weaker at sudden onset.
  • conflictforecast.org (Mueller/Rauh): News-based armed conflict risk scores from the Conflict Forecast project. Provides armed conflict risk at 3-month and 12-month horizons (cf_armed_conflict_risk_3m, cf_armed_conflict_risk_12m) and violence intensity outlook (cf_violence_intensity_3m). Better at detecting shifts and escalation signals from media coverage patterns.
  • ACLED CAST (Conflict Alert System Tool): Event-count forecasts from ACLED's CAST API, disaggregated by event type: total events (cast_total_events), battles (cast_battles_events), explosions/remote violence (cast_erv_events), and violence against civilians (cast_vac_events) at 6-month lead. CAST forecasts event counts (not fatalities) and the event-type breakdown reveals shifts in the character of violence that aggregate measures would miss. See resolver/connectors/acled_cast.py.

Qualitative source (prompt-time web research)

  • ICG CrisisWatch: International Crisis Group's monthly CrisisWatch bulletin. Per-country directional assessments (Deteriorated/Improved/Unchanged) and forward-looking "On the Horizon" flags (~3 conflict risks + ~1 resolution opportunity per month). Primary data source: scripts/refresh_crisiswatch.py runs monthly in CI via refresh-crisiswatch.yml, fetching the server-rendered page from the Internet Archive's Wayback Machine (--source wayback — the ICG site's Cloudflare blocks direct CI fetches) and parsing with BeautifulSoup; a local Playwright path (--source live) is kept for manual refreshes. Secondary: Gemini grounding calls during HS runs. Data injected into RC and triage evidence for ACE hazards.

Storage and injection

  • Table: conflict_forecasts in Pythia DuckDB (keyed on source, iso3, hazard_code, metric, lead_months, forecast_issue_date).
  • HS integration: Automatically loaded into RC and triage prompts for ACE hazards via horizon_scanner/conflict_forecasts.py. Includes staleness warnings for data >45 days old. CAST section notes that it measures event counts rather than fatalities.
  • Forecaster integration: Injected into research prompts for ACE questions as structured quantitative anchors.
  • Run manually: python -m resolver.tools.fetch_conflict_forecasts (or --sources views conflictforecast_org acled_cast, --dry-run to preview).

Structured data connectors

Pythia pulls structured humanitarian, climate, and conflict-forecast data from authoritative APIs, stores it in DuckDB, and injects it into pipeline prompts (HS triage, RC assessment, and/or forecaster SPD). This replaced the former Research LLM stage with deterministic, reproducible evidence injection.

Data sources

Source Module What it provides DuckDB table(s)
ReliefWeb horizon_scanner/reliefweb.py Humanitarian situation reports (45-day window, up to 15/country) reliefweb_reports
ACAPS INFORM Severity pythia/acaps.py Crisis severity scores + trend acaps_inform_severity, acaps_inform_severity_trend
ACAPS Risk Radar pythia/acaps.py Forward-looking risk with triggers acaps_risk_radar
ACAPS Daily Monitoring pythia/acaps.py Analyst-curated daily updates acaps_daily_monitoring
ACAPS Humanitarian Access pythia/acaps.py Access constraint scores (HS triage only) acaps_humanitarian_access
ACLED Political Events pythia/acled_political.py Event-level political data (ACE/DI hazards only) acled_political_events
NMME Seasonal Forecasts resolver/tools/ingest_nmme.py Temp/precip anomalies, 7-month lead (DR/FL/TC) seasonal_forecasts
ENSO State/Forecast horizon_scanner/enso/ ENSO phase + strength COMPUTED from the ONI (NOAA ERDDAP → CPC weekly → CPC ONI table), plus the IRI Quick Look's 9-season outlook, plume and IOD (DR/FL/TC) enso_state (DB-first; a run that resolves no index carries the last good record forward, labelled stale)
Seasonal TC Forecasts horizon_scanner/seasonal_tc/ Basin-level TC activity forecasts from TSR, NOAA CPC, BoM, Météo-France La Réunion (SWI), and an IMD/NIO climatology block across 8 basins (TC only) seasonal_tc_outlooks, seasonal_tc_context_cache (DB-first)
HDX Signals horizon_scanner/hdx_signals.py OCHA automated crisis alerts: conflict, displacement, food insecurity, agricultural stress (all hazards) hdx_signals (DB-first, falls back to CSV)
VIEWS resolver/connectors/views.py ML-based conflict fatality predictions (ACE, 1–6 month leads) conflict_forecasts
conflictforecast.org resolver/connectors/conflictforecast.py News-based conflict risk scores (ACE, 3m/12m) conflict_forecasts
ACLED CAST resolver/connectors/acled_cast.py Event-count forecasts by type: total/battles/ERV/VAC (ACE, 6-month lead) conflict_forecasts
FEWS NET IPC resolver/connectors/fewsnet_ipc.py Phase 3+ population estimates (DR hazard; Current Situation + Most Likely). The stored value is the LOWER bound of FEWS NET's published range (e.g. 1,000,000 for "1.0 - 2.49 million") and value_high its upper bound, which the Phase 3+ prompts print beside it; probe_fewsnet_values.yml re-checks this facts_resolved (via Resolver pipeline)
IDMC conflict displacement resolver/ingestion/idmc_conflict.py Conflict displacement per country and month, the ACE/PA resolution series: IDMC's recommended figures only (triangulations dropped), records spanning more than 31 days and months above the population held out. A month resolves 90 days after it ends; a missing month is zero only for a country reported in 8 of the 12 months before and bracketed by a later report (base_rate_spd.resolve_conflict_month); probe_idmc_conflict.yml re-measures the lag facts_resolved (hazard ACE, metrics new_displacements / new_displacements_held)
GDACS resolver/connectors/gdacs.py Disaster population exposure + event occurrence (FL/DR/TC) facts_resolved (via Resolver pipeline)
ICG CrisisWatch horizon_scanner/crisiswatch.py + scripts/refresh_crisiswatch.py Expert conflict arrows + "On the Horizon" flags (ACE RC + triage + SPD). Wayback Machine scraper (monthly, in CI) + Gemini grounding (runtime fallback) crisiswatch_entries
GDELT pythia/gdelt.py Media-derived conflict intensity indicators from GDELT 1.0 daily event exports (CAMEO-tiered, Goldstein, tone; ACE only) gdelt_conflict_indicators

Adversarial evidence checks

For RC Level 1+ cases, pythia/adversarial_check.py runs counter-evidence web searches and synthesizes results into structured output (counter-evidence, historical analogs, stabilizing factors, net assessment). Stored in hs_adversarial_checks.

Calibration advice

pythia/tools/generate_calibration_advice.py generates per-hazard/metric calibration guidance from historical scores. Stored in calibration_advice and injected into forecasting prompts.

Env vars

  • ACAPS_EMAIL, ACAPS_PASSWORD: ACAPS API credentials (required for ACAPS feeds).
  • PYTHIA_DB_URL: DuckDB path (shared with all connectors).

Data model / DuckDB tables

Key tables (see pythia/db/schema.py and SCHEMAS.md for canonical fields):

  • HS: hs_runs, hs_triage (includes RC columns), hs_country_reports, hs_hazard_tail_packs, hs_adversarial_checks
  • Seasonal climate: seasonal_forecasts (NMME country-level temp/precip anomalies)
  • Conflict forecasts: conflict_forecasts (VIEWS + conflictforecast.org + ACLED CAST predicted fatalities, risk scores, and event counts)
  • Structured data: reliefweb_reports, acled_political_events, acaps_inform_severity, acaps_inform_severity_trend, acaps_risk_radar, acaps_daily_monitoring, acaps_humanitarian_access
  • Context sources: enso_state, seasonal_tc_outlooks, seasonal_tc_context_cache, hdx_signals, crisiswatch_entries
  • Questions + research: questions, question_research
  • Forecasts: forecasts_raw, forecasts_ensemble (both include reasoning_trace_json)
  • Sibyl: sibyl_runs (run coverage, cost, budget_capped), sibyl_forecasts (per-question trial traces, pooled quantiles, JS divergences)
  • Resolutions + scoring: resolutions, scores, eiv_scores
  • Calibration: calibration_weights, calibration_advice, bucket_centroids, bucket_definitions
  • Diagnostics: llm_calls, question_run_metrics
  • Resolver: facts_resolved, facts_deltas, snapshots, manifests (the last two are written only by the snapshot export CLI, never by the pipeline)
  • Hazard resolution machine (resolver/hazard_resolution/): haz_raw_* per-source caches, haz_triggers, haz_impact_candidates, haz_resolutions, haz_revisions, haz_base_rates_occurrence, haz_base_rates_severity — created by haz-migrate (or any machine entry point); behaviour configured in rulebook.yaml. CLIs: resolve-hazards --hazard {cyclone,flood,drought} --month YYYY-MM [--summary-out PATH], haz-backcast [--time-budget-min N], haz-base-rates, haz-acceptance. Runs in production since Aug 2026 (shadow mode): resolver_update.yml Phase 2.5 resolves the trailing 3 months each cycle, and the nightly haz_backcast.yml fills history in time-boxed chunks; its answers do not yet feed compute_resolutions — that flip is gated on the monthly acceptance report in the backfill-diagnostics artifact

Model management

All model choices are centralized in pythia/config.yaml under llm.models (the model registry: one alias per model family) and llm.profiles.prod (the SPD ensemble list plus a roles block assigning an alias to every other purpose).

llm:
  models:                # THE single place to swap a model family
    gpt:          openai:gpt-6-sol
    gpt_mini:     openai:gpt-6-luna
    claude:       anthropic:claude-opus-5-5
    gemini_pro:   google:gemini-3.1-pro-preview
    gemini_flash: google:gemini-3.5-flash
    gemini_lite:  google:gemini-2.5-flash

  profile: "prod"        # override with PYTHIA_LLM_PROFILE

  profiles:
    prod:
      ensemble:          # SPD forecast ensemble — entries reference aliases
        - model: gpt
          thinking: high
        - model: claude
        # ... (see config.yaml for the full list and per-model params)
      roles:             # every non-ensemble purpose -> alias
        hs_default:      gemini_flash
        rc_pass1:        gemini_flash
        track2_spd:      gemini_flash
        scenario_writer: gemini_flash
        # ... (see config.yaml for all roles)

Changing models

  • Upgrade a model family (e.g. GPT-5.6 → GPT-6): change the id on one line in llm.models, then add a cost entry in pythia/model_costs.json. Every ensemble member and role that references the alias picks up the new model.
  • Add/remove an ensemble member: add or delete a - model: <alias> entry (define the alias in llm.models first). Multiple members from the same provider are supported.
  • Point a role at a different model: edit the roles: block. Values are registry aliases or explicit provider:model_id refs.
  • Add a new provider: requires a call_<provider>() function and dispatch branch in forecaster/providers.py, plus an entry in _PROVIDER_ENV_KEYS.

pythia/tests/test_model_registry.py validates the registry: every alias/role resolves and every reachable model has a cost entry.

Model costs

Per-model cost rates are stored in pythia/model_costs.json as [input, output] cost per 1,000,000 (1M) tokens in USD — one entry per model id (provider-prefixed lookups are normalized automatically). When switching to a new model id, add its cost entry to this file; models without an entry log $0 cost, and the registry test fails to remind you.

Purpose-specific env overrides

Every role can still be overridden at runtime without touching config: PYTHIA_HS_FALLBACK_MODEL_SPECS, PYTHIA_SCENARIO_MODEL_ID, PYTHIA_TRIAGE_MODEL_PASS1/2, PYTHIA_RC_MODEL_PASS1/2, HS_MODEL_ID, PYTHIA_GROUNDING_MODEL_ID, PYTHIA_WEB_RESEARCH_MODEL_ID. Env vars always win over config roles.

Env var overrides

  • PYTHIA_SPD_ENSEMBLE_SPECS: overrides the entire SPD ensemble at runtime (comma-separated provider:model_id pairs).
  • PYTHIA_BLOCK_PROVIDERS: comma-separated provider names to exclude (e.g. google).
  • PYTHIA_SPD_GOOGLE_MODEL_ID: overrides all Google model IDs in the SPD ensemble.

Quickstart (local)

1) Clone + install

git clone <YOUR_REPO_URL>
cd Pythia
python -m venv .venv
source .venv/bin/activate
pip install -r python_library_requirements.txt

2) Configure the DB and profile

export PYTHIA_DB_URL="duckdb:///data/resolver.duckdb"
export PYTHIA_LLM_PROFILE="test"  # test or prod

3) Ensure the schema exists

python - <<'PY'
from pythia.db.schema import ensure_schema
ensure_schema()
print("schema ready")
PY

4) (Optional) Enable tail packs

export PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1  # tail packs are off by default
# Note: The question-level web research pipeline is deprecated.
# SPD prompts receive structured data directly via _load_structured_data().
# Grounding uses OpenAI web search (primary) + Gemini (fallback).

5) Run Horizon Scanner

PYTHIA_LLM_PROFILE=test python -m horizon_scanner.horizon_scanner

6) Run the forecaster (bounded)

PYTHIA_LLM_PROFILE=test python -m forecaster.cli --mode pythia --limit 20 --purpose local_smoke

7) Inspect outputs

  • DuckDB: open the database at PYTHIA_DB_URL (app.db_url default is duckdb:///data/resolver.duckdb).
  • Debug bundle (optional):
python -m scripts.dump_pythia_debug_bundle \
  --db "${PYTHIA_DB_URL}" \
  --hs-run-id "<HS_RUN_ID>" \
  --forecaster-run-id "<FORECASTER_RUN_ID>"
  • Bundles default to debug/pytia_debug_bundle__<run_id>.md (note: filename uses pytia_).

Running in GitHub Actions

Workflows

  • Production pipeline: run_horizon_scanner.yml runs HS + forecaster end-to-end.
  • Resolver Update (backfill): resolver_update.yml — primary data ingestion (11th monthly, two days ahead of the forecast pipeline on the 13th; until 2026-10-05 the 28th and the 1st). Single-job workflow with 5 sequential phases: (1) Resolver connectors (fatal), (2) Resolution sources, (3) Structured data, (4) Context sources, (5) Verify + export artifact.
  • Structured data refresh: ingest-structured-data.yml — mid-cycle refresh for fast-changing sources (weekly Sunday 03:00 UTC): conflict forecasts, GDACS, ReliefWeb, ACLED political.
  • Post-forecast pipeline: compute_resolutions.yml, compute_scores.yml, compute_calibration_pythia.yml. compute_resolutions prints, per hazard and metric, how many horizon-months resolved from a source, defaulted to zero, or stayed open for lag; compute_scores drops scores whose resolution was withdrawn.
  • Canonical DB guards: every workflow that uploads pythia-resolver-db runs .github/actions/guard-canonical-upload first and stops red if another run uploaded a newer one after the copy it downloaded; the chained workflows (resolutions, scores, calibration, publish, inspect) skip when their trigger uploaded nothing; and the pipeline gate (scripts/ci/check_pipeline_active.py) answers in flight, idle or unknown, proceeding with a warning on a schedule and stopping on a manual dispatch when it cannot tell.
  • Maintenance: compact_resolver_db.yml compacts the canonical DB (on the 4th and the 20th, and by hand), and its conflict_displacement_reset mode exports and deletes every ACE/PA resolution, score and learned row, then starts the resolution chain.
  • Forecaster CI: forecaster-ci.yml covers SPD unit tests and optional compare artifacts.
  • ENSO / Seasonal TC refresh: fetched fresh and stored to the DB by resolver_update.yml Phase 4 (fetch_and_store_enso() / fetch_and_store_seasonal_tc()). Run manually via python -m horizon_scanner.enso.enso_module / python -m horizon_scanner.seasonal_tc.seasonal_tc_runner. (The standalone refresh-enso.yml / refresh-seasonal-tc.yml workflows were removed July 2026 — Phase 4 does the real fetch+store.)
  • CrisisWatch refresh: refresh-crisiswatch.yml fetches ICG CrisisWatch data from the Internet Archive's Wayback Machine (Save Page Now capture + CDX snapshot walk — the ICG site's Cloudflare blocks direct CI fetches) on the 3rd/5th/7th/10th of each month, parses with BeautifulSoup, and commits updated JSON to horizon_scanner/data/crisiswatch_latest.json only when a newer edition is available. crisiswatch-ci.yml runs the refresh-script tests.
  • SPEI-3 drought feed: spei3_refresh.yml produces resolver/data/spei3_country_means.csv on the 10th of each month — the only drought indicator reaching back before the HDX and NMME ingests began. It is a global raster and zonal statistics do not belong in a resolution run, so the CSV is the product and the rulebook's tabular provider reads the committed file. Five stages (plan / fetch / reduce / validate / promote in scripts/build_spei3_country_means.py), plus a dispatch-only describe/probe pair that asks the CDS what a request enum accepts and does nothing else. It asks BOTH ERA5-Drought releases, one per month: the consolidated one for the settled end of the record and the intermediate one for the recent months it does not yet hold, which is what lets the feed reach the month the drought backcast needs. Every row records which release it came from, so a provisional intermediate value is re-asked for as soon as its settled version exists rather than waiting on a trailing window; dependencies are pinned in its OWN requirements-spei3.txt so the pipeline never carries cdsapi. Needs the CDSAPI_KEY secret. It sits in its own concurrency group, never pythia-resolver-db, so it cannot contend with the nightly backcast or the monthly ingest — which is why it commits its resume-ledger restale request into resolver/data/spei3_status.json and the nightly backcast applies it. Branch protection refuses a CI push to main, so the commit needs a SPEI3_COMMIT_TOKEN with bypass — there is no pull-request fallback, because one opened by GITHUB_TOKEN never triggers its own checks and so could never merge. That token carries no expiry, so it is guarded on the REFUSAL: a failed push is classified, the run fails through a step naming the class, and the run issue register reports a refused credential at degraded.
  • DuckDB inspection: inspect_resolver_duckdb.yml — 7 data quality checks on the DB artifact, triggered after Resolver Update.
  • Release (publish): publish_latest_data.yml uploads the canonical DB to the pythia-data-latest release and then runs the same DB inspection inline. The forecast pipeline releases in the order Horizon Scanner Triage → Sibyl → publish → inline DB inspect (a single release per cycle): HS Triage completion triggers Sibyl, and Sibyl dispatches publish at the end of its run. The post-scoring chain (Resolver Update → resolutions → scores → calibration) also dispatches publish at the end of calibration; publish has no other trigger. If a Sibyl run breaks, recover the forecast release by manually dispatching publish_latest_data.yml.

Canonical DB artifacts + signature guardrails

  • Workflow downloads the canonical DB artifact (pythia-resolver-db by default), appends new HS + forecaster rows, and uploads the updated DB as a new artifact.
  • It validates a signature (scripts/ci/db_signature.py) and writes:
    • diagnostics/db_signature_before.json
    • diagnostics/db_signature_after.json
  • If signature validation fails, the workflow rejects the candidate DB and tries the next run. A missing/failed signature is a hard-stop; confirm required tables exist or re-run with a reset DB artifact for bootstrap runs.

Artifact outputs

  • Debug bundle: pythia-debug-bundle — debug/pythia_debug_bundle__<run_id>.zip, written in the Sibyl job (run_sibyl.yml) at the end of the forecast cycle, beside the current-run and forecast attribution bundles (see docs/bundles.md for the four bundles and their join keys). Read anomalies__<run>.json first: severity, subsystem, one line on what is wrong, and the file holding the evidence. BUNDLE_MANIFEST.json names every other file and what it holds — per-question metrics, SPD tables, every LLM call, the provider batch lifecycle and the raw provider objects, this cycle's Actions logs, the resolved environment and LLM profile, source copies of the files most often implicated in a bad run, connector freshness, prompt-cache and retry reports, per-member forecast completeness, and this run beside the previous two. Build locally with:
python -m scripts.dump_pythia_debug_bundle --db duckdb:///data/resolver.duckdb \
  --hs-run-id hs_20260901T040916 --forecaster-run-id fc_1788237725
# PYTHIA_BUNDLE_WORKFLOW_LOGS=0 and PYTHIA_BUNDLE_FETCH_PROVIDER_OBJECTS=0 skip
# the two collectors that need credentials; both degrade cleanly without them.

When the zip exceeds 25 MB the workflow logs are split into pythia-debug-bundle-workflow-logs rather than dropped, and the manifest says so.

  • Forecast attribution bundle: pythia-forecast-attribution-bundle — ai_bundle/forecast_attribution__<run_id>.zip, built in the same Sibyl job right after the debug bundle. A signal ledger parsed from every model's reasoning trace (where probability mass moved and what the model said moved it, classed by a versioned regex taxonomy), prior anchoring against the base-rate anchor, RC assessment, trace quality, an input inventory per question, evidence items linked to signals, prompt section fingerprints and token share, model and Sibyl contrasts, run-over-run deltas and a brief per hazard. Its guide opens by saying the ledger is claimed attribution, not measured influence. Build locally with:
python -m scripts.ai_bundle.build_forecast_attribution_bundle --db duckdb:///data/resolver.duckdb \
  --out-dir ai_bundle [--run-id fc_1788237725] [--hs-run-id hs_20260901T040916]
  • Resolver debug bundle: resolver-debug-bundle-<run_id> — one zip per Resolver Update run, built by if: always() so the runs worth diagnosing are covered. It answers, from the zip alone: which URL each connector called and what came back (http/requests.jsonl, http/envelopes/ — including the response fields the connectors themselves discard); what thresholds and windows the run used (config/rulebook.yaml, run/env_effective.json); why any unresolved hazard cell produced no row (hazard/cell_ledger.csv, one reason code per cell); what ceiling a rejected figure was measured against and where that number came from (hazard/figures_ledger.csv); whether the extraction budget was exhausted and by whom (hazard/extraction_budget.csv); which code produced the run (code/, run/git.json); and whether any two sources in the run contradict each other (checks/contradictions.md). Start at README.md inside the zip. Build locally with:
python -m scripts.build_resolver_debug_bundle --db duckdb:///data/resolver.duckdb \
  --diagnostics-dir diagnostics --out debug_bundle/resolver-debug-bundle.zip
# Set PYTHIA_RUN_LOG_DIR before the ingest to capture HTTP requests, response
# envelopes and the machine's per-cell / per-figure ledgers; without it the
# bundle reports what the logs say and no more.

Hard ceiling of 80 MB compressed. When it binds, code/ is dropped before any evidence is and the manifest names what went; nothing is truncated silently. Every file is scanned for every secret in the build environment before the zip is written, and a hit fails the build.

  • Legacy monolithic bundle: python -m scripts.dump_pythia_debug_bundle --legacy still writes the single markdown file.
  • SPD compare JSON: debug/spd_compare_smoke or debug/spd_compare_tests
  • Diagnostics: diagnostics/ (DB signatures, compare JSON, latency summaries)
  • HS triage coverage: diagnostics/hs_triage_coverage__<HS_RUN_ID>.csv and diagnostics/hs_triage_failures__<HS_RUN_ID>.json
  • AI analysis bundle: pythia-ai-analysis-bundle (calibration workflow) — one zip per scoring round joining forecast reasoning to realized outcomes for every scored question; hand it to an AI analyst. Build locally with:
python -m scripts.ai_bundle.build_scored_forecast_bundle --db duckdb:///data/resolver.duckdb --out-dir ai_bundle
# options: --months-back 12, --n-case-studies 10, --include-test,
#          --include-sibyl-trials {case-studies|all|none}, --keep-staging

DB migrations

python scripts/migrate_llm_calls_telemetry.py --db duckdb:///path/to/resolver.duckdb

HS triage reruns

  • After the HS run completes, stdout includes HS_TRIAGE_RERUN_ISO3S=.... To re-run just those countries:
PYTHIA_HS_ONLY_COUNTRIES="<ISO3S>" python -m horizon_scanner.horizon_scanner

Dashboard

Run the API locally

PYTHIA_DB_URL="duckdb:///data/resolver.duckdb" uvicorn pythia.api.app:app --reload --port 8000

Production (Render) runs a single worker — multiple workers each open their own DuckDB connection and caches, multiplying memory on the 2GB instance:

uvicorn pythia.api.app:app --host 0.0.0.0 --port $PORT --workers 1

Recommended Render env: PYTHIA_MAX_CONCURRENT_HEAVY=1, MALLOC_ARENA_MAX=2. The /v1/run endpoint is disabled by default; set PYTHIA_ALLOW_INPROCESS_RUN=1 locally if you want the API to launch pipeline runs in-process.

Run the web UI locally

cd web
npm install
NEXT_PUBLIC_PYTHIA_API_BASE=http://localhost:8000/v1 npm run dev

Pages and RC visibility

  • Forecast Index (Overview): shows the Humanitarian Impact Forecast Index, with RC KPI counts and RC map markers (highest RC level per country).
  • Forecasts (/questions): latest forecasts table includes RC score, triage fields, and track assignment when latest_only=true.
  • HS Triage (/hs-triage): displays per-run triage with RC likelihood/direction/magnitude/score and tier (quiet/priority).
  • Question detail (/questions/[questionId]): RC fields appear alongside research and SPD outputs.
  • Countries (/countries): includes highest RC level/score per country from the latest HS run.
  • Performance (/performance): forecast evaluation with KPI cards, ensemble selector dropdown (compare ensemble_mean vs. ensemble_bayesmc), median and average scores (Brier, Log, CRPS), and views by Total, Hazard, Run, and Model.
  • Downloads (/downloads): forecast/triage exports with RC fields, plus per-question score CSVs, model-level summary CSVs, and rationale exports.
  • About (/about): versioned prompt snapshots and system overview history.

Downloads / exports

  • Forecast SPD & EIV export: /v1/downloads/forecasts.csv and /v1/downloads/forecasts.xlsx include RC probability/direction/magnitude/score columns and track assignment per row.
  • Score exports: per-question CSV for each named ensemble (full 6-month x 5-bin grid, expected impact values, resolutions, and Brier/Log/CRPS scores), model-level summary CSV (avg/median/min/max per model and hazard), and rationale export (human-readable LLM explanations by hazard).
  • Countries endpoint: /v1/countries includes highest_rc_level/highest_rc_score (latest HS run).
  • Questions endpoint: /v1/questions?latest_only=true includes RC fields (regime_change_*) and track.

See PUBLIC_APIS.md for canonical API contracts.

Configuration

Canonical config

  • pythia/config.yaml is the authoritative configuration.
  • There is no pythia/config.py; all runtime defaults live in the YAML + env vars.
  • app.db_url defines DuckDB; PYTHIA_DB_URL overrides at runtime.
  • llm.profile selects the default model bundle; override with PYTHIA_LLM_PROFILE.
  • llm.profiles.<name>.ensemble defines the forecast ensemble; see Model management.
  • pythia/model_costs.json contains per-model cost rates.

Key env vars

  • Provider keys: OPENAI_API_KEY, GEMINI_API_KEY, ANTHROPIC_API_KEY, EXA_API_KEY, PERPLEXITY_API_KEY.
  • Structured data: ACAPS_EMAIL, ACAPS_PASSWORD (ACAPS API auth); ACLED_ACCESS_KEY, ACLED_EMAIL (ACLED CAST API auth).
  • Concurrency:
    • PYTHIA_LLM_CONCURRENCY (global LLM call cap)
    • HS_MAX_WORKERS
    • FORECASTER_RESEARCH_MAX_WORKERS
    • FORECASTER_SPD_MAX_WORKERS
  • Grounding:
    • PYTHIA_GROUNDING_PRIMARY_BACKEND: openai (default) or gemini — controls which backend is tried first for RC/triage grounding.
    • PYTHIA_ADVERSARIAL_CHECK_ENABLED: Enable adversarial checks for RC L1+ (default 1).
    • Note: The legacy PYTHIA_WEB_RESEARCH_ENABLED, PYTHIA_RETRIEVER_ENABLED, and PYTHIA_HS_RESEARCH_WEB_SEARCH_ENABLED are deprecated (set to 0 in workflow).
  • SPD tuning / timeouts (Gemini 3):
    • PYTHIA_GOOGLE_SPD_THINKING_LEVEL_FLASH, PYTHIA_GOOGLE_SPD_THINKING_LEVEL_PRO
    • PYTHIA_GOOGLE_SPD_TIMEOUT_FLASH_SEC, PYTHIA_GOOGLE_SPD_TIMEOUT_PRO_SEC
    • PYTHIA_GOOGLE_SPD_RETRIES
  • SPD ensemble override: PYTHIA_SPD_ENSEMBLE_SPECS (e.g. openai:gpt-6-sol,google:gemini-3.5-flash).
  • HS triage resilience:
    • PYTHIA_HS_FALLBACK_MODEL_SPECS (defaults to hs_fallback from the active profile, then openai:gpt-6-sol; keeps HS triage running when Gemini fails).
    • PYTHIA_HS_ONLY_COUNTRIES (comma-separated ISO3s/names to rerun HS triage for a subset).
    • PYTHIA_PROVIDER_FAILURE_THRESHOLD, PYTHIA_PROVIDER_COOLDOWN_SECONDS, PYTHIA_PROVIDER_RESET_ON_SUCCESS
    • PYTHIA_LLM_RETRY_TIMEOUTS (set 0 to opt out of timeout retries outside HS triage)
    • PYTHIA_HS_LLM_MAX_ATTEMPTS (defaults to 3 for HS triage retries)
    • PYTHIA_HS_GEMINI_TIMEOUT_SEC (defaults to 120s for HS triage)

GitHub Secrets setup

Secret Used by Required for
OPENAI_API_KEY Forecaster SPD ensemble OpenAI models in ensemble
GEMINI_API_KEY HS, retriever, forecaster Gemini models (HS + SPD + retriever)
ANTHROPIC_API_KEY Forecaster SPD ensemble + Sibyl Anthropic models
EXA_API_KEY Web research Exa backend (optional)
PERPLEXITY_API_KEY Web research Perplexity backend (optional)
ACAPS_EMAIL Structured data ACAPS API auth (INFORM, Risk Radar, etc.)
ACAPS_PASSWORD Structured data ACAPS API auth
ACLED_ACCESS_KEY Structured data ACLED CAST API auth
ACLED_EMAIL Structured data ACLED CAST API auth
GITHUB_TOKEN Actions Artifact download + summary updates

Minimum set for full pipeline: GEMINI_API_KEY + at least one of (OPENAI_API_KEY, ANTHROPIC_API_KEY). Missing keys disable their providers, and the run proceeds with a partial ensemble.

Operational notes

  • Partial ensembles are expected: providers can timeout; the ensemble uses available members and records ensemble_meta.
  • Timeouts cap tail risk: per-provider timeouts are enforced; increase them only if needed.
  • Shared retriever: evidence packs are cached and reused to stabilize sources and reduce cost.
  • HS triage always writes: HS stores triage outputs even if no questions are produced.

Known limitations and challenges

  • Prediction market retriever disabled: Metaculus returns 403, Polymarket returns 422. Do not re-enable until upstream APIs are fixed.
  • IPC API unavailable: The IPC connector requires IPC_API_KEY which is not available; FEWS NET Phase 3+ data is used instead.
  • IFRC Montandon sparsity: natural hazard PA data may be sparse for some hazards/countries, reducing base-rate strength.
  • DI, CU, HW fully silenced: Displacement Influx, Civil Unrest, and Heatwave hazards are blocked at the hazard catalog level and never enter RC, triage, grounding, or question generation.
  • Occasional “200 but ungrounded”: some web search calls can return success without verified sources. OpenAI is the primary grounding backend; Gemini is the fallback.

Troubleshooting

  • No questions generated: HS may mark all tiers quiet. Check hs_triage in DuckDB and confirm your country list (horizon_scanner/hs_country_list.txt) and hazards_allowed in config.
  • No hazard tail packs: confirm PYTHIA_HS_HAZARD_TAIL_PACKS_ENABLED=1 and RC Level ≥1 for the hazard. Tail packs are limited to 2 hazards per country.
  • RC fields missing in API: ensure hs_triage has regime_change_* columns (pythia/db/schema.py) and that latest_only=true is set on /v1/questions.
  • No active models: verify PYTHIA_LLM_PROFILE, llm.profiles ensemble config in pythia/config.yaml, and provider API keys.
  • Debug bundle too large for step summary: artifacts still exist under debug/ even if GitHub Step Summary truncates.
  • Slow runs: Gemini tails can dominate latency. Tune PYTHIA_LLM_CONCURRENCY, FORECASTER_*_MAX_WORKERS, and SPD timeouts.
  • Interpreting question_run_metrics: question_run_metrics (if present) records wall-clock vs compute vs queue time per question; see scripts/dump_pythia_debug_bundle.py.

Cross-links

Contributing

See CONTRIBUTING.md for coding standards and workflow notes. Archived READMEs live in docs/archive/README_INDEX.md.

License

This repository follows the licensing terms bundled with the codebase (see LICENSE if present or repository metadata).

About

An experimental humanitarian impact forecasting system, combining parallel structured data injected ensemble forecasts and an agentic deep research forecasting approach. Known as Fred on public interface side

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages