Skip to content

feat(history): comparable runs and trustworthy regression alerts - #24

Merged
jeffbrennan merged 1 commit into
mainfrom
feat/history-comparisons
Sep 13, 2026
Merged

jeffbrennan merged 1 commit into
mainfrom
feat/history-comparisons

Conversation

@jeffbrennan

Copy link
Copy Markdown
Owner

Summary

Implements brief 06 — comparable history and trustworthy regression alerts, in three increments.

Increment 1 — versioned records & deterministic storage

  • RunRecord v2 (RUN_RECORD_VERSION): provenance (status, backend, transport, compute_type, access_mode, runtime_version, client_version), workload_fingerprint, coverage, config; nullable measures; separate shuffle_read_bytes/shuffle_write_bytes and memory_bytes_spilled/disk_bytes_spilled; cumulative_time_s kept distinct from wall-clock duration_s.
  • Legacy rows read as record_version=1 with unknown provenance and are excluded from baselines; historical zeros are never reinterpreted as measured.
  • Deterministic format selection: .jsonl stays JSONL, existing Delta stays Delta; auto only uses the package heuristic for new stores.
  • JSONL reads infer across the whole store so a long legacy prefix cannot hide v2 columns.

Increment 2 — baseline selection and rule assessments

  • select_baseline_cohort requires the same workload identity (fingerprint, configurable), comparable coverage, and supports runtime/compute dimensions; filters null measurements before applying the window.
  • check_alerts returns a status-bearing AlertAssessment per rule (triggered/clean/insufficient_data/unsupported/skipped), with configurable min_samples, explicit true-zero handling, and finite-threshold/positive-window validation.
  • Capture forwards alert_output_path for on_trigger="file" and exposes alert_assessments.

Increment 3 — comparison report

  • workload_fingerprint (stable across node ids, changes with join type/scan path).
  • compare_runs + sparkparse compare CLI: compatible metrics, cohort/sample counts, excluded measures with reasons, and plan_changed.

Validation

  • just ci — ruff, ruff format, pyrefly (0 errors), 477 passed / 2 skipped.
  • Reviews addressed: cumulative time deduplicated per query; mixed-version JSONL schema inference; window counts only valid measurements.

Version RunRecord with provenance, coverage, configuration, workload
fingerprint, and distinct wall-clock/cumulative timing and memory/disk
spill measures. Missing measures stay null.

Make history storage selection deterministic so a .jsonl path stays JSONL
after installing deltalake, and read legacy records as version 1 rather
than treating historical zeros as measured values.

Select baseline cohorts by workload identity, coverage, and optional
runtime/compute dimensions with a configurable minimum sample count.
Evaluate every alert rule into a status-bearing assessment, validate
metric/window/threshold inputs, and forward alert output paths from the
capture API.

Add plan fingerprinting and a compare_runs report exposed through a new
`sparkparse compare` command.
@jeffbrennan
jeffbrennan merged commit 523e1cd into main Sep 13, 2026
4 checks passed
@jeffbrennan
jeffbrennan deleted the feat/history-comparisons branch September 13, 2026 01:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant