Skip to content

Harden exact solve status handling and reduce DW memory - #47

Merged
berendmarkhorst merged 1 commit into
mainfrom
fix/status-safe-exact-solves
Sep 4, 2026
Merged

berendmarkhorst merged 1 commit into
mainfrom
fix/status-safe-exact-solves

Conversation

@berendmarkhorst

@berendmarkhorst berendmarkhorst commented Sep 4, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • make HiGHS and Gurobi termination handling explicit across the core cut loop, SAP, prize-collecting, budget, and MWCSPB runners
  • return only independently connectivity-valid incomplete incumbents, always with an unknown (inf) gap unless an independent bound provides a conservative nonzero certificate
  • discard stale pre-cut values when the global deadline expires before the required re-solve; public APIs now raise when no valid incomplete incumbent exists
  • preserve historical tuple shapes by making internal status fields opt-in via return_status=True
  • reject negative edge/arc costs at edge-cost Steiner entry points while retaining negative MWCS node weights and topology-only MWCSPB edge attributes
  • clamp the MWCSP-to-PCSTP shift at zero so all-positive node-weight instances no longer create invalid negative transformed edge costs
  • reduce Dreyfus-Wagner forward-pass memory to label arrays only and recompute split/predecessor data along the final reconstruction path

Status and validation coverage

Regression tests cover tiny positive time limits, no incumbent, proven infeasibility, valid-but-unproven incumbents, expiry immediately after violated cuts, preprocessing, SAP/transformed and specialized paths, and both HiGHS and licensed Gurobi backends. They also cover the disconnected negative-cycle counterexample, custom arc-weight attributes, transform entry points, MWCS negative node weights, and Dreyfus-Wagner tied paths, zero weights, disconnected graphs, and ILP-oracle comparisons.

Dreyfus-Wagner benchmark

Compared origin/main at 90a587d with this implementation using fresh subprocesses, fixed seed 20260904, three repetitions per case, and identical Python 3.13.4 / NetworkX 3.5 / SciPy 1.17.0 environments. f is the target edge factor (m = f*n). Every objective matched.

k n f main ms candidate ms cand/main main RSS MiB candidate RSS MiB RSS delta MiB
6 300 3 3.026 3.782 1.250 71.1 71.2 +0.0
6 300 8 3.681 4.703 1.278 72.0 72.1 +0.0
6 1200 3 10.060 12.475 1.240 73.7 73.3 -0.4
6 1200 8 13.820 17.335 1.254 76.6 76.4 -0.2
7 300 3 5.720 6.249 1.092 71.3 71.3 -0.0
7 300 8 6.411 7.151 1.115 72.2 71.9 -0.3
7 1200 3 18.875 20.336 1.077 74.6 73.8 -0.8
7 1200 8 23.344 26.450 1.133 77.7 76.7 -1.0
8 300 3 11.063 11.056 0.999 71.8 71.2 -0.6
8 300 8 11.928 12.363 1.036 72.7 72.4 -0.4
8 1200 3 38.098 39.278 1.031 76.3 74.3 -2.0
8 1200 8 42.566 45.193 1.062 79.4 77.6 -1.8
9 300 3 22.404 21.937 0.979 72.9 71.9 -1.0
9 300 8 23.738 23.526 0.991 73.7 72.6 -1.1
9 1200 3 78.033 76.759 0.984 79.8 75.7 -4.1
9 1200 8 83.400 82.588 0.990 83.0 78.7 -4.2
10 300 3 49.507 45.519 0.919 74.7 72.5 -2.3
10 300 8 48.145 45.599 0.947 75.6 73.4 -2.2
10 1200 3 166.402 151.494 0.910 86.4 78.7 -7.8
10 1200 8 169.626 164.295 0.969 89.6 81.7 -7.9

The median runtime ratio across the matrix was 1.034, so this PR makes no general runtime-improvement claim. Median process RSS changed by -1.0 MiB across all cases; the largest measured reduction was 7.9 MiB at k=10, n=1200.

Nested-cut profiling

On a fixed n=600, m=3600, 60-terminal separation workload (nested=1, one separation thread, 20 profiled calls), CSR construction accounted for 0.163 s of 2.850 s total (5.7%), versus 1.617 s cumulative max-flow time. A prototype per-thread CSR cache reduced construction calls from 1200 to 20, but nine unprofiled repetitions were indistinguishable: 0.085434 s cached versus 0.085517 s on main. The prototype was reverted and this PR contains no nested-cut change.

Validation

  • python3 -m pytest -q: 795 passed, 26 skipped, 20 xpassed
  • GitHub Actions: Python 3.8, 3.9, 3.10, 3.11, and 3.12 jobs all pass
  • new standalone files: Black clean; flake8 clean with E501 ignored to match the configured Black line length
  • full flake8 baseline comparison (E501 ignored): 152 findings here versus 156 on main
  • full mypy baseline comparison (--ignore-missing-imports): 69 existing findings here versus 71 on main; isolated new files pass with --follow-imports=skip
  • git diff --check and compileall: clean
  • warning-strict Sphinx was attempted, but the local iCloud-backed virtual environment blocked at zero CPU while importing a SciPy binary extension; it was interrupted before document parsing and emitted no documentation warning

@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 76.10922% with 70 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
steinerpy/mathematical_model.py 68.39% 61 Missing ⚠️
steinerpy/objects.py 85.00% 9 Missing ⚠️

📢 Thoughts on this report? Let us know!

@berendmarkhorst
berendmarkhorst marked this pull request as ready for review September 4, 2026 14:41
@berendmarkhorst
berendmarkhorst merged commit 82f0b66 into main Sep 4, 2026
6 checks passed
@berendmarkhorst
berendmarkhorst deleted the fix/status-safe-exact-solves branch September 4, 2026 14:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants