Per-target models that predict ligand/peptide binding from Boltz-2 affinity-module embeddings, benchmarked against Boltz-2's own scalar outputs and (Part 1) against a physically-motivated LRIP feature set. Two parts:
- Part 1 — ULVSH small molecules (closed). Per-target classification/ regression on 10 ULVSH targets from Boltz-2 embeddings, vs raw Boltz-2 and vs LRIP interaction profiles. Honest result: the learned embeddings roughly match (and do not decisively beat) raw Boltz-2 in aggregate; LRIP adds real signal over raw scalars but is largely redundant with the embeddings.
- Part 2 — peptide/mutation robustness (active). Does Boltz-2 track the
effect of mutations on binding? Current data is a 13-system SKEMPI subset
under
data/peptide_systems/, now with both Boltz embeddings and LRIP profiles. LRIP scores far belowpair_meanhere (median Spearman 0.037 vs 0.522) — but see the blocker below before reading anything into that. An earlier BH3/p53 peptide-ligand arm is on hold underdata/peptides/.
The modeling code is clean — joins, alignment, folds and nulls all check out. The problem is upstream in how the matrices were generated:
- LRIP was computed on Boltz-2 cofolds, a fresh receptor conformation per row. Both source papers hold one receptor fixed for every compound and minimize under backbone restraints; per-residue columns are only comparable across rows if the rows share a receptor frame.
- The peptide matrices use a −1.0 kcal/mol key-residue cutoff instead of the papers' −0.1, which deletes most of the residues SKEMPI actually mutates (183 of 3BT1's 240 single mutants have no column for the residue they mutate).
- Two previously published readings were wrong: "0.037 ≈ nothing" ignores that the protocol's scramble null is −0.119, not 0, and "LRIP is actively harmful when concatenated" is reproduced by same-shape Gaussian noise.
Reproduce with
PYTHONPATH=. python scripts/diagnose_lrip.py --part1 --part2. Full detail: the ACTIVE BLOCKER section at the top ofAGENTS.md.
The embedding exporter lives in the neighboring Boltz fork (../boltz/README.md):
it writes affinity_embeddings_<ligand>.npz files holding the pooled affinity
representation from immediately before the Boltz-2 scalar affinity heads.
See AGENTS.md for the deep project context, results, and caveats.
- ULVSH affinity labels, docking scores, structures: https://lab.drugdesign.unistra.fr/datasets/ulvsh/
- Boltz-2 paper input files (published examples/benchmarks, not the ULVSH label source): https://zenodo.org/records/16946890
- Methodology reference: Bret, Sindt, Rognan, Assessing Boltz-2 Performance for
the Binding Classification of Docking Hits, J. Chem. Inf. Model. 2026, 66,
1511-1521. PDF in
papers/. Per-target ROC AUC of raw Boltz-2 vs ULVSH active/inactive labels; AUC > 0.65 "acceptable", median across 10 targets 0.763. - Boltz-2 model: Passaro et al., Boltz-2, bioRxiv 2025.06.14.659707. PDF in
papers/. - LRIP / interaction-profile scoring: Ji et al., Briefings in Bioinformatics
22(5) 2021 (
papers/bbab054.pdf); Niu et al., LRIP-SF (papers/aef2177_CombinedPDF_v1.pdf).
pip install -r requirements.txt
pip install -e .boltz2_aff/ repository root
├── boltz2_aff/ importable Python package
│ ├── data.py ULVSH label loader, p_affinity derivation
│ ├── features.py embedding/scalar discovery, feature-set selection
│ ├── modeling.py train_classifier / train_regressor / boltz_baseline_metrics
│ ├── pipeline.py Part-1 CLI driver (python -m boltz2_aff.pipeline)
│ └── peptide_pipeline.py Part-2 per-system nested-CV Ridge CLI
├── scripts/
│ ├── sweep_embedding_keys.py Part-1 embedding-component sweep
│ ├── nested_cv.py Part-1 nested-CV (unbiased combo selection)
│ ├── diagnose_lrip.py LRIP audit -- RUN BEFORE ANY LRIP MODELING
│ ├── model_lrip.py Part-1 LRIP vs embeddings vs raw Boltz
│ ├── model_lrip_combined.py Part-1 "does LRIP add on top of Boltz?" increments
│ ├── make_boltz_inputs_peptide_systems.py active Part-2 input generator
│ ├── _build_aff_emb.py post-hoc embedding reconstruction from trunk z
│ └── parse_*/make_*/part2_*.py on-hold BH3/p53 Part-2 ingest + analysis
├── data/
│ ├── ulvsh/ Part 1
│ │ ├── source/<TARGET>/ imported ULVSH: raw/{vitro,scores}.tsv, minimized/<ZINC>/
│ │ ├── reference_boltz/inputs/ transferred paper-run YAMLs + job file
│ │ └── modeling/ model-ready inputs
│ │ ├── labels.tsv one normalized label row per (target, ligand_id)
│ │ ├── manifest.tsv per-target coverage
│ │ └── features/
│ │ ├── boltz_scalars.tsv B2-A/B2-C scalar table (all variants)
│ │ ├── boltz_embeddings/<TARGET>/ affinity_embeddings_<ligand>.npz
│ │ └── lrip/<TARGET>.dat per-residue interaction profiles (+ README)
│ ├── peptide_systems/ active Part 2 (13-system SKEMPI subset)
│ │ ├── source/<PDB>/ curated structures, FASTA, mutation tables
│ │ ├── boltz/inputs/<system>/ generated cofold YAMLs + measurement manifests
│ │ └── modeling/ consolidated embeddings + LRIP + labels
│ └── peptides/ on-hold BH3/p53 arm (source, boltz inputs, embeddings)
├── runs/ all model outputs (git-tracked summaries)
│ ├── sweep_embeddings/ sweep_combined/ Part-1 embedding sweeps
│ ├── lrip/ lrip_combined/ Part-1 LRIP results
│ ├── peptide_systems/ridge/ pre-LRIP Part-2 baseline
│ ├── peptide_systems/ridge_lrip/ active Part-2 median-label results
│ ├── peptide_systems/ridge_lrip_mean/ Part-2 mean-label sensitivity
│ └── peptide_embeddings/ on-hold Part-2 results
├── papers/ reference PDFs (incl. papers/peptides/ for Part 2)
├── AGENTS.md deep AI-facing project context + full results
└── README.md this file
Part 1 data is organized by processing stage under data/ulvsh/: imported
assets in source/, transferred paper-run inputs in reference_boltz/, and
compact model-ready labels/scalars/embeddings/LRIP in modeling/.
Default feature set is Boltz affinity embeddings. Run one, several, or all 10
targets with --targets (omit for all):
python -m boltz2_aff.pipeline --targets ROCK1 --out-dir runs/rock1_embeddings
python -m boltz2_aff.pipeline --out-dir runs/all_embeddings # all targetsOutputs under --out-dir:
dataset.csv— merged labels, metadata, selected feature columns.manifest.json— data coverage, target/variant counts, metrics summary (including the raw Boltz-2 baseline AUC on the same rows).models/classifier.joblib,models/regressor.joblib.models/metrics_*.json— cross-validation metrics.models/predictions_*.csv— cross-validated predictions when possible.
Use --feature-set to choose the modeling input:
embeddings(default): flattenedaffinity_embeddings_*.npzarrays only.boltz: six scalar Boltz affinity fields from the compact reference table.ulvsh_scores: original ULVSH docking/physics score columns (~25-28 numeric features per ligand: glide, vina, MMGB/SA, etc.).combined: concatenation of the three blocks above (~1055 cols on ROCK1).
Whenever Boltz scalar records are present they are also merged into dataset.csv
as metadata, so the raw Boltz-2 baseline ROC AUC is reported on the same rows
regardless of feature set. (LRIP is evaluated separately — see below.)
python -m boltz2_aff.pipeline --feature-set boltz --out-dir runs/boltz_scalar
python -m boltz2_aff.pipeline --feature-set ulvsh_scores --out-dir runs/ulvsh_scores
python -m boltz2_aff.pipeline --feature-set combined --out-dir runs/combinedEach affinity_embeddings_<ligand>.npz contains four arrays from the Boltz-2
affinity module:
| Key | Dims | Description |
|---|---|---|
pair_mean1 |
128 | Pooled receptor-ligand interface pair representation, ensemble member 1, before the scalar heads. |
pair_mean2 |
128 | Same pooling, ensemble member 2. |
head1 |
384 | Representation after the final affinity MLP, ensemble member 1, before the scalar prediction heads. |
head2 |
384 | Same, ensemble member 2. |
By default all four are concatenated (1024 dims). Restrict with --embedding-keys:
python -m boltz2_aff.pipeline --embedding-keys pair_mean1 --out-dir runs/pm1There is no universal best component across targets; pair_mean1 is a
reasonable fixed default. See the sweep section below.
The pipeline fits two tasks by default:
- Classification uses the ULVSH
Activecolumn. Rows with nonnumeric or percent-style affinity (<40%) are kept here. - Regression uses only uncensored numeric affinity (
Ki/EC50/IC50/Kdor providedpki) and trains onp_affinity = 6 - log10(value_uM)so larger = stronger binding.
Rows are grouped by target::ligand_id in cross-validation so multiple Boltz
variants of one ligand cannot leak across folds. Classification uses a no-PCA
RandomForestClassifier; regression uses RidgeCV in Boltz-residual mode. Key
reported metrics:
classification.cv_roc_auc— the headline metric (the paper's primary).regression.cv_roc_auc— a screening AUC: per fold train the regressor on uncensored-affinity rows, predictp_affinityfor every active-labeled test row (incl. censored inactives), AUC vsactive_bool. Mirrors the paper.boltz_baseline— raw Boltz-2 per-target AUC (no model fit): B2-Aboltz_affinity_pred_value(lower = stronger) and B2-Cboltz_affinity_probability_binary(higher = active). This is the bar to beat.
scripts/sweep_embedding_keys.py runs the pipeline per target per embedding
combination; scripts/nested_cv.py adds unbiased inner-fold combo selection.
python scripts/sweep_embedding_keys.py --feature-set combined --out-root runs/sweep_combined
python scripts/nested_cv.py --out runs/nested_cv.jsonHonest headline: the learned embeddings roughly match raw Boltz-2 and do not
decisively beat it in aggregate. With nested CV, the combined feature set beats
raw B2-C on 8/10 targets (median AUC ~0.80 vs 0.74), but against the better of
B2-A/B2-C per target the margin largely disappears. ROCK1 is the standout
(~0.90) but is a favorable, unrepresentative target. Persistent failures at tiny
n: ADRA2B (n=13, noise) and MTR1A (n=36). Regression is only meaningful for
ROCK1 and CASR; four targets have no uncensored numeric rows and are skipped.
Full per-target tables and caveats in AGENTS.md.
Per-residue ligand–receptor interaction energies (LRIP / IP-SF, Junmei Wang lab)
for all 10 targets in data/ulvsh/modeling/features/lrip/<TARGET>.dat (format +
join notes in that dir's README). Scored with the same RF / StratifiedGroupKFold
/ cv_roc_auc methodology as the embeddings, compared on the same rows.
python scripts/diagnose_lrip.py --part1 # AUDIT -- run this first -> stdout
python scripts/model_lrip.py # standalone LRIP vs emb vs raw Boltz -> runs/lrip/
python scripts/model_lrip_combined.py # does LRIP add on top of Boltz? -> runs/lrip_combined/Standalone LRIP is a negative: median AUC 0.612, below the embeddings
(0.744) and raw Boltz (best-of-B2A/B2C 0.793); clears >0.65 on 4/10. ROCK1 again
the standout (0.869). Caveat from the 2026-07-29 audit: cv_roc_auc is
cross-validated with a permutation null of ~0.44, while the raw Boltz baseline is
not cross-validated at all (null exactly 0.50) — the margin is overstated by
~0.06. The ordering holds; the size does not.
LRIP carries real signal, and the redundancy claim is not yet established.
Added to raw Boltz scalars it helps (median 0.666 → 0.699, +0.057, helps
7/10). Added on top of the learned stack (embeddings + Boltz scalars) it does
essentially nothing (0.748 → 0.750, +0.008). But that +0.008 is structurally
suppressed: RF's max_features="sqrt" gives the ~85 LRIP columns only ~7.6% of
per-split feature sampling against ~1030 embedding columns. Concluding that
Boltz-2 subsumes LRIP needs a feature-group-aware model and a CI on the
increment. Full breakdown and future work in AGENTS.md.
The active set is a curated SKEMPI subset of 13 protein–protein systems with
single and combinatorial mutation ΔΔG measurements, under
data/peptide_systems/source/<PDB>/. It replaces the BH3/p53 arm.
python scripts/make_boltz_inputs_peptide_systems.py # -> data/peptide_systems/boltz/inputs/
python -m boltz2_aff.peptide_pipeline
python -m boltz2_aff.peptide_pipeline --label mean --out-dir runs/peptide_systems/ridge_lrip_meanThe generator produces 1,705 sequence-only cofold YAMLs (1,692 unique mutants +
13 WT) from 2,123 mutant measurement rows, deduplicating structures while keeping
every observation in a one-to-many measurements.tsv. Production embeddings came
from the post-hoc path (scripts/_build_aff_emb.py pools the binder/partner
interface from retained trunk z) into data/peptide_systems/modeling/; that
bundle now also contains 1,705 transferred LRIP rows. The poses and MM-GBSA
protocol behind LRIP are not included, so end-to-end LRIP provenance remains
incomplete and raw Boltz scalars remain unavailable. The per-system pipeline
validates every feature/label/measurement join, builds WT-difference embeddings
and LRIP profiles, and fits one nested-CV Ridge model per system and feature
view; out-of-fold outputs report Spearman, ΔΔG-sign agreement, MAE, RMSE,
Pearson, and R².
LRIP result (2026-07-26), reinterpreted by the 2026-07-29 audit. Median
per-system Spearman is 0.037 for LRIP, 0.522 for pair_mean, 0.728
for mutation-only, and 0.773 for mutation+pair_mean. The mean-label
sensitivity pass is effectively identical. Evaluation is within-series Spearman
and ΔΔG sign, not pooled AUC.
Those numbers reproduce, but two conclusions drawn from them were wrong:
- 0.037 is not "no signal." The row-scramble null for this nested-Ridge protocol is −0.119, not 0. LRIP beats its own scrambled control on 8/13 systems — weak-but-real, concentrated in 1AO7 (0.616), 1CHO (0.356), 3HFM (0.217) and 1JTG (0.192).
- The negative concatenation increment is dimensionality, not LRIP. Adding
real LRIP to
pair_meancosts −0.146; same-shape Gaussian noise costs −0.150 and row-scrambled LRIP costs −0.222.
What actually varies across systems is matrix quality:
Spearman(locality, LRIP performance) = +0.637 against −0.02 for the
pair_mean control, where locality measures whether the LRIP change for a point
mutation lands on the residue that was mutated. Reproduce with
python scripts/diagnose_lrip.py --part2. Contract and full table in
data/peptide_systems/modeling/README.md; scientific detail and the producer
ask in AGENTS.md.
An earlier peptide-ligand arm (BH3 ↔ Mcl-1/Bcl-xL/Bfl-1, p53 TAD ↔ MDM2/MDMX)
lives under data/peptides/ with 2,139 extracted embeddings. First results
(embedding-model arm only): within-series Spearman 0.66–0.79 on BH3, |ΔΔG|
magnitude Spearman 0.65–0.92 on p53, and captured Bcl-2-family selectivity — the
mutational signal is present in the representation feeding Boltz-2's scalar
heads. The raw-Boltz scalar baseline (the direct Rognan comparison) is built but
blocked on running Boltz-2 over the peptide YAMLs. Retained but not active; full
detail in AGENTS.md.
See AGENTS.md for the full per-target breakdowns, caveats, LRIP future work,
and the Part-2 plan. A top-level notebooks/ directory can be added later for
exploratory visualization and figures — model fitting and metrics should stay in
reproducible package/script code, with notebooks reading saved out-of-fold
predictions and summaries from runs/.