Experiment runs

Metrics, ceiling headroom, and audit flags for every experiment.

24 runs published

Score trajectory

Overall VCC score across runs — the story is the slope, not any single point.

Cross-experiment matrix

Every run, side by side — how accurate, how much room is left, and whether it passes our biology checks.

Run ID Strategy Status Progress Score PDS Audit
k001-mean-shift-baseline mean-shift probe · mock data
38%
0.120 0.080 2 warn Inspect →
k002-vcc2025-validation-mean-shift mean-shift
58%
0.000 clean Inspect →
k003-mean-shift-validation mean-shift
6%
-0.948 0.000 clean Inspect →
k004-layer-a-b-validation layer-a-b
85%
-0.149 0.002 clean Inspect →
k004-real-resampling-validation control-cell resampling
70%
-0.304 -0.009 clean Inspect →
k005-atlas-prior-validation atlas-prior
85%
clean Inspect →
k006-replogle-prior-validation replogle-prior
98%
-0.021 0.265 clean Inspect →
k007-neighbor-prior-validation replogle-prior+neighbor-imputation
99%
-0.016 0.269 clean Inspect →
k008-kd-heterogeneity-validation replogle-prior+neighbor-imputation+heterogeneous-kd
99%
-0.011 0.272 clean Inspect →
k008-kd-s0p7-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=0.7)
99%
0.001 0.282 clean Inspect →
k008-kd-s1p0-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.0)
99%
0.015 0.297 clean Inspect →
k008-kd-s1p3-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.3)
99%
0.029 0.311 clean Inspect →
k008-kd-s1p7-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.7)
99%
0.043 0.326 clean Inspect →
k008-kd-s2p0-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)
99%
0.051 0.336 clean Inspect →
k009-gamma-kd-s1p4-validation replogle-prior+neighbor-imputation+gamma-kd(kd_std=1.4, eta_max=5.0, library_cap=median)
99%
-0.001 0.204 clean Inspect →
k009-gamma-kd-s2p0-validation replogle-prior+neighbor-imputation+gamma-kd(kd_std=2.0, eta_max=5.0, library_cap=median)
99%
0.007 0.208 clean Inspect →
k010-mean-corrected-kd-s2p0-validation replogle-prior+neighbor-imputation+mean-corrected-kd(kd_std=2.0)
99%
0.009 0.288 clean Inspect →
k010-mean-corrected-kd-s4p0-validation replogle-prior+neighbor-imputation+mean-corrected-kd(kd_std=4.0)
99%
0.012 0.290 clean Inspect →
k011-delta-scale-x1p3-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(1.3)
99%
0.056 0.304 clean Inspect →
k011-delta-scale-x1p7-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(1.7)
99%
0.060 0.339 clean Inspect →
k011-delta-scale-x2p0-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(2.0)
99%
0.057 0.356 clean Inspect →
k012-lineage-score mean-shift
6%
1 warn Inspect →
k012-transfer-loo mean-shift
6%
clean Inspect →
k013-context-scale-validation replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+per_context_delta_scale(A=1.7,B=2.65,C=0.75)
99%
0.031 0.306 clean Inspect →
k013-context-scale-validation 99% per-context delta_scale {A:1.7, B:2.65, C:0.75} scored +0.0312 (rank 560) — clean negative: redistributing the magnitude budget by lineage ratios lost vs uniform x1.7 (+0.0596) on every component Overall score 0.031 · Perturbation discrimination (pds) 0.306 · +5 more
k012-transfer-loo 6% LOO eval on 47 K562/hESC paired deltas: raw transfer cosine is only ~0.13 — context transfer is fundamentally weak; low-rank map lifts sign accuracy to 0.60 but destroys magnitude (deP 0.10); fitted scalar s=0.44 says hESC deltas are WEAKER than K562, opposite the 2026-context direction n pairs 47.000 · identity mean cosine 0.126 · +8 more
k012-lineage-score 6% context identity resolved on discriminative genes: A is clearly Jurkat-like (r=0.65 vs next best 0.40), B weakly RPE1-leaning (0.37), C unresolved/hESC-leaning — global Pearson is inflated by housekeeping+assay similarity and misleads A jurkat discriminative pearson 0.649 · A next best discriminative 0.399 · 1 audit warn · +6 more
k011-delta-scale-x2p0-validation 99% delta_scale=2.0 — overall +0.0566 (rank 515); the scale curve has bent: pds best-yet 0.356 and fid ~0, but nmae -0.125 outweighs the gains. x1.7 (+0.0596) stands as champion; the scalar magnitude axis is now bracketed Overall score 0.057 · Perturbation discrimination (pds) 0.356 · +5 more -0.00 Overall score vs prior
k011-delta-scale-x1p7-validation 99% delta_scale=1.7 — overall +0.0596 (rank 486), new best; fid nearly closed (-0.006) and pds back to 0.339, but nmae cratered to -0.076: overscaled DE magnitudes are now the cost Overall score 0.060 · Perturbation discrimination (pds) 0.339 · +5 more +0.00 Overall score vs prior
k011-delta-scale-x1p3-validation 99% explicit delta_scale=1.3 on the champion config — overall +0.0559 (rank 502), new best score; fid improved to -0.031 but nmae flipped negative (-0.023) and pds slipped to 0.304 Overall score 0.056 · Perturbation discrimination (pds) 0.304 · +5 more +0.04 Overall score vs prior
k010-mean-corrected-kd-s4p0-validation 99% mean-corrected kd_std=4.0 — overall +0.0122 (rank 564); more mean-1 spread helps marginally (s2.0 +0.0090) but stays far under the biased champion (+0.0511). Applied delta magnitude, not dispersion-at-mean-1, is what the score rewards. Overall score 0.012 · Perturbation discrimination (pds) 0.290 · +5 more +0.00 Overall score vs prior
k010-mean-corrected-kd-s2p0-validation 99% mean-corrected trunc-normal kd_std=2.0 — overall +0.0090 (rank 565); nmae +0.017 as predicted, but fid -0.305 / pds 0.288: the applied mean inflation was load-bearing, not a bug Overall score 0.009 · Perturbation discrimination (pds) 0.288 · +5 more +0.00 Overall score vs prior
k009-gamma-kd-s2p0-validation 99% gamma kd_std=2.0 — overall +0.0075 (rank 544); nmae best-ever +0.022 but pds 0.21 / fid -0.23 still far below trunc-normal s2.0 (+0.0511). The skewed eta shape itself is the problem, not the mean. Overall score 0.007 · Perturbation discrimination (pds) 0.208 · +5 more +0.01 Overall score vs prior
k009-gamma-kd-s1p4-validation 99% gamma kd_std=1.4 — overall -0.0011 (rank 565); nmae best-ever +0.020 confirms mean-bias fix, but pds 0.20 and fid -0.27 regress: near-zero-eta mass makes cells look unperturbed Overall score -0.001 · Perturbation discrimination (pds) 0.204 · +5 more -0.05 Overall score vs prior
k008-kd-s2p0-validation 99% kd_std=2.0 — overall +0.0511 (rank 462); fid -0.111 but gains are shrinking and nmae is eroding — the scalar optimum is near Overall score 0.051 · Perturbation discrimination (pds) 0.336 · +5 more +0.01 Overall score vs prior
k008-kd-s1p7-validation 99% kd_std=1.7 — overall +0.0429 (rank 477); fid -0.148, the sweep curve still has not bent Overall score 0.043 · Perturbation discrimination (pds) 0.326 · +5 more +0.01 Overall score vs prior
k008-kd-s1p3-validation 99% kd_std=1.3 — overall +0.0291 (rank 490); fid -0.209, still no bend in the sweep curve Overall score 0.029 · Perturbation discrimination (pds) 0.311 · +5 more +0.01 Overall score vs prior
k008-kd-s1p0-validation 99% kd_std=1.0 sweep point — overall +0.0152 (rank 509); fid gain accelerates as spread approaches Atlas-measured levels Overall score 0.015 · Perturbation discrimination (pds) 0.297 · +5 more +0.01 Overall score vs prior
k008-kd-s0p7-validation 99% kd_std=0.7 sweep point — first positive overall score (+0.0007); fid responds strongly to more KD heterogeneity Overall score 0.001 · Perturbation discrimination (pds) 0.282 · +5 more +0.01 Overall score vs prior
k008-kd-heterogeneity-validation 99% Heterogeneous per-cell knockdown (eta ~ N(1, 0.4)) on k007 priors — fid responds, overall improves to -0.0113 Overall score -0.011 · Perturbation discrimination (pds) 0.272 · +5 more +0.00 Overall score vs prior
k007-neighbor-prior-validation 99% STRING-neighbor imputation for the 28 unscreened targets — 18 recovered with real signature shape Overall score -0.016 · Perturbation discrimination (pds) 0.269 · +5 more +0.01 Overall score vs prior
k006-replogle-prior-validation 98% Replogle K562 GWPS + Atlas prior — real signatures for 272/300 targets Overall score -0.021 · Perturbation discrimination (pds) 0.265 · +5 more
k005-atlas-prior-validation 85% 2025 Atlas prior — real per-target signatures for the 4 overlapping 2026 targets Overall score undefined · Perturbation discrimination (pds) undefined
k004-real-resampling-validation 70% Real control-cell resampling — VCC 2026 validation, 360k cells, 300 targets Overall score -0.304 · Perturbation discrimination (pds) -0.009 · +5 more -0.15 Overall score vs prior
k004-layer-a-b-validation 85% Context-conditioned Layer A/B — hand-tuned target knockdown + log1p transport Overall score -0.149 · Perturbation discrimination (pds) 0.002 · +5 more +0.80 Overall score vs prior
k003-mean-shift-validation 6% Sparse mean-shift pipeline test — VCC 2026 validation, 360k cells, 300 targets Overall score -0.948 · Perturbation discrimination (pds) 0.000 · +5 more
k002-vcc2025-validation-mean-shift 58% First real cell-eval run — mean-shift floor vs VCC 2025 validation (H1 hESC, 50 targets) DE gene recall 0% ceiling · Pearson Δ undefined
k001-mean-shift-baseline 38% Basal mean-shift baseline — ceiling headroom probe DE gene recall 27% ceiling · Pearson Δ 26% ceiling · 2 audit warns probe · mock data