k001-mean-shift-baselineBasal mean-shift baseline — ceiling headroom probek002-vcc2025-validation-mean-shiftFirst real cell-eval run — mean-shift floor vs VCC 2025 validation (H1 hESC, 50 targets)k003-mean-shift-validationSparse mean-shift pipeline test — VCC 2026 validation, 360k cells, 300 targetsk004-layer-a-b-validationContext-conditioned Layer A/B — hand-tuned target knockdown + log1p transportk004-real-resampling-validationReal control-cell resampling — VCC 2026 validation, 360k cells, 300 targetsk005-atlas-prior-validation2025 Atlas prior — real per-target signatures for the 4 overlapping 2026 targetsk006-replogle-prior-validationReplogle K562 GWPS + Atlas prior — real signatures for 272/300 targetsk007-neighbor-prior-validationSTRING-neighbor imputation for the 28 unscreened targets — 18 recovered with real signature shapek008-kd-heterogeneity-validationHeterogeneous per-cell knockdown (eta ~ N(1, 0.4)) on k007 priors — fid responds, overall improves to -0.0113k008-kd-s0p7-validationkd_std=0.7 sweep point — first positive overall score (+0.0007); fid responds strongly to more KD heterogeneityk008-kd-s1p0-validationkd_std=1.0 sweep point — overall +0.0152 (rank 509); fid gain accelerates as spread approaches Atlas-measured levelsk008-kd-s1p3-validationkd_std=1.3 — overall +0.0291 (rank 490); fid -0.209, still no bend in the sweep curvek008-kd-s1p7-validationkd_std=1.7 — overall +0.0429 (rank 477); fid -0.148, the sweep curve still has not bentk008-kd-s2p0-validationkd_std=2.0 — overall +0.0511 (rank 462); fid -0.111 but gains are shrinking and nmae is eroding — the scalar optimum is neark009-gamma-kd-s1p4-validationgamma kd_std=1.4 — overall -0.0011 (rank 565); nmae best-ever +0.020 confirms mean-bias fix, but pds 0.20 and fid -0.27 regress: near-zero-eta mass makes cells look unperturbedk009-gamma-kd-s2p0-validationgamma kd_std=2.0 — overall +0.0075 (rank 544); nmae best-ever +0.022 but pds 0.21 / fid -0.23 still far below trunc-normal s2.0 (+0.0511). The skewed eta shape itself is the problem, not the mean.k010-mean-corrected-kd-s2p0-validationmean-corrected trunc-normal kd_std=2.0 — overall +0.0090 (rank 565); nmae +0.017 as predicted, but fid -0.305 / pds 0.288: the applied mean inflation was load-bearing, not a bugk010-mean-corrected-kd-s4p0-validationmean-corrected kd_std=4.0 — overall +0.0122 (rank 564); more mean-1 spread helps marginally (s2.0 +0.0090) but stays far under the biased champion (+0.0511). Applied delta magnitude, not dispersion-at-mean-1, is what the score rewards.k011-delta-scale-x1p3-validationexplicit delta_scale=1.3 on the champion config — overall +0.0559 (rank 502), new best score; fid improved to -0.031 but nmae flipped negative (-0.023) and pds slipped to 0.304k011-delta-scale-x1p7-validationdelta_scale=1.7 — overall +0.0596 (rank 486), new best; fid nearly closed (-0.006) and pds back to 0.339, but nmae cratered to -0.076: overscaled DE magnitudes are now the costk011-delta-scale-x2p0-validationdelta_scale=2.0 — overall +0.0566 (rank 515); the scale curve has bent: pds best-yet 0.356 and fid ~0, but nmae -0.125 outweighs the gains. x1.7 (+0.0596) stands as champion; the scalar magnitude axis is now bracketedk012-lineage-scorecontext identity resolved on discriminative genes: A is clearly Jurkat-like (r=0.65 vs next best 0.40), B weakly RPE1-leaning (0.37), C unresolved/hESC-leaning — global Pearson is inflated by housekeeping+assay similarity and misleadsk012-transfer-looLOO eval on 47 K562/hESC paired deltas: raw transfer cosine is only ~0.13 — context transfer is fundamentally weak; low-rank map lifts sign accuracy to 0.60 but destroys magnitude (deP 0.10); fitted scalar s=0.44 says hESC deltas are WEAKER than K562, opposite the 2026-context directionk013-context-scale-validationper-context delta_scale {A:1.7, B:2.65, C:0.75} scored +0.0312 (rank 560) — clean negative: redistributing the magnitude budget by lineage ratios lost vs uniform x1.7 (+0.0596) on every component

Day 26 of 78 · Run #14 · 2026-09-14

kd_std=2.0 — overall +0.0511 (rank 462); fid -0.111 but gains are shrinking and nmae is eroding — the scalar optimum is near

Overall score 0.051 · Perturbation discrimination (pds) 0.336 · +5 more

Leaderboard rank#462
Overall score0.051
Cells predicted360,000
Compute cost$0.90
Strategyreplogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)

kd_std=2.0 scored overall +0.0511 (rank 462): fid -0.148 -> -0.111 (+0.037, down from +0.061/step), pds 0.326 -> 0.336, reach 0.087 -> 0.091, jac -0.016 -> -0.014 — all still improving but the curve is bending. nmae eroded +0.008 -> +0.004 for the second consecutive step: overspread is blurring the DE signal. The scalar-eta optimum sits around 2.0-2.3; further scalar pushes trade nmae for fid at a worsening rate. The right next move is per-target spread, not a bigger global value.

Scorecard

All six VCC metrics — best-in-series is highlighted.

MetricValue
Overall score+0.0511
Perturbation discrimination (pds)+0.3356
Expression accuracy (mse)+0.0000best
DE log-FC accuracy (nmae)+0.0042
DE direction fidelity (fid)-0.1111
DE direction reach (reach)+0.0912
Significance overlap (jac)-0.0136
What’s next
Audit & metrics Overall score 0.051 · Perturbation discrimination (pds) 0.336 · +5 more +

No audit flags.

Metrics vs ceiling (all scores)

Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
  • kd_std=2.0 probes near the Atlas global eta std (2.32) — expected to be at or past the fid optimum
  • nmae erosion should continue as overspread blurs per-target DE log-FC signal
Evidence Literature & entity enrichment +

Literature

0 genes

Tavily · auxiliary, not scored

Literature pending — run tools/enrich_literature.py.

Field context

VCC / perturbation research

Field research pending — run tools/enrich_newsroom.py.

Biomedical NER

0 entities

Pioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline

NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.

Narrative digest run digest · traces to facts.json +

Narrative pending — run tools/render_narrative.py.

Trust & provenance Self-tests + reproduce command +
🤖
Dr. Kytos Biological Observer Pipeline Execution Trace

6 pipeline steps · sourced from committed artifacts only

cell_eval.score_perturbations()

No metrics CSVs found — run cell-eval and commit results.

kytos.audit.evaluate_rules()

Ran 5 deterministic rules · 0 flags raised (none fired).

tavily.search_literature()

Literature enrichment not yet run — execute tools/enrich_literature.py.

pioneer.ner_extract_entities()

NER extraction not yet run — execute tools/pioneer_ner.py.

narrative.synthesize_digest()

Narrative not yet generated — execute tools/render_narrative.py.

verification.self_audit()

Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.

PENDINGPlanted-signal self-testnot run yet

Run python tools/planted_signal.py --json …/verification/planted_signal.json.

PENDINGIndependent agent auditnot run yet

Run python tools/holo_audit.py --run experiments/<run-id>.

PENDINGNarrative grounding checknot run yet

Run python tools/check_narrative.py --run experiments/<run-id> after the digest is generated.