Day 28 of 78 · Run #17 · 2026-09-16
mean-corrected trunc-normal kd_std=2.0 — overall +0.0090 (rank 565); nmae +0.017 as predicted, but fid -0.305 / pds 0.288: the applied mean inflation was load-bearing, not a bug
Overall score 0.009 · Perturbation discrimination (pds) 0.288 · +5 more
k010 s2.0 scored overall +0.0090 (rank 565): nmae +0.004 -> +0.017 (mean correction worked), but pds 0.336 -> 0.288 and fid -0.111 -> -0.305. Removing the trunc-normal's clipped-mean inflation (eta mean 1.40 -> 1.00) cost more in fid/pds than it gained in nmae — the champion's apparent bias was load-bearing. Two readings: (a) the metric rewards dispersion magnitude per se, or (b) our borrowed deltas are systematically ~40% too weak and the inflation was compensating. s4.0 (mean 1.0, std 1.24) disambiguates: if it still loses, spread-at-mean-1 can't buy back what inflated mean provided, pointing at delta-scale calibration as the real lever.
Scorecard
All six VCC metrics — best-in-series is highlighted.
| Metric | Value | |
|---|---|---|
| Overall score | +0.0090 | |
| Perturbation discrimination (pds) | +0.2882 | |
| Expression accuracy (mse) | +0.0000 | best |
| DE log-FC accuracy (nmae) | +0.0168 | |
| DE direction fidelity (fid) | -0.3047 | |
| DE direction reach (reach) | +0.0742 | |
| Significance overlap (jac) | -0.0206 |
Audit & metrics Overall score 0.009 · Perturbation discrimination (pds) 0.288 · +5 more
No audit flags.
Metrics vs ceiling (all scores)
Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
- Mean-1 eta at the winning symmetric shape should hold pds/fid near k008-s2.0 while recovering nmae
- If it underperforms, the +40% applied mean inflation is compensating for systematically weak borrowed deltas
Evidence Literature & entity enrichment
Literature
0 genesTavily · auxiliary, not scored
Literature pending — run tools/enrich_literature.py.
Field context
VCC / perturbation researchField research pending — run tools/enrich_newsroom.py.
Biomedical NER
0 entitiesPioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline
NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.
Narrative digest run digest · traces to facts.json
Narrative pending — run tools/render_narrative.py.
Trust & provenance Self-tests + reproduce command
6 pipeline steps · sourced from committed artifacts only
No metrics CSVs found — run cell-eval and commit results.
Ran 5 deterministic rules · 0 flags raised (none fired).
Literature enrichment not yet run — execute tools/enrich_literature.py.
NER extraction not yet run — execute tools/pioneer_ner.py.
Narrative not yet generated — execute tools/render_narrative.py.
Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.