Day 27 of 78 · Run #15 · 2026-09-15
gamma kd_std=1.4 — overall -0.0011 (rank 565); nmae best-ever +0.020 confirms mean-bias fix, but pds 0.20 and fid -0.27 regress: near-zero-eta mass makes cells look unperturbed
Overall score -0.001 · Perturbation discrimination (pds) 0.204 · +5 more
gamma s1.4 scored overall -0.0011 (rank 565): nmae +0.004 -> +0.020 (best yet — the mean-preserving fix worked as designed), but pds 0.336 -> 0.204 and fid -0.111 -> -0.272 regressed hard, reach 0.091 -> 0.065, jac -0.014 -> -0.024. Interpretation: at kd_std=1.4 the capped gamma has shape 0.51, concentrating ~40%+ of eta mass near zero — those cells are effectively unperturbed, which directly erodes discrimination (pds) and direction fidelity (fid). The trunc-normal keeps mass centered on eta~1, which is what the metric rewards. Net: mean-preservation helps nmae but the skewed shape costs more than the bias did.
Scorecard
All six VCC metrics — best-in-series is highlighted.
| Metric | Value | |
|---|---|---|
| Overall score | -0.0011 | |
| Perturbation discrimination (pds) | +0.2037 | |
| Expression accuracy (mse) | +0.0000 | best |
| DE log-FC accuracy (nmae) | +0.0199 | |
| DE direction fidelity (fid) | -0.2717 | |
| DE direction reach (reach) | +0.0653 | |
| Significance overlap (jac) | -0.0235 |
Audit & metrics Overall score -0.001 · Perturbation discrimination (pds) 0.204 · +5 more
No audit flags.
Metrics vs ceiling (all scores)
Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
- Removing the trunc-normal's +17% mean inflation should recover nmae
- If fid/pds hold at matched spread, mean-inflation was the nmae cost; if they regress, the leaderboard rewards broad dispersion around eta=1 rather than a skewed eta density
Evidence Literature & entity enrichment
Literature
0 genesTavily · auxiliary, not scored
Literature pending — run tools/enrich_literature.py.
Field context
VCC / perturbation researchField research pending — run tools/enrich_newsroom.py.
Biomedical NER
0 entitiesPioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline
NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.
Narrative digest run digest · traces to facts.json
Narrative pending — run tools/render_narrative.py.
Trust & provenance Self-tests + reproduce command
6 pipeline steps · sourced from committed artifacts only
No metrics CSVs found — run cell-eval and commit results.
Ran 5 deterministic rules · 0 flags raised (none fired).
Literature enrichment not yet run — execute tools/enrich_literature.py.
NER extraction not yet run — execute tools/pioneer_ner.py.
Narrative not yet generated — execute tools/render_narrative.py.
Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.