Day 27 of 78 · Run #16 · 2026-09-15
gamma kd_std=2.0 — overall +0.0075 (rank 544); nmae best-ever +0.022 but pds 0.21 / fid -0.23 still far below trunc-normal s2.0 (+0.0511). The skewed eta shape itself is the problem, not the mean.
Overall score 0.007 · Perturbation discrimination (pds) 0.208 · +5 more
gamma s2.0 scored overall +0.0075 (rank 544): nmae +0.020 -> +0.022 (best yet), fid -0.272 -> -0.233, pds 0.204 -> 0.208, reach 0.065 -> 0.072. More spread helps directionally but the gamma A/B is conclusive: at matched kd_std=2.0 the mean-preserving gamma scores +0.0075 vs the biased trunc-normal's +0.0511. The leaderboard rewards symmetric dispersion centered on eta~1 — gamma's shape-0.25 density piles mass near eta~0 (unperturbed-looking cells) which directly costs discrimination. The nmae lesson is real though: remove the trunc-normal's mean inflation while KEEPING its symmetric shape — mean-corrected trunc-normal is the next variant.
Scorecard
All six VCC metrics — best-in-series is highlighted.
| Metric | Value | |
|---|---|---|
| Overall score | +0.0075 | |
| Perturbation discrimination (pds) | +0.2083 | |
| Expression accuracy (mse) | +0.0000 | best |
| DE log-FC accuracy (nmae) | +0.0225 | best |
| DE direction fidelity (fid) | -0.2329 | |
| DE direction reach (reach) | +0.0721 | |
| Significance overlap (jac) | -0.0252 |
Audit & metrics Overall score 0.007 · Perturbation discrimination (pds) 0.208 · +5 more
No audit flags.
Metrics vs ceiling (all scores)
Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
- Doubling gamma spread should recover some fid/pds vs s1.4 if dispersion is what the metric rewards
- nmae should stay strong since mean-preservation is shape-independent
Evidence Literature & entity enrichment
Literature
0 genesTavily · auxiliary, not scored
Literature pending — run tools/enrich_literature.py.
Field context
VCC / perturbation researchField research pending — run tools/enrich_newsroom.py.
Biomedical NER
0 entitiesPioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline
NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.
Narrative digest run digest · traces to facts.json
Narrative pending — run tools/render_narrative.py.
Trust & provenance Self-tests + reproduce command
6 pipeline steps · sourced from committed artifacts only
No metrics CSVs found — run cell-eval and commit results.
Ran 5 deterministic rules · 0 flags raised (none fired).
Literature enrichment not yet run — execute tools/enrich_literature.py.
NER extraction not yet run — execute tools/pioneer_ner.py.
Narrative not yet generated — execute tools/render_narrative.py.
Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.