Day 26 of 78 · Run #14 · 2026-09-14
kd_std=2.0 — overall +0.0511 (rank 462); fid -0.111 but gains are shrinking and nmae is eroding — the scalar optimum is near
Overall score 0.051 · Perturbation discrimination (pds) 0.336 · +5 more
kd_std=2.0 scored overall +0.0511 (rank 462): fid -0.148 -> -0.111 (+0.037, down from +0.061/step), pds 0.326 -> 0.336, reach 0.087 -> 0.091, jac -0.016 -> -0.014 — all still improving but the curve is bending. nmae eroded +0.008 -> +0.004 for the second consecutive step: overspread is blurring the DE signal. The scalar-eta optimum sits around 2.0-2.3; further scalar pushes trade nmae for fid at a worsening rate. The right next move is per-target spread, not a bigger global value.
Scorecard
All six VCC metrics — best-in-series is highlighted.
| Metric | Value | |
|---|---|---|
| Overall score | +0.0511 | |
| Perturbation discrimination (pds) | +0.3356 | |
| Expression accuracy (mse) | +0.0000 | best |
| DE log-FC accuracy (nmae) | +0.0042 | |
| DE direction fidelity (fid) | -0.1111 | |
| DE direction reach (reach) | +0.0912 | |
| Significance overlap (jac) | -0.0136 |
Audit & metrics Overall score 0.051 · Perturbation discrimination (pds) 0.336 · +5 more
No audit flags.
Metrics vs ceiling (all scores)
Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
- kd_std=2.0 probes near the Atlas global eta std (2.32) — expected to be at or past the fid optimum
- nmae erosion should continue as overspread blurs per-target DE log-FC signal
Evidence Literature & entity enrichment
Literature
0 genesTavily · auxiliary, not scored
Literature pending — run tools/enrich_literature.py.
Field context
VCC / perturbation researchField research pending — run tools/enrich_newsroom.py.
Biomedical NER
0 entitiesPioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline
NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.
Narrative digest run digest · traces to facts.json
Narrative pending — run tools/render_narrative.py.
Trust & provenance Self-tests + reproduce command
6 pipeline steps · sourced from committed artifacts only
No metrics CSVs found — run cell-eval and commit results.
Ran 5 deterministic rules · 0 flags raised (none fired).
Literature enrichment not yet run — execute tools/enrich_literature.py.
NER extraction not yet run — execute tools/pioneer_ner.py.
Narrative not yet generated — execute tools/render_narrative.py.
Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.