Day 17 of 78 · Run #3 · 2026-09-05
Sparse mean-shift pipeline test — VCC 2026 validation, 360k cells, 300 targets
Overall score -0.948 · Perturbation discrimination (pds) 0.000 · +5 more
First real submission — a sparse top-300 mean-shift slice proved the pipeline end-to-end. Scored at the degenerate floor (-0.948) exactly as predicted: every target shares one control slice, so perturbation discrimination is zero.
Audit & metrics Overall score -0.948 · Perturbation discrimination (pds) 0.000 · +5 more
No audit flags.
Metrics vs ceiling (all scores)
Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
- A constant per-target sparse slice will score near the bottom of the leaderboard
- Perturbation discrimination will be zero because every target predicts the same control slice
Evidence Literature & entity enrichment
Literature
0 genesTavily · auxiliary, not scored
Literature pending — run tools/enrich_literature.py.
Field context
VCC / perturbation researchField research pending — run tools/enrich_newsroom.py.
Biomedical NER
0 entitiesPioneer GLiNER2 · fine-tuned on Tavily literature when available · regex fallback offline
NER pending — run tools/pioneer_ner.py --train once, then tools/pioneer_ner.py --run experiments/<run-id>.
Narrative digest run digest · traces to facts.json
Narrative pending — run tools/render_narrative.py.
Trust & provenance Self-tests + reproduce command
6 pipeline steps · sourced from committed artifacts only
No metrics CSVs found — run cell-eval and commit results.
Ran 5 deterministic rules · 0 flags raised (none fired).
Literature enrichment not yet run — execute tools/enrich_literature.py.
NER extraction not yet run — execute tools/pioneer_ner.py.
Narrative not yet generated — execute tools/render_narrative.py.
Verification artifacts not yet committed — run planted_signal.py, holo_audit.py, check_narrative.py.