k001-mean-shift-baselineBasal mean-shift baseline — ceiling headroom probek002-vcc2025-validation-mean-shiftFirst real cell-eval run — mean-shift floor vs VCC 2025 validation (H1 hESC, 50 targets)k003-mean-shift-validationSparse mean-shift pipeline test — VCC 2026 validation, 360k cells, 300 targetsk004-layer-a-b-validationContext-conditioned Layer A/B — hand-tuned target knockdown + log1p transportk004-real-resampling-validationReal control-cell resampling — VCC 2026 validation, 360k cells, 300 targetsk005-atlas-prior-validation2025 Atlas prior — real per-target signatures for the 4 overlapping 2026 targetsk006-replogle-prior-validationReplogle K562 GWPS + Atlas prior — real signatures for 272/300 targetsk007-neighbor-prior-validationSTRING-neighbor imputation for the 28 unscreened targets — 18 recovered with real signature shapek008-kd-heterogeneity-validationHeterogeneous per-cell knockdown (eta ~ N(1, 0.4)) on k007 priors — fid responds, overall improves to -0.0113k008-kd-s0p7-validationkd_std=0.7 sweep point — first positive overall score (+0.0007); fid responds strongly to more KD heterogeneityk008-kd-s1p0-validationkd_std=1.0 sweep point — overall +0.0152 (rank 509); fid gain accelerates as spread approaches Atlas-measured levelsk008-kd-s1p3-validationkd_std=1.3 — overall +0.0291 (rank 490); fid -0.209, still no bend in the sweep curvek008-kd-s1p7-validationkd_std=1.7 — overall +0.0429 (rank 477); fid -0.148, the sweep curve still has not bentk008-kd-s2p0-validationkd_std=2.0 — overall +0.0511 (rank 462); fid -0.111 but gains are shrinking and nmae is eroding — the scalar optimum is neark009-gamma-kd-s1p4-validationgamma kd_std=1.4 — overall -0.0011 (rank 565); nmae best-ever +0.020 confirms mean-bias fix, but pds 0.20 and fid -0.27 regress: near-zero-eta mass makes cells look unperturbedk009-gamma-kd-s2p0-validationgamma kd_std=2.0 — overall +0.0075 (rank 544); nmae best-ever +0.022 but pds 0.21 / fid -0.23 still far below trunc-normal s2.0 (+0.0511). The skewed eta shape itself is the problem, not the mean.k010-mean-corrected-kd-s2p0-validationmean-corrected trunc-normal kd_std=2.0 — overall +0.0090 (rank 565); nmae +0.017 as predicted, but fid -0.305 / pds 0.288: the applied mean inflation was load-bearing, not a bugk010-mean-corrected-kd-s4p0-validationmean-corrected kd_std=4.0 — overall +0.0122 (rank 564); more mean-1 spread helps marginally (s2.0 +0.0090) but stays far under the biased champion (+0.0511). Applied delta magnitude, not dispersion-at-mean-1, is what the score rewards.k011-delta-scale-x1p3-validationexplicit delta_scale=1.3 on the champion config — overall +0.0559 (rank 502), new best score; fid improved to -0.031 but nmae flipped negative (-0.023) and pds slipped to 0.304k011-delta-scale-x1p7-validationdelta_scale=1.7 — overall +0.0596 (rank 486), new best; fid nearly closed (-0.006) and pds back to 0.339, but nmae cratered to -0.076: overscaled DE magnitudes are now the costk011-delta-scale-x2p0-validationdelta_scale=2.0 — overall +0.0566 (rank 515); the scale curve has bent: pds best-yet 0.356 and fid ~0, but nmae -0.125 outweighs the gains. x1.7 (+0.0596) stands as champion; the scalar magnitude axis is now bracketedk012-lineage-scorecontext identity resolved on discriminative genes: A is clearly Jurkat-like (r=0.65 vs next best 0.40), B weakly RPE1-leaning (0.37), C unresolved/hESC-leaning — global Pearson is inflated by housekeeping+assay similarity and misleadsk012-transfer-looLOO eval on 47 K562/hESC paired deltas: raw transfer cosine is only ~0.13 — context transfer is fundamentally weak; low-rank map lifts sign accuracy to 0.60 but destroys magnitude (deP 0.10); fitted scalar s=0.44 says hESC deltas are WEAKER than K562, opposite the 2026-context directionk013-context-scale-validationper-context delta_scale {A:1.7, B:2.65, C:0.75} scored +0.0312 (rank 560) — clean negative: redistributing the magnitude budget by lineage ratios lost vs uniform x1.7 (+0.0596) on every component

Day 3 of 78 · Run #1 · 2026-08-22

Basal mean-shift baseline — ceiling headroom probe

probe · mock data

DE gene recall 27% ceiling · Pearson Δ 26% ceiling · 2 audit warns

Audit & metrics DE gene recall 27% ceiling · Pearson Δ 26% ceiling · 2 audit warns +

Metrics vs ceiling (all scores)

Differential expression volcano plot (log2FC vs -log10 p)
Pre-registered hypotheses (2)
  • Basal co-expression alone explains <20% of cross-context transfer on DESigGenesRecall
  • Mean-shift baseline will sit below 50% of ceiling on pearson_delta
Evidence 6 genes · 77 entities · Tavily + Pioneer +

Literature

6 genes

Tavily · auxiliary, not scored

ACTB2/3
GAPDH2/2
IFIT12/3
ISG152/3
MX12/2
OAS12/3

Field context

VCC / perturbation research

Biomedical NER

77 entities

Pioneer regex fallback · 77 entities across 6 genes · 34 gene, 29 perturbation type, 8 pathway, 3 cell type

77 entities6 genes
ACTB 10 entities · regex fallback
MOIgene PSgene PCCgene DEgene CRISPRgene CRISPRiperturbation type CRISPRperturbation type Knockdownperturbation type+ 2 more
GAPDH 14 entities · regex fallback
MCFgene S2gene S1gene MCF7gene PCRgene RTgene GAPDHgene NTCgene+ 6 more
IFIT1 19 entities · regex fallback
IIgene IIIgene L193gene TPR4gene R187gene IFIT1gene CRISPRgene UCSFgene+ 11 more
ISG15 9 entities · regex fallback
ISG15gene erkpathway autophagypathway CRISPRiperturbation type CRISPRperturbation type Cas9perturbation type knockdownperturbation type knockoutperturbation type+ 1 more
MX1 13 entities · regex fallback
CRISPRgene UCSFgene UPRgene ERgene unfolded protein responsepathway UPRpathway CRISPRiperturbation type CRISPRperturbation type+ 5 more
OAS1 12 entities · regex fallback
LINCSgene L1000gene CMAPgene OAS1gene CRISPRgene ATPgene type I interferonpathway CRISPRperturbation type+ 4 more
Narrative digest run digest · traces to facts.json +

> Fallback digest rendered deterministically from `facts.json` (no LLM call).

Full digest

Headline

Basal mean-shift baseline — ceiling headroom probe

Metrics

| metric | value | ceiling | |---|---|---| | DESigGenesRecall | 0.12 | 0.45 | | pearson_delta | 0.08 | 0.31 |

Provenance

- commit: `2fe7bd7f4f0902b454d0f2face73f917e6e065d5` - seed: 0 - code hash: `k001-mean-shift-v0` - hypotheses pre-registered: ['Basal co-expression alone explains <20% of cross-context transfer on DESigGenesRecall', 'Mean-shift baseline will sit below 50% of ceiling on pearson_delta']

Trust & provenance Self-tests + reproduce command +
🤖
Dr. Kytos Biological Observer Pipeline Execution Trace

6 pipeline steps · sourced from committed artifacts only

cell_eval.score_perturbations()

Evaluated 4 cell-eval metrics vs 2 ceiling bounds · data_status=probe.

kytos.audit.evaluate_rules()

Ran 5 deterministic rules · 2 flags raised (housekeeping_shift, pathway_coherence).

tavily.search_literature()

TAVILY_API_KEY not set — enrichment not attempted.

pioneer.ner_extract_entities()

Regex fallback extractor — 6 files enriched (no API key or all model calls failed).

narrative.synthesize_digest()

Deterministic fallback digest (no API key or LLM call failed).

verification.self_audit()

Passed: planted-signal 13/13 caught · holo-agent 4/4 passed · narrative-check 4/4 checks passed.

PASSPlanted-signal self-test13/13 caught · 13/13

We plant known-answer failures through our own audit rules. If it can't catch what we planted, it can't be trusted to catch what we didn't. 13/13 caught (13/13 cases).

python tools/planted_signal.py

PASSIndependent agent audit4/4 passed · 4/4

An autonomous agent (h/web-surfer-flash) browsed the live page and verified 4/4 passed (4/4 fields) against facts.json — the surface is honest, not just claimed.

Holo agent screenshot of the live run page

python tools/holo_audit.py --run experiments/k001-mean-shift-baseline

PASSNarrative grounding check4/4 checks passed · 4/4

The LLM digest is grounded by construction: every number in it must trace back to facts.json, verified deterministically. 4/4 checks passed (4/4 checks).

python tools/check_narrative.py --run experiments/k001-mean-shift-baseline