Open virtual-cell research
Kytos Observatory
We publish every prediction — and every place biology says we’re wrong. Kytos predicts how cells respond when a gene is silenced; for every run we publish the score, the biological failures the score misses, the evidence behind them, and the commands to reproduce it.
Rank 560 · Score 0.031 · VCC 2026 validation · 46 days to submit
Latest run
2026-09-18per-context delta_scale {A:1.7, B:2.65, C:0.75} scored +0.0312 (rank 560) — clean negative: redistributing the magnitude budget by lineage ratios lost vs uniform x1.7 (+0.0596) on every component
Overall score 0.031 · Perturbation discrimination (pds) 0.306 · +5 more
Inspect run →Score trajectory
The point of publishing every run is that the trajectory is honest — including the dead ends.
Score trajectory: -0.948 → 0.031 (Δ +0.979 over 19 runs)
Run log
What we tried, what happened, what it cost — in order.
- k003-mean-shift-validation-0.948
First real submission — a sparse top-300 mean-shift slice proved the pipeline end-to-end. Scored at the degenerate floor (-0.948) exactly as predicted: every target shares one control slice, so perturbation discrimination is zero.
- k004-real-resampling-validation-0.304
Swapped the sparse slice for real single-cell resampling — 360k control cells with natural dispersion. Score jumped to -0.304: dispersion matters, but the targets were still indistinguishable.
- k004-layer-a-b-validation-0.149
First target-specific model — a context-conditioned knockdown prior with log1p transport over real control cells. pds went positive for the first time (0.002); score reached -0.149.
- k005-atlas-prior-validationnot submitted
Atlas-prior attempt — the 2025 validation set only overlaps 4 of 300 2026 targets. Built and verified but never submitted: coverage, not the prior, was the bottleneck. This run is why we went broad.
- k006-replogle-prior-validation-0.021
Replogle K562 genome-wide Perturb-seq — real CRISPRi signatures for 272 of 300 targets. Score -0.021, rank 534, pds 0.265. The 28 targets outside Replogle still use the hand-tuned prior.
- k007-neighbor-prior-validation-0.016
The 28 fallback genes are genuinely unscreened (absent from K562 GWPS, K562 essential, and RPE1 screens; no alias recovery). Neighbor imputation from STRING partners recovered real signature shape for 18 of them (PSMB9->proteasome, TAF4->TFIID, MAPK7->MAP2K5/MEF2, MLKL->necroptosis). Score -0.0159 vs k006 -0.0210; nmae swung -0.074 -> +0.002; fid worsened slightly (-0.36 -> -0.40), sharpening the case that distribution shape — not signature means — is the deficit.
- k008-kd-heterogeneity-validation-0.011
k007 priors unchanged (dispatch real 816 / neighbor 54 / fallback 30); only the sampler changed: perturbed_i = basal_i + eta_i * delta + eps, eta ~ N(1, 0.4) clipped at 0, eps ~ N(0, 0.05). Overall -0.0159 -> -0.0113. fid improved modestly (-0.403 -> -0.388), nmae +0.002 -> +0.007, pds 0.269 -> 0.272. Hypothesis partially confirmed: scalar KD heterogeneity is real signal but small — it models spread along the delta axis only. The residual fid deficit likely lives in off-direction covariance, which needs the full Layer B sampler.
- k008-kd-s0p7-validation+0.001
Identical pipeline to k008 s0.4 except kd_std=0.7. Overall crossed zero: -0.0113 -> +0.0007, rank 519. fid -0.388 -> -0.334 (the largest single-run fid gain so far), nmae +0.007 -> +0.010, pds 0.272 -> 0.282, reach 0.062 -> 0.067. The scalar-eta mechanism has real headroom — the Atlas-measured spread (~1.1 median) suggests trying 1.0 next before concluding the mechanism is exhausted.
- k008-kd-s1p0-validation+0.015
kd_std=1.0 scored overall +0.0152 (rank 509): fid -0.334 -> -0.270 (largest gain yet), pds 0.282 -> 0.297, reach 0.067 -> 0.073, nmae +0.011. Every component improved again — the mechanism is not exhausted at the Atlas median. Note the measured eta distribution is heavy-tailed (global std 2.32, p75 2.51) so values >1 may still help; s1.3 submitted as the bracket point.
- k008-kd-s1p3-validation+0.029
kd_std=1.3 scored overall +0.0291 (rank 490): fid -0.270 -> -0.209, pds 0.297 -> 0.311, reach 0.073 -> 0.079, jac -0.019 -> -0.018; nmae flat (+0.011). fid gains per step: 0.4->0.7 gave +0.054, 0.7->1.0 gave +0.064, 1.0->1.3 gave +0.061 — still linear, not bending. The heavy right tail in measured eta (global std 2.32) explains why the optimum sits above the median; next probes 1.7 and ~2.0.
- k008-kd-s1p7-validation+0.043
kd_std=1.7 scored overall +0.0429 (rank 477): fid -0.209 -> -0.148 (+0.061, same step size as before), pds 0.311 -> 0.326, reach 0.079 -> 0.087, jac improving. nmae dipped +0.011 -> +0.008 — the first sign of a trade-off: overspread is starting to blur the DE log-FC signal. The optimum is near; 2.0 is the ceiling probe.
- k008-kd-s2p0-validation+0.051
kd_std=2.0 scored overall +0.0511 (rank 462): fid -0.148 -> -0.111 (+0.037, down from +0.061/step), pds 0.326 -> 0.336, reach 0.087 -> 0.091, jac -0.016 -> -0.014 — all still improving but the curve is bending. nmae eroded +0.008 -> +0.004 for the second consecutive step: overspread is blurring the DE signal. The scalar-eta optimum sits around 2.0-2.3; further scalar pushes trade nmae for fid at a worsening rate. The right next move is per-target spread, not a bigger global value.
- k009-gamma-kd-s1p4-validation-0.001
gamma s1.4 scored overall -0.0011 (rank 565): nmae +0.004 -> +0.020 (best yet — the mean-preserving fix worked as designed), but pds 0.336 -> 0.204 and fid -0.111 -> -0.272 regressed hard, reach 0.091 -> 0.065, jac -0.014 -> -0.024. Interpretation: at kd_std=1.4 the capped gamma has shape 0.51, concentrating ~40%+ of eta mass near zero — those cells are effectively unperturbed, which directly erodes discrimination (pds) and direction fidelity (fid). The trunc-normal keeps mass centered on eta~1, which is what the metric rewards. Net: mean-preservation helps nmae but the skewed shape costs more than the bias did.
- k009-gamma-kd-s2p0-validation+0.007
gamma s2.0 scored overall +0.0075 (rank 544): nmae +0.020 -> +0.022 (best yet), fid -0.272 -> -0.233, pds 0.204 -> 0.208, reach 0.065 -> 0.072. More spread helps directionally but the gamma A/B is conclusive: at matched kd_std=2.0 the mean-preserving gamma scores +0.0075 vs the biased trunc-normal's +0.0511. The leaderboard rewards symmetric dispersion centered on eta~1 — gamma's shape-0.25 density piles mass near eta~0 (unperturbed-looking cells) which directly costs discrimination. The nmae lesson is real though: remove the trunc-normal's mean inflation while KEEPING its symmetric shape — mean-corrected trunc-normal is the next variant.
- k010-mean-corrected-kd-s2p0-validation+0.009
k010 s2.0 scored overall +0.0090 (rank 565): nmae +0.004 -> +0.017 (mean correction worked), but pds 0.336 -> 0.288 and fid -0.111 -> -0.305. Removing the trunc-normal's clipped-mean inflation (eta mean 1.40 -> 1.00) cost more in fid/pds than it gained in nmae — the champion's apparent bias was load-bearing. Two readings: (a) the metric rewards dispersion magnitude per se, or (b) our borrowed deltas are systematically ~40% too weak and the inflation was compensating. s4.0 (mean 1.0, std 1.24) disambiguates: if it still loses, spread-at-mean-1 can't buy back what inflated mean provided, pointing at delta-scale calibration as the real lever.
- k010-mean-corrected-kd-s4p0-validation+0.012
k010 s4.0 scored overall +0.0122 (rank 564): nmae +0.019, pds 0.290, fid -0.292, reach 0.077 — marginal gain over s2.0 (+0.0090) but far below the uncorrected champion (+0.0511). The k009/k010 pair is now conclusive: at mean-1, no eta shape or spread tested recovers what the clipped trunc-normal's ~40% mean inflation delivered. The effective driver is applied delta magnitude — borrowed K562/Atlas signatures appear systematically weak in the 2026 contexts, and the 'bias' was compensating. Next honest lever: explicit delta-scale calibration (or context-conditioned Layer A scaling), not more eta-distribution search.
All runs
View all →| Run ID | Strategy | Status | Progress | Score | PDS | Audit | |
|---|---|---|---|---|---|---|---|
| k001-mean-shift-baseline | mean-shift |
probe · mock data | 0.120 | 0.080 | 2 warn | Inspect → | |
| k002-vcc2025-validation-mean-shift | mean-shift |
0.000 | — | clean | Inspect → | ||
| k003-mean-shift-validation | mean-shift |
-0.948 | 0.000 | clean | Inspect → | ||
| k004-layer-a-b-validation | layer-a-b |
-0.149 | 0.002 | clean | Inspect → | ||
| k004-real-resampling-validation | control-cell resampling |
-0.304 | -0.009 | clean | Inspect → | ||
| k005-atlas-prior-validation | atlas-prior |
— | — | clean | Inspect → | ||
| k006-replogle-prior-validation | replogle-prior |
-0.021 | 0.265 | clean | Inspect → | ||
| k007-neighbor-prior-validation | replogle-prior+neighbor-imputation |
-0.016 | 0.269 | clean | Inspect → | ||
| k008-kd-heterogeneity-validation | replogle-prior+neighbor-imputation+heterogeneous-kd |
-0.011 | 0.272 | clean | Inspect → | ||
| k008-kd-s0p7-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=0.7) |
0.001 | 0.282 | clean | Inspect → | ||
| k008-kd-s1p0-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.0) |
0.015 | 0.297 | clean | Inspect → | ||
| k008-kd-s1p3-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.3) |
0.029 | 0.311 | clean | Inspect → | ||
| k008-kd-s1p7-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=1.7) |
0.043 | 0.326 | clean | Inspect → | ||
| k008-kd-s2p0-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0) |
0.051 | 0.336 | clean | Inspect → | ||
| k009-gamma-kd-s1p4-validation | replogle-prior+neighbor-imputation+gamma-kd(kd_std=1.4, eta_max=5.0, library_cap=median) |
-0.001 | 0.204 | clean | Inspect → | ||
| k009-gamma-kd-s2p0-validation | replogle-prior+neighbor-imputation+gamma-kd(kd_std=2.0, eta_max=5.0, library_cap=median) |
0.007 | 0.208 | clean | Inspect → | ||
| k010-mean-corrected-kd-s2p0-validation | replogle-prior+neighbor-imputation+mean-corrected-kd(kd_std=2.0) |
0.009 | 0.288 | clean | Inspect → | ||
| k010-mean-corrected-kd-s4p0-validation | replogle-prior+neighbor-imputation+mean-corrected-kd(kd_std=4.0) |
0.012 | 0.290 | clean | Inspect → | ||
| k011-delta-scale-x1p3-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(1.3) |
0.056 | 0.304 | clean | Inspect → | ||
| k011-delta-scale-x1p7-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(1.7) |
0.060 | 0.339 | clean | Inspect → | ||
| k011-delta-scale-x2p0-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+delta_scale(2.0) |
0.057 | 0.356 | clean | Inspect → | ||
| k012-lineage-score | mean-shift |
— | — | 1 warn | Inspect → | ||
| k012-transfer-loo | mean-shift |
— | — | clean | Inspect → | ||
| k013-context-scale-validation | replogle-prior+neighbor-imputation+heterogeneous-kd(kd_std=2.0)+per_context_delta_scale(A=1.7,B=2.65,C=0.75) |
0.031 | 0.306 | clean | Inspect → |
CHRONICLE · 2026-09-18
The edge of the map
Borrowed biology kept winning — until we asked which cells we were actually predicting.
Our best score (+0.0596, rank 486 of 984) came from borrowing real perturbation signatures — but a lineage check showed the challenge cells are not the cells those signatures came from: context A looks Jurkat-like, B weakly RPE1, C unresolved. Three honest attempts to fix the mismatch all failed: moving K562 signatures keeps only 13% of their direction in a new cell type, lineage-matched scaling scored worse than the uniform guess, and no public screen of the right cells covers any of the 300 targets. The lookup has a ceiling — so we stopped borrowing and started training our own model on a rented GPU.
CHRONICLE · 2026-09-11
Borrowed biology beat invented biology
We moved a genome-wide CRISPRi screen into a cell type it never saw — 272 of 300 targets got real signatures.
The Virtual Cell Challenge asks: switch off one gene, predict what the other 18,000 do — in a cell type your model has never seen. Instead of inventing how genes respond, we borrowed a real catalog: Replogle's genome-wide CRISPRi screen, run in K562 cells, moved into H1 hESC. It gave 272 of the 300 targets genuine signatures, and our leaderboard score went from -0.149 to -0.021 — rank 534.