Kytos Observatory — why we publish every prediction, and every time biology says we're wrong
The claim 01 / 03
Why this matters
Looks right. Is wrong.
Imagine a student takes a test and gets every question “correctly formatted” — right structure, no blank answers, nothing flagged by the grading software.
Imagine the answers are all wrong.
That’s what most AI model evals miss: they check if the output looks right, not if it is right. It’s like spell-check passing a grammatically perfect essay that says nothing true.
We’re building tools that predict how cells react to drugs — before you test them in a real lab. So “looks right but is wrong” isn’t a typo, it’s a drug that doesn’t work, discovered a year too late.
The evidence 02 / 03
🏆 VEED Summer Lock-In · Aug 2026
{Tech: Europe} × VEED Hackathon
Observatory Milestone 0 — the public transparency layer for our Virtual Cell Challenge entry, shipping from this repo through Nov 5.
Opt-in closed … · we keep building in public
📡 78-day public build
Virtual Cell Challenge 2026
The vessel
The κύτος vessel is a live readout of each run’s health — not decoration. Fill level tracks ceiling headroom; membrane stress and vesicles map to audit severity.
fill · 99% headroomcracks · 0 audit warnings
The substance 03 / 03
Why the Observatory
The Virtual Cell Challenge produces high-dimensional predictions and leaderboard scores. It does not publish when a model is biologically wrong while still scoring acceptably. We do — run by run, in the open, for the full competition window.
Live example (run #24): Overall score 0.031 · Perturbation discrimination (pds) 0.306 · +5 more — plus audit flags the headline metrics never name.
Official infrastructure
- Leaderboard +
cell-evalsix-metric scoring - Zero-shot across six unseen cell contexts (2026)
- High-quality Perturb-seq ground truth (~1k cells / perturbation)
Observatory adds
- Ceiling headroom (% of best achievable per metric)
- Biological audit flags separate from score
- Literature + NER + provenance + reproduce commands
Sources & further reading
- Arc — Virtual Cell Challenge 2026 (six metrics, biological vs numerical)
- Arc — 2025 wrap-up (5,000+ registrants, Generalist Prize)
- cell-eval2 (2026 metric scale 0–1)
- Nature Methods 2025 — perturbation models vs linear baselines
- Repo:
docs/substantiation.md(maintainer crib sheet)