Pretrained models and PyTorch tools for Cerberus, along with its predecessors Borzoi and Borzoi Prime. These deep neural networks predict regulatory activity (e.g. chromatin accessibility and gene expression) from DNA sequence.
Cerberus replaces Borzoi's transformer with bidirectional state space (Hydra) blocks. It reads 786 kb of sequence and predicts 8,361 human and 3,102 mouse functional genomics tracks at 32 bp resolution, including RNA-seq, CAGE, DNase/ATAC, ChIP, CLIP, and 3′ RNA-seq. It is more accurate and cheaper to run than Borzoi. See releases/cerberus.
The TensorFlow version of baskerville, used by borzoi, has moved to baskerville-tf.
Requires Python 3.11+ and, for gene and SNP scoring, bedtools on PATH
(e.g. apt install bedtools or conda install -c bioconda bedtools).
git clone https://github.com/calico/baskerville.git
cd baskerville
pip install .Cerberus needs an NVIDIA GPU and the cuda extra (mamba-ssm). For a
development setup, install in editable mode with the dev extras:
pip install -e ".[dev,cuda]" # drop ,cuda on CPU-only hosts (Borzoi only)The *_folds.py scripts run jobs locally by default, or on GCP Batch with
--backend gcp (see GCP configuration). --backend slurm
needs slurmrunner, which is not publicly available.
The hdwig extra (calico/hdwig, not yet
public) is needed only to read .hw coverage files in hound_data.
On a GPU host, download one pretrained Cerberus replicate and check that it
reproduces the published forward pass. hound_verify reads the weights from the
clone, so install it editable (pip install -e ".[cuda]"):
cd releases/cerberus
./download.sh 0 # -> models/f0c0/model_best.pth
hound_verify --family cerberusSee Model releases for all weights, and the guides below for training and scoring.
For detailed instructions on dataset construction, training, and evaluating models, see:
Training
- Training models
- Cross‑fold training
- Transfer learning from a pretrained model
- Ensemble distillation (precomputed targets)
- Online distillation (live teacher predictions)
Attribution
SNP analysis
Reference
- Targets table
- Updating batch-norm statistics
- Pretrained Borzoi trunk block
- Pretrained Borzoi head block
- GCP Batch execution (gcprunner)
Training, evaluation, SNP scoring, and ISM fold scripts support Google Cloud Batch
via --backend gcp. Gradient and distillation fold scripts support local and
Slurm execution.
There are no built-in GCP resources: before using --backend gcp, point the
runner at your own project and bucket with these environment variables (or the
matching --gcp_* flags). Commands fail fast with a clear error if a required
one is unset.
| Variable | Required | Meaning |
|---|---|---|
GCPRUNNER_PROJECT |
yes | GCP project that runs the Batch jobs (--gcp_project) |
GCPRUNNER_CACHE_PREFIX |
yes | gs:// prefix for content-addressed staged inputs |
GCPRUNNER_OUTPUT_PREFIX |
yes | gs:// prefix for run outputs (unless --gcp_output_dir is given) |
GCPRUNNER_REGION |
no | Batch region; defaults to us-central1 (--gcp_region) |
GCPRUNNER_IMAGE_BASE |
no | Image repo; defaults to <region>-docker.pkg.dev/<project>/baskerville/baskerville |
GCPRUNNER_IMAGE |
no | Image tag or full URI (--gcp_image); else the newest commit-<sha> image built from main |
export GCPRUNNER_PROJECT=my-gcp-project
export GCPRUNNER_CACHE_PREFIX=gs://my-bucket/cache
export GCPRUNNER_OUTPUT_PREFIX=gs://my-bucket/outputBuild the image from dockerfiles/baskerville.Dockerfile
and push it to that repo, or pass --gcp_image. Tag it
commit-$(git rev-parse HEAD): the default image and --gcp_branch only find
images tagged that way. See
docs/gcprunner.md §8 for the one-time
project, bucket, and image setup.
hound_snp_folds \
--backend gcp \
--gcp_output_dir gs://my-bucket/runs/2026-05-snp \
-q l4 \
-j 1024 \
-p 64 \
-f hg38.fa \
-t targets.txt \
params.json models/borzoi vcf.vcfTo run the latest built image from a specific git branch, pass its name (the image is resolved to that branch's newest built commit and digest-pinned):
--gcp_branch my-featureTo pin an arbitrary image instead, pass a full URI, a custom tag name, or a
numeric build id via --gcp_image (which overrides --gcp_branch):
--gcp_image my-tagFor complete setup instructions (IAM, registry configurations, etc.), see GCP Batch execution.
Pretrained weights are hosted in a public GCS bucket and distributed via
releases/. See each family's README for download commands and
verification instructions.
| Family | Replicates | Species | Input | Output | Tracks (human / mouse) |
|---|---|---|---|---|---|
| Cerberus | 8 | one model, human+mouse heads | 786 kb | 32 bp | 8,361 / 3,102 |
| Borzoi | 4 | separate human, mouse models | 524 kb | 32 bp | 7,611 / 2,608 |
| Borzoi Prime | 4 | separate human, mouse models | 524 kb | 16 bp | 5,431 / 1,774 |
Cerberus requires a GPU; Borzoi and Borzoi Prime also run on CPU.
If you use Cerberus, please cite:
@article{kelley2026cerberus,
title={Cerberus: bidirectional state space blocks improve accuracy and efficiency of regulatory sequence models},
author={Kelley, David R and Yuan, Han and Huang, Xingfan and Linder, Johannes},
journal={bioRxiv},
year={2026}
}If you use Borzoi or Borzoi Prime, please cite:
@article{linder2025predicting,
title={Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation},
author={Linder, Johannes and Srivastava, Divyanshi and Yuan, Han and Agarwal, Vikram and Kelley, David R},
journal={Nature Genetics},
pages={1--13},
year={2025},
publisher={Nature Publishing Group US New York},
doi={10.1038/s41588-024-02053-6}
}@article{linder2025predictingcelltype,
title={Predicting cell type-specific coverage profiles from DNA sequence},
author={Linder, Johannes and Yuan, Han and Kelley, David R},
journal={bioRxiv},
pages={2025--06},
year={2025},
publisher={Cold Spring Harbor Laboratory}
}This repository accompanies the Cerberus, Borzoi, and Borzoi Prime papers. Bug reports and questions are welcome as GitHub issues; pull requests are reviewed on a best-effort basis.
To run the checks CI runs:
pip install -e ".[dev]"
pytest -m "not slow"
uvx ruff@0.15 format --check .
npx prettier@3.8.2 --check .This project is licensed under the Apache License 2.0; see LICENSE. Third-party code and its licenses are listed in NOTICE.