Skip to content

Latest commit

 

History

18 Commits

Folders and files

Repository files navigation

baskerville

Pretrained models and PyTorch tools for Cerberus, along with its predecessors Borzoi and Borzoi Prime. These deep neural networks predict regulatory activity (e.g. chromatin accessibility and gene expression) from DNA sequence.

Cerberus replaces Borzoi's transformer with bidirectional state space (Hydra) blocks. It reads 786 kb of sequence and predicts 8,361 human and 3,102 mouse functional genomics tracks at 32 bp resolution, including RNA-seq, CAGE, DNase/ATAC, ChIP, CLIP, and 3′ RNA-seq. It is more accurate and cheaper to run than Borzoi. See releases/cerberus.

The TensorFlow version of baskerville, used by borzoi, has moved to baskerville-tf.

Installation

Requires Python 3.11+ and, for gene and SNP scoring, bedtools on PATH (e.g. apt install bedtools or conda install -c bioconda bedtools).

git clone https://github.com/calico/baskerville.git
cd baskerville
pip install .

Cerberus needs an NVIDIA GPU and the cuda extra (mamba-ssm). For a development setup, install in editable mode with the dev extras:

pip install -e ".[dev,cuda]"   # drop ,cuda on CPU-only hosts (Borzoi only)

The *_folds.py scripts run jobs locally by default, or on GCP Batch with --backend gcp (see GCP configuration). --backend slurm needs slurmrunner, which is not publicly available.

The hdwig extra (calico/hdwig, not yet public) is needed only to read .hw coverage files in hound_data.

Quickstart

On a GPU host, download one pretrained Cerberus replicate and check that it reproduces the published forward pass. hound_verify reads the weights from the clone, so install it editable (pip install -e ".[cuda]"):

cd releases/cerberus
./download.sh 0           # -> models/f0c0/model_best.pth
hound_verify --family cerberus

See Model releases for all weights, and the guides below for training and scoring.

Training, Attribution & Evaluation

For detailed instructions on dataset construction, training, and evaluating models, see:

Training

Attribution

SNP analysis

Reference

GCP Batch Execution

Training, evaluation, SNP scoring, and ISM fold scripts support Google Cloud Batch via --backend gcp. Gradient and distillation fold scripts support local and Slurm execution.

GCP configuration

There are no built-in GCP resources: before using --backend gcp, point the runner at your own project and bucket with these environment variables (or the matching --gcp_* flags). Commands fail fast with a clear error if a required one is unset.

Variable Required Meaning
GCPRUNNER_PROJECT yes GCP project that runs the Batch jobs (--gcp_project)
GCPRUNNER_CACHE_PREFIX yes gs:// prefix for content-addressed staged inputs
GCPRUNNER_OUTPUT_PREFIX yes gs:// prefix for run outputs (unless --gcp_output_dir is given)
GCPRUNNER_REGION no Batch region; defaults to us-central1 (--gcp_region)
GCPRUNNER_IMAGE_BASE no Image repo; defaults to <region>-docker.pkg.dev/<project>/baskerville/baskerville
GCPRUNNER_IMAGE no Image tag or full URI (--gcp_image); else the newest commit-<sha> image built from main
export GCPRUNNER_PROJECT=my-gcp-project
export GCPRUNNER_CACHE_PREFIX=gs://my-bucket/cache
export GCPRUNNER_OUTPUT_PREFIX=gs://my-bucket/output

Build the image from dockerfiles/baskerville.Dockerfile and push it to that repo, or pass --gcp_image. Tag it commit-$(git rev-parse HEAD): the default image and --gcp_branch only find images tagged that way. See docs/gcprunner.md §8 for the one-time project, bucket, and image setup.

Example Usage

hound_snp_folds \
    --backend gcp \
    --gcp_output_dir gs://my-bucket/runs/2026-05-snp \
    -q l4 \
    -j 1024 \
    -p 64 \
    -f hg38.fa \
    -t targets.txt \
    params.json models/borzoi vcf.vcf

To run the latest built image from a specific git branch, pass its name (the image is resolved to that branch's newest built commit and digest-pinned):

    --gcp_branch my-feature

To pin an arbitrary image instead, pass a full URI, a custom tag name, or a numeric build id via --gcp_image (which overrides --gcp_branch):

    --gcp_image my-tag

For complete setup instructions (IAM, registry configurations, etc.), see GCP Batch execution.

Model releases

Pretrained weights are hosted in a public GCS bucket and distributed via releases/. See each family's README for download commands and verification instructions.

Family Replicates Species Input Output Tracks (human / mouse)
Cerberus 8 one model, human+mouse heads 786 kb 32 bp 8,361 / 3,102
Borzoi 4 separate human, mouse models 524 kb 32 bp 7,611 / 2,608
Borzoi Prime 4 separate human, mouse models 524 kb 16 bp 5,431 / 1,774

Cerberus requires a GPU; Borzoi and Borzoi Prime also run on CPU.

Citation

If you use Cerberus, please cite:

@article{kelley2026cerberus,
  title={Cerberus: bidirectional state space blocks improve accuracy and efficiency of regulatory sequence models},
  author={Kelley, David R and Yuan, Han and Huang, Xingfan and Linder, Johannes},
  journal={bioRxiv},
  year={2026}
}

If you use Borzoi or Borzoi Prime, please cite:

@article{linder2025predicting,
  title={Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation},
  author={Linder, Johannes and Srivastava, Divyanshi and Yuan, Han and Agarwal, Vikram and Kelley, David R},
  journal={Nature Genetics},
  pages={1--13},
  year={2025},
  publisher={Nature Publishing Group US New York},
  doi={10.1038/s41588-024-02053-6}
}
@article{linder2025predictingcelltype,
  title={Predicting cell type-specific coverage profiles from DNA sequence},
  author={Linder, Johannes and Yuan, Han and Kelley, David R},
  journal={bioRxiv},
  pages={2025--06},
  year={2025},
  publisher={Cold Spring Harbor Laboratory}
}

Contributing

This repository accompanies the Cerberus, Borzoi, and Borzoi Prime papers. Bug reports and questions are welcome as GitHub issues; pull requests are reviewed on a best-effort basis.

To run the checks CI runs:

pip install -e ".[dev]"
pytest -m "not slow"
uvx ruff@0.15 format --check .
npx prettier@3.8.2 --check .

License

This project is licensed under the Apache License 2.0; see LICENSE. Third-party code and its licenses are listed in NOTICE.

About

PyTorch models of DNA sequence to regulatory activity

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages