From be804044a6696acaea04452e2d4e56de136fae48 Mon Sep 17 00:00:00 2001 From: Canberk Pitirli Date: Thu, 1 Oct 2026 17:55:33 +0300 Subject: [PATCH] rework the readme --- README.md | 380 +++++++++++++++++++++++++++++++++--------------------- 1 file changed, 236 insertions(+), 144 deletions(-) diff --git a/README.md b/README.md index ae26266..f2813c7 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,87 @@ +
+ # FastNN +**A deep learning library in Rust, with CUDA kernels for the parts that matter.** + [![CI](https://github.com/CanReader/FastNN/actions/workflows/ci.yml/badge.svg)](https://github.com/CanReader/FastNN/actions/workflows/ci.yml) +[![crates.io](https://img.shields.io/crates/v/fastnn.svg)](https://crates.io/crates/fastnn) +[![docs.rs](https://img.shields.io/docsrs/fastnn)](https://docs.rs/fastnn) +[![MSRV](https://img.shields.io/badge/rustc-1.87+-blue.svg)](https://github.com/CanReader/FastNN/blob/master/Cargo.toml) +[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE) + +[Quick start](#quick-start) · +[Examples](#examples) · +[API docs](https://docs.rs/fastnn) · +[Releases](https://github.com/CanReader/FastNN/releases) · +[Contributing](CONTRIBUTING.md) + +
+ +FastNN is a from-scratch deep learning library: tensors with reverse-mode +autograd, the usual layers, transformers with KV-cached generation, modern +optimizers, data loading, resumable checkpoints, and safetensors interchange. +It runs on the CPU out of the box and on NVIDIA GPUs with one feature flag. + +It is built to be read. Every file is one idea and none of them are long, so you +can follow a `loss.backward()` call from the training loop down to the kernel +and back without getting lost. + +* **Readable.** 86 source files, about 12k lines. Forward ops and their backward + rules live in mirrored files, one per category. +* **Verified.** Every backward rule is checked against finite differences of its + own forward, and every CUDA kernel against the CPU path. +* **Built for real runs.** Exact resume including optimizer state, NaN tracing + to the op that produced it, frozen layers that skip their gradient entirely. + +## Contents + +* [Features](#features) +* [Installation](#installation) +* [Quick start](#quick-start) +* [Examples](#examples) +* [Platform support](#platform-support) +* [How it fits together](#how-it-fits-together) +* [Long training runs](#long-training-runs) +* [Extending FastNN](#extending-fastnn) +* [Error handling](#error-handling) +* [Testing](#testing) +* [CUDA](#cuda) +* [Project layout](#project-layout) +* [Limitations](#limitations) +* [Contributing](#contributing) +* [License](#license) + +## Features + +| Area | What is in it | +|---|---| +| **Tensors** | Broadcasting arithmetic, matmul (batched, transposed forms), reductions, indexing, views, `cat`/`stack`/`narrow`, CPU or CUDA storage | +| **Autograd** | Reverse mode with no tape and no global state, `no_grad`, `detect_anomaly`, gradient clipping by norm or value | +| **Layers** | `Linear`, `Conv1d`, `Conv2d` (dilation, groups, depthwise), `ConvTranspose1d/2d`, pooling, `BatchNorm2d`, `LayerNorm`, `RMSNorm`, `Dropout`, `Embedding`, `LSTM`, `GRU` | +| **Transformers** | `MultiHeadAttention` with causal and padding masks, encoder and decoder stacks, full encoder-decoder `Transformer` | +| **Generation** | `KvCache` for incremental decoding, `Sampler` (temperature, top-k, top-p, repetition penalty), `BeamSearch` | +| **Optimizers** | `SGD`, `Adam`, `AdamW`, `RMSprop`, `Adagrad`, `Adadelta`, `RAdam`, `Lion`, plus `Lookahead` and `Ema` wrappers | +| **Schedules** | `Warmup`, `CosineAnnealing`, `OneCycle`, `StepDecay`, `Constant` | +| **Losses** | Cross-entropy (label smoothing, class weights, ignore index), MSE, MAE, BCE, Huber, KL, NLL, focal, dice, Gaussian and Poisson NLL, triplet, contrastive, InfoNCE | +| **Data** | `Dataset` trait, shuffling `DataLoader` with device placement, built-in MNIST | +| **Checkpoints** | Native weights, full training state (weights + optimizer + step), safetensors read and write with F16/BF16 widening | + +## Installation + +```bash +cargo add fastnn # CPU only, no CUDA toolkit needed +cargo add fastnn --features cuda # with the CUDA kernels +``` -A deep learning library in Rust, with CUDA kernels for the parts that matter. +| Requirement | Version | +|---|---| +| Rust | 1.87 or newer | +| CUDA toolkit (only with `--features cuda`) | 11.8 or newer, for GPUs from Turing to Hopper (compute capability 7.5 to 9.0) | -Tensors with autograd, the usual layers, optimizers, losses, data loading, -checkpoints, KV-cached text generation, and safetensors interchange — small -enough to read end to end, and it actually trains. +## Quick start + +A complete MNIST training loop: ```rust use fastnn::prelude::*; @@ -35,52 +110,71 @@ save(&model, "mnist.fdl")?; ``` No gradient-mode flags, no borrow dance around the optimizer. That is the whole -training loop. - -## Getting started +training loop. To train on a GPU, move the model and the loader: -As a dependency: - -```bash -cargo add fastnn # CPU only — no CUDA toolkit needed -cargo add fastnn --features cuda # with the CUDA kernels +```rust +let device = Device::best(); // GPU if there is one, else CPU +model.to_device(device); +let loader = DataLoader::new(&train, 128).to_device(device); ``` -Running the examples from a checkout: +## Examples ```bash -# CPU only — no CUDA toolkit needed -cargo run --example simple_mlp --release - -# With a GPU -cargo run --example mnist_mlp --release --features cuda +cargo run --release --example simple_mlp +cargo run --release --example mnist_cnn --features cuda ``` -Always use `--release`. Debug builds are roughly 50× slower. +Always use `--release`. Debug builds are roughly 50x slower. | Example | What it shows | |---|---| -| `simple_mlp` | The smallest complete training loop (XOR) | -| `mnist_mlp` | Dataset loading, batching, evaluation, checkpointing | +| `simple_mlp` | The smallest complete training loop, an MLP learning XOR | +| `mnist_mlp` | Dataset download, batching, evaluation, checkpointing | | `mnist_cnn` | Convolutions, batch norm, pooling, an LR schedule | -| `char_lm` | A GPT-style transformer, trained from scratch, that generates text | +| `char_lm` | A small GPT trained from scratch that generates text, with resume on Ctrl-C | +| `finetune` | Pretrain, freeze the backbone, train a new head | +| `gan` | Two networks trained against each other | +| `vae` | A variational autoencoder on 2-D data | +| `reinforce` | Policy-gradient reinforcement learning on a bandit | +| `multi_task` | One shared trunk, two heads, one combined loss | -`char_lm` takes a corpus path, or trains on a small embedded one: +`char_lm` trains on a small embedded corpus by default, or on any text file: ```bash curl -o shakespeare.txt \ https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt -cargo run --example char_lm --release -- shakespeare.txt +cargo run --release --example char_lm -- shakespeare.txt ``` -### Prebuilt examples +### Prebuilt binaries -Every [release](https://github.com/CanReader/FastNN/releases) ships the -example programs prebuilt for Linux (x86_64, ARM64), macOS (Intel, Apple -Silicon), and Windows, plus a Linux build with CUDA. Download the archive for -your platform and run `simple_mlp`, `char_lm`, or `mnist_cnn` directly, no Rust -toolchain needed. The CUDA build links against the CUDA 13 runtime, so it needs -that installed. +Every [release](https://github.com/CanReader/FastNN/releases) ships all of the +examples prebuilt, so you can try FastNN without installing Rust. Download the +archive for your platform, extract it, and run any example directly. + +| Platform | Archive | +|---|---| +| Linux x86_64 | `fastnn--x86_64-linux.tar.gz` | +| Linux ARM64 | `fastnn--aarch64-linux.tar.gz` | +| Linux x86_64 with CUDA | `fastnn--x86_64-linux-cuda.tar.gz` | +| macOS Apple Silicon | `fastnn--aarch64-macos.tar.gz` | +| macOS Intel | `fastnn--x86_64-macos.tar.gz` | +| Windows x86_64 | `fastnn--x86_64-windows.zip` | + +Each release includes a `SHA256SUMS` file and a build provenance attestation +for every archive, which you can check with +`gh attestation verify --repo CanReader/FastNN`. The CUDA build links +against the CUDA 13 runtime, so it needs that installed. + +## Platform support + +| Platform | CPU | CUDA | +|---|---|---| +| Linux x86_64 | Tested in CI | Built in CI, parity tested on real hardware | +| Linux ARM64 | Built for releases | Not supported | +| macOS (Intel, Apple Silicon) | Tested in CI | Not available on macOS | +| Windows x86_64 | Tested in CI | Should build, not tested | ## How it fits together @@ -98,23 +192,17 @@ that installed. ``` Everything starts on the CPU. `model.to_device(device)` moves a model, -`DataLoader::to_device` moves its batches, and mixing the two panics rather than -copying behind your back. +`DataLoader::to_device` moves its batches, and mixing devices panics rather +than copying behind your back. -```rust -let device = Device::best(); // GPU if there is one, else CPU -model.to_device(device); -let loader = DataLoader::new(&dataset, 64).to_device(device); -``` - -### Two ideas worth knowing +Two design decisions shape the rest of the library. **The graph is the tensors.** There is no tape and no global state. A tensor produced by a differentiable op carries a node naming the rule that made it and the tensors it consumed, so `loss.backward()` just walks that structure. Drop the loss and the graph frees itself. A tensor with no node is plain data, so there is -no `requires_grad` flag to keep in sync — wrap inference in `no_grad(|| ...)` -when you want to skip building a graph at all. +no `requires_grad` flag to keep in sync. Wrap inference in `no_grad(|| ...)` to +skip building a graph at all. **Parameters are shared slots.** `Param` is a handle, not a value. The model and the optimizer hold handles to the same weight and the same gradient, which is why @@ -122,15 +210,15 @@ the optimizer hold handles to the same weight and the same gradient, which is wh without special support: use the same handle twice and both paths' gradients add up on their own. -## Running unattended +## Long training runs -Four things a long run needs, none of which the training loop above shows. +Four things a long, unattended run needs that the quick start does not show. -**Resume where you stopped.** `save()` writes weights, which is what you want for -a finished model. Resuming needs the optimizer too — its momentum, Adam's two -moment estimates, and the step count that drives bias correction. Restore weights -alone and the optimizer restarts at step 1, so the first update after resuming -lands far harder than it should. +**Resume exactly where you stopped.** `save()` writes weights, which is right for +a finished model. Resuming also needs the optimizer: momentum, Adam's moment +estimates, and the step count that drives bias correction. Restore weights alone +and the optimizer restarts at step 1, so the first update after resuming lands +far harder than it should. ```rust let mut step = load_training(&model, &mut opt, "run.fdl").unwrap_or(0); @@ -142,13 +230,12 @@ while step < total { } ``` -Loading a plain checkpoint with `load_training` is an error rather than a silent -optimizer reset. `char_lm` does this, so Ctrl-C costs at most a few hundred steps. +Loading a plain checkpoint with `load_training` is an error, not a silent +optimizer reset. **Find the NaN at the op that made it.** One non-finite gradient becomes a -non-finite weight, and every activation downstream is NaN from then on — the loss -only *prints* as NaN some steps later, by which point the checkpoint is poisoned -too and the culprit is long gone. +non-finite weight, and every activation downstream is NaN from then on. By the +time the loss prints NaN the checkpoint is poisoned and the culprit is long gone. ```rust detect_anomaly(|| { @@ -161,17 +248,17 @@ if !opt.gradients_are_finite() { continue; } // cheap guard: skip the batch Both read every gradient, so they are debugging and guard-rail tools, not something to leave on in a healthy loop. -**Freeze a backbone.** Frozen parameters hand out a detached tensor, so the -backward pass stops there — no gradient is computed only to be discarded. +**Freeze a backbone.** A frozen parameter hands out a detached tensor, so the +backward pass stops there and no gradient is computed only to be thrown away. ```rust backbone.freeze(); let mut opt = Adam::new(model.parameters(), 1e-3); // updates only the head ``` -**Report a bad shape instead of dying.** Ordinary ops panic, which is right -inside a model. At the edge of an app — a shape from a config file, a batch from -an upload — use the `try_` variants. +**Report a bad shape instead of crashing.** Ordinary ops panic, which is right +inside a model. At the edge of an application, where a shape comes from a config +file or an upload, use the `try_` variants. ```rust let y = x.try_matmul(&w)?; // Error::Shape, not a panic @@ -182,57 +269,10 @@ table.try_index_select(&ids)?; They validate and then delegate, so there is still exactly one implementation of each operation, and they build the same graph. -## Layout - -Every file is one idea, and none are long. 86 source files, ~11k lines. - -``` -src/ - tensor/ the array type and everything you can do to it - core.rs Tensor: shape, storage, graph link - checked.rs try_* variants that report instead of panicking - shape.rs strides, broadcasting, index math - storage.rs the bytes, on one device or the other - device.rs Device::cuda(0) -> Result - init.rs zeros, randn, kaiming, xavier, ... - ops/ one file per category - arith unary activation matmul reduce view index - conv (im2col/col2im) pool norm - - autograd/ reverse-mode differentiation - node.rs Backward trait, graph nodes, gradient slots - engine.rs the reverse pass - mode.rs no_grad - anomaly.rs detect_anomaly - ops/ one backward rule per forward op, same file names - - nn/ layers, all implementing Module - module.rs param.rs sequential.rs linear.rs - conv.rs conv1d.rs conv_transpose.rs pooling.rs - norm.rs activation.rs dropout.rs shape.rs embedding.rs - attention.rs transformer.rs decoder.rs rnn.rs - cache.rs (kv cache) sample.rs (temperature/top-k/top-p) beam.rs - loss.rs metric_losses.rs - - optim/ sgd adam rmsprop adagrad radam lion lookahead ema - schedule.rs state.rs - data/ dataset.rs loader.rs mnist.rs - serialize/ checkpoint.rs training.rs (weights + optimizer) - safetensors.rs (interchange, reads F32/F16/BF16) half.rs - cuda/ ffi.rs (raw bindings) kernels.rs (safe wrappers) buffer.rs - rng.rs error.rs lib.rs - -cuda/kernels.cu all GPU kernels -``` - -`tensor/ops/` and `autograd/ops/` mirror each other file for file: the forward -for `relu` is in `tensor/ops/activation.rs`, its derivative in -`autograd/ops/activation.rs`. - -## Extending it +## Extending FastNN **A new layer.** Implement `forward`, plus `named_parameters` if it has weights. -Device placement, parameter counting, and gradient zeroing come from those. +Device placement, parameter counting, and gradient zeroing all come from those. ```rust struct Residual { inner: Linear } @@ -247,8 +287,9 @@ impl Module for Residual { } ``` -**A new op.** Write the forward in `tensor/ops/`, the rule in `autograd/ops/`, -attach it with `with_grad`, and add a check to `tests/gradcheck.rs`. +**A new op.** Write the forward in `tensor/ops/`, the backward rule in +`autograd/ops/`, attach it with `with_grad`, and add a case to +`tests/gradcheck.rs`. ```rust // tensor/ops/unary.rs @@ -266,15 +307,17 @@ impl Backward for SoftplusBackward { } ``` -The rule's gradients must come back in the same order as the inputs. Save -detached tensors — a saved value that still carries its own node would pin the -graph that produced it. +The rule returns gradients in the same order as the inputs. Save detached +tensors: a saved value that still carries its own node would pin the graph that +produced it. **A new dataset.** Implement `Dataset` and hand it to a `DataLoader`. Items are written into caller-provided slices rather than returned as tensors, so batching 60,000 images does not build 60,000 throwaway ones. -## Errors +[CONTRIBUTING.md](CONTRIBUTING.md) has the full checklist for adding an op. + +## Error handling Tensor maths panics on shape and device mistakes. Those are bugs in the calling code, like indexing past the end of a slice, and threading `Result` through every @@ -283,67 +326,116 @@ code, like indexing past the end of a slice, and threading `Result` through ever `Result` is for what genuinely fails at runtime: ```rust -let device = Device::cuda(0)?; // no GPU -let data = Mnist::load(Split::Train)?; // download or parse failed -load(&model, "model.fdl")?; // missing file, or a shape that moved +let device = Device::cuda(0)?; // no GPU +let data = Mnist::load(Split::Train)?; // download or parse failed +load(&model, "model.fdl")?; // missing file, or a shape that moved ``` ## Testing ```bash +cargo test # everything cargo test --test gradcheck # every backward rule vs finite differences cargo test --test robustness # resume, anomalies, freezing -cargo test --test generation # kv-cached decoding, masks, beam, sampling +cargo test --test generation # kv-cached decoding, masks, beam search, sampling cargo test --test convergence # every layer family and optimizer actually learns cargo test --test losses # loss values against hand-computed references cargo test --features cuda --test cuda_parity -- --test-threads=1 # CPU vs GPU -cargo test # everything cargo bench # throughput ``` -`gradcheck` checks every backward rule against central finite differences of its -own forward. It is the reason the autograd layer can be trusted; a new rule -without a test there is not finished. +`gradcheck` compares every backward rule with central finite differences of its +own forward. It is why the autograd layer can be trusted, and a new rule without +a case there is not finished. -`cuda_parity` runs each op on both devices and compares. It skips itself with a -note when there is no GPU, so a CPU-only machine still gets a green run. It found -a real bug in the softmax kernel: the block reduction assumed a power-of-two -thread count and silently dropped the tail of every row whose width was not one — -which included MNIST's ten classes and any odd sequence length. +`cuda_parity` runs each op on both devices and compares the results. It skips +itself when there is no GPU, so a CPU-only machine still gets a green run. It has +already caught a real bug: the softmax kernel's block reduction assumed a +power-of-two thread count and silently dropped the tail of every row whose width +was not one, which included MNIST's ten classes. + +Every pull request runs formatting, clippy, the full suite on Linux, macOS, and +Windows, an MSRV build, docs, packaging, and a CUDA compile. A nightly job reruns +everything in release mode on stable and beta and trains the examples end to end. ## CUDA The `cuda` feature is opt-in, so a plain `cargo add fastnn` never needs the -toolkit. Enabling it needs the NVIDIA CUDA Toolkit; set `CUDA_PATH` or -`CUDA_HOME` if it is not in the default location. `build.rs` compiles -`cuda/kernels.cu` with `nvcc` and links `cudart`, `cublas`, and `curand`. -Compute capabilities 7.5 through 9.0 (Turing through Hopper). If `nvcc` rejects -your system compiler as too new, point `FASTNN_NVCC_CCBIN` at one it accepts. +toolkit. With the feature on, `build.rs` compiles `cuda/kernels.cu` with `nvcc` +and links `cudart`, `cublas`, and `curand`. -Without the feature, `cuda/stubs.c` supplies the symbols, `Device::cuda(0)` -returns an error, and everything runs on the CPU. +* Set `CUDA_PATH` or `CUDA_HOME` if the toolkit is not in the default location. +* If `nvcc` rejects your system compiler as too new, point `FASTNN_NVCC_CCBIN` + at one it accepts. +* Without the feature, `cuda/stubs.c` supplies the symbols, `Device::cuda(0)` + returns an error, and everything runs on the CPU. Two things carry most of the GPU performance. Matrix multiplication goes through cuBLAS, and the two transposed forms the backward pass needs (`matmul_nt`, -`matmul_tn`) use a transpose flag rather than building a transposed copy. GPU +`matmul_tn`) use a transpose flag instead of building a transposed copy. GPU allocations come from a size-keyed free list, so the many short-lived temporaries a training step creates cost a hash lookup instead of a driver round trip. +## Project layout + +``` +src/ + tensor/ the array type and everything you can do to it + core.rs Tensor: shape, storage, graph link + checked.rs try_* variants that report instead of panicking + shape.rs strides, broadcasting, index math + storage.rs the bytes, on one device or the other + device.rs Device::cuda(0) -> Result + init.rs zeros, randn, kaiming, xavier, ... + ops/ arith unary activation matmul reduce view + index conv pool norm + + autograd/ reverse-mode differentiation + node.rs Backward trait, graph nodes, gradient slots + engine.rs the reverse pass + mode.rs no_grad + anomaly.rs detect_anomaly + ops/ one backward rule per forward op, same file names + + nn/ layers, all implementing Module + linear conv conv1d conv_transpose pooling norm activation + dropout embedding shape attention transformer decoder rnn + cache (kv cache) sample (sampling) beam (beam search) + loss metric_losses module param sequential + + optim/ sgd adam rmsprop adagrad radam lion lookahead + ema schedule state + data/ dataset loader mnist + serialize/ checkpoint training (weights + optimizer) + safetensors half (F16/BF16 conversion) + cuda/ ffi (raw bindings) kernels (safe wrappers) buffer + +cuda/kernels.cu every GPU kernel +``` + +`tensor/ops/` and `autograd/ops/` mirror each other file for file: the forward +for `relu` is in `tensor/ops/activation.rs` and its derivative is in +`autograd/ops/activation.rs`. + ## Limitations -- `f32` compute. Checkpoints stored as `F16`/`BF16` load fine (widened on - read), but the maths runs in single precision. -- `LSTM` and `GRU` are built from ordinary differentiable ops, one graph node per - gate per timestep. Correct, and fine for short sequences — use a transformer - for long ones. -- Tensors are always contiguous. `permute` and `expand` write a new buffer rather +* Compute is `f32`. Checkpoints stored as F16 or BF16 load fine and are widened + on read, but the maths runs in single precision. +* `LSTM` and `GRU` are built from ordinary differentiable ops, one graph node per + gate per timestep. Correct, and fine for short sequences; use a transformer for + long ones. +* Tensors are always contiguous. `permute` and `expand` write a new buffer rather than returning a strided view. ## Contributing -Bug reports and PRs are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for how -the code is laid out, what CI checks, and how to add an op. +Bug reports, ideas, and pull requests are all welcome. [CONTRIBUTING.md](CONTRIBUTING.md) +covers the setup, the conventions, what CI checks, and how to add an op. Issues +labelled [`good first issue`](https://github.com/CanReader/FastNN/labels/good%20first%20issue) +are a good place to start. + +Please report security problems privately, as described in [SECURITY.md](SECURITY.md). ## License -MIT. +FastNN is released under the [MIT License](LICENSE).