Educational notebooks exploring production-grade agentic systems and LLMs pipelines from first principles. Learn the core patterns behind tools like Claude Code, Cursor's agent mode, and autonomous research assistants by implementing them yourself.
These notebooks teach you to build the scaffolding that transforms an LLM from a text generator into a truly useful everyday tool.
| Notebooks | Blog post | Content |
|---|---|---|
| 01 - Basic Agentic Harness.ipynb | Building a Basic Agentic Harness | Start here to understand the fundamentals |
| 02 - Small Language Model.ipynb | Build a (small) language model by counting | Build an n-gram language model from scratch and generate text |
| 03 - Advanced Agentic Harness.ipynb | Building an Advanced Agentic Harness | Production-grade patterns: DAG orchestration, memory, verification, and multi-agent roles |
| None | Self hosting LLMs with Llama.cpp | Run LLMs on your own hardware: install llama.cpp, serve models over an OpenAI-compatible API, explore GGUF files, and choose the right quantization |
| 04 - Evaluating Agentic Harnesses.ipynb | Evaluating Your Agentic Harness | Eval suites, cost and latency measurement, failure modes, and three measured upgrades |
| 05 - MCP Server.ipynb | An MCP Server from Scratch | Build an MCP server from scratch: OpenAlex data pipeline, SQLite + FTS5, raw JSON-RPC over stdio, and a from-scratch client |
| None | LLMs Are the New Wikipedia | Every old criticism is new again: why the reliability debate around LLMs mirrors Wikipedia's early years, and why quantitative evaluation is what settles it |
| 06 - Local Inference with vLLM.ipynb | Local LLM Inference at Scale with vLLM | Run an open-weight model with vLLM: choosing a model from memory bandwidth, continuous batching, prefix caching, and structured outputs over thousands of abstracts |
- Start with Notebook 1 (
01 - Basic Agentic Harness.ipynb) to understand the core concepts - Continue with Notebook 2 (
02 - Small Language Model.ipynb) to see how language models work from the ground up - Then Notebook 3 (
03 - Advanced Agentic Harness.ipynb) to upgrade the harness with production-grade patterns - Continue with Notebook 4 (
04 - Evaluating Agentic Harnesses.ipynb) to measure whether any of it actually works - Continue with Notebook 5 (
05 - MCP Server.ipynb) to hand real data to an agent through a from-scratch MCP server - Finish with Notebook 6 (
06 - Local Inference with vLLM.ipynb) to run an open-weight model on your own GPU and process thousands of documents with it
Notebooks 3 and 4 share a module, d4sci_harness.py — notebook 3 builds it
step by step, notebook 4 imports it. See The harness as a module below.
Notebook 5 likewise keeps its executable pieces in mcp_server/ — see
The MCP server as scripts below. Notebook 6 reads the OpenAlex
database that notebook 5 builds, and needs a Linux machine with an NVIDIA GPU — see
Hardware for notebook 6 below.
- Almost no API keys — Only notebook 1 requires a hosted model. Notebooks 3 and 4 switch to a rule-based mock backend with one line, notebooks 2 and 5 don't call an LLM at all, and notebook 6 runs an open-weight model on your own GPU
- Production-ready patterns — Learn the same techniques used in Claude Code, Cursor, and Devin
- Hands-on implementation — Build everything from scratch to understand every design decision
- Measured, not asserted — The eval suite in notebook 4 turns "it worked once" into pass rates, cost, and failure-mode distributions, and notebook 6 measures every inference-engine claim on the hardware in front of you
Build a minimal but complete harness from scratch — the foundation of any autonomous agent system.
Core concepts:
- The five components of a harness's core state (goal, trace, memory, budget, status)
- Implementing a control loop that drives an LLM through multi-step tasks
- Defining typed tools the LLM can call safely
- Validating LLM-proposed actions against schemas before execution
- Inspecting execution traces for debugging
What you'll build: A single-agent harness that solves multi-step research tasks by repeatedly composing context, asking the LLM what to do next, executing tool calls, and updating state until the goal is met.
Build a complete language model from first principles — counting, not neural networks — to understand what every LLM is really doing under the hood.
Core concepts:
- The language modeling task: estimating
P(next word | previous words) - Building an n-gram (4-gram) model over the WikiText-103 corpus (~103M words)
- Why text is sparse: word-frequency distributions and Zipf's law
- Counting three-word contexts and the single-continuation problem
- Autoregressive generation — feeding the output back in to predict the next token
- Temperature and sampling — greedy decoding vs. probabilistic sampling
What you'll build: A 4-gram model trained on Wikipedia that generates text from a prompt, with a temperature switch that mirrors the same decoding knob exposed by modern LLMs.
Upgrade every component toward production-grade systems like Claude Code, Devin, or modern research agents.
Advanced topics:
- Typed tools with Pydantic — Auto-generated JSON schemas and robust validation
- DAG orchestration — Parallel execution of independent tasks instead of sequential processing
- Multi-tier memory — Working, episodic, and semantic memory with retrieval
- Verification hierarchy — Cheap deterministic checks first, LLM-as-judge only when they pass
- Multi-agent roles — Planner / Worker / Critic specialization for robustness
- Multi-dimensional budgeting — Graceful degradation under token, time, and cost constraints
- Error taxonomy — Transient failures retry in place; missing information escalates to a re-plan
- Structured tracing — Full observability and replay capability
What you'll build: A three-city comparison agent whose plan contains nine independent fetches — executed in parallel under a concurrency cap — plus a final aggregation step and verifiable output, with trace plots for per-step latency, budget pressure, and tokens by role.
One good demo proves the harness can work. An eval suite proves it usually works — and prices the times it doesn't.
Core concepts:
- Eval suites vs. demos — the smallest useful loop: run every task, record pass/fail, tokens, cost, and status
- Adversarial tasks — one task asks for a city that doesn't exist, to exercise the recovery path rather than the happy path
- Reading results like a CI dashboard — failure-mode distribution (
failed_executevs.failed_verify), regression baselines, and which verification tier produced each verdict - Trace-derived plots — cost per task by outcome, per-role latency against actual wall clock (the gap is the parallelism payoff), tokens by role, and budget-pressure trajectories against the degradation threshold
- Honest instruments — why the token and cost columns are estimates derived from plan shape, and what that does and doesn't buy you
Closing three gaps, each shipped with its own mini-benchmark:
- Re-planning on failure — an unknown entity amends the plan instead of aborting the run
- Real embeddings —
all-MiniLM-L6-v2vs. Jaccard retrieval, benchmarked over a labeled query set rather than swapped on faith - Specialized workers — a
FetcherAgent/WriterAgentsplit routed by capability, because millisecond lookups and multi-second LLM calls do not deserve the same concurrency and retry policy
What you'll build: A four-task eval suite run against the harness module, producing a metrics table and four diagnostic plots built entirely from the structured trace the orchestrator already emits — no extra instrumentation.
Build a complete MCP server from scratch — raw JSON-RPC over stdio, no SDK — that gives an agent queryable access to a real bibliographic database.
Core concepts:
- Data acquisition done right — pull a laptop-sized subset of OpenAlex via cursor pagination and
select=field trimming, caching the raw JSON separately from the database - A normalized SQLite schema — 9 tables mirroring OpenAlex's entity graph, plus BM25 full-text search via FTS5, because the joins are the point
- EDA as an audit — know the data's truncations, gaps, and traps (double-counting joins, closed-world citations, right-censored years) before the agent finds them
- The MCP wire format — newline-delimited JSON-RPC 2.0 over stdio: discovery, per-request
_meta, result envelopes, cache hints, and error classification, all made visible - The latest protocol revision — the server is native to MCP
2026-07-28, the current spec version, which made the core stateless:server/discoverreplaces theinitializehandshake, and every request carries its own version and capabilities in_meta - Defense in depth — read-only connections at the capability level, server-side argument validation, query deadlines, and expiring pagination handles
- A from-scratch client — exercise the server one exchange at a time, including hostile calls and version rejection, so you see every byte on the wire
What you'll build: A five-tool, one-resource MCP server over an OpenAlex subset — list_tables, describe_table, query with handle-based pagination, fetch_page, and BM25 search_works — plus the launch configuration to plug it into a real MCP host.
Stop treating the model as a remote service: run an open-weight LLM yourself with vLLM, the engine most self-hosted deployments run on, and measure what it does.
Core concepts:
- Why naive generation wastes a GPU — decode is memory-bandwidth bound, so batching is nearly free until the KV cache runs out
- Choosing a model for the machine — tokens/s ≈ bandwidth ÷ bytes read per token, and why a Mixture-of-Experts model in FP8 (
Qwen3.5-35B-A3B-FP8) beats dense models on a 128 GB DGX Spark. - Anatomy before loading — computing KV cache cost per token from the model config, including Qwen3.5's hybrid full/linear attention layers
- The offline engine — what
LLM(...)does at startup (weight loading, memory profiling,torch.compile, CUDA graphs) and where the memory goes - Continuous batching, measured — throughput against batch size on the same requests
- Prefix caching, measured — shared system prompts reused block by block, and why a prefix shorter than one block is never shared
- Structured outputs — grammar-constrained decoding from a Pydantic schema, against an unconstrained baseline that only asks for JSON
- From notebook to server — the same engine behind
vllm serveand an OpenAI-compatible API
What you'll build: A structured extraction over 2,000 "scaling laws" abstracts sampled from notebook 5's OpenAlex database, turned into an analysis of which fields use the term, whether the model agrees with OpenAlex's own topic labels, and how often a "scaling law" is actually a power law.
Everything notebook 3 builds step by step — typed tools, the plan DAG, the parallel executor, multi-tier memory, the verification hierarchy, multi-dimensional budgets, structured tracing, and the Orchestrator that composes them — also lives in d4sci_harness.py as an importable module. Notebook 3 constructs it; notebook 4 imports it, which is how you would consume it in a real project:
import d4sci_harness as dh
from d4sci_harness import TOOLS, MemoryStore, Orchestrator
llm = dh.set_provider("anthropic") # or "mock" for offline runs
store = MemoryStore()
orch = Orchestrator(provider=llm, tools=TOOLS, memory=store)
result = await orch.run("Compare Paris and Tokyo.", ["paris", "tokyo"])
print(result.status, result.budget.tokens_used, result.verdict.tier)| Subsystem | Key names |
|---|---|
| LLM providers | LLMProvider, AnthropicProvider, MockProvider, set_provider |
| Typed tools | TypedTool, TOOLS, CityArgs, AggregateArgs |
| Plan as a DAG | PlanDAG, PlanNode, NodeStatus, PlanDAG.plot |
| Parallel execution | execute_dag, MAX_CONCURRENT, MAX_NODE_RETRIES |
| Memory | MemoryStore, WorkingMemory, build_context |
| Verification | Verdict, ReportCheck, deterministic_check_report, llm_judge_report, verify_report |
| Agent roles | PlannerAgent, WorkerAgent, CriticAgent |
| Budget + recovery | BudgetMulti, ErrorClass, classify_error, retry_with_backoff |
| Tracing | TraceEvent, Tracer |
| Composition root | Orchestrator, RunResult |
Two things worth knowing before you extend it:
- Backends are a one-line swap.
MockProvideris fully offline and rule-based;AnthropicProvideruses Claude. The code path is identical either way, which is what makes the eval suite runnable in CI. - Verification is pluggable.
Orchestrator.run()accepts acheckcallable that replaces the deterministic tier wholesale, so you can verify something other than "does this text mention these terms" without touching the escalation logic above it.
Notebook 5 explains every design decision in prose, but the executable pieces live in mcp_server/ so they can run outside the notebook — from a shell, a Makefile, or an MCP host's launch configuration:
| File | Role |
|---|---|
mcp_openalex_server.py |
The dependency-free MCP stdio server: five tools + one resource over the read-only OpenAlex database |
create_openalex_db.py |
The build pipeline: raw JSONL cache → parse_work → two-pass bulk load → FTS index. Run as python create_openalex_db.py [raw_path] [db_path] |
mcp_openalex_adapter.py |
A dual-era adapter that answers the legacy (pre-2026) initialize handshake itself and bridges older clients to the stateless server without modifying it |
test_mcp_openalex_server.py |
Wire-level tests that spawn the server as a subprocess and exercise it exactly as a host would |
Everything on the server's critical path is standard library only (json, sqlite3, sys), so it runs as a bare subprocess wherever a host launches it. The server is modern-only: it requires a host that speaks MCP 2026-07-28, the latest protocol revision, and rejects the legacy initialize handshake explicitly — hosts still on an earlier version should launch mcp_openalex_adapter.py instead. To register it with a stdio host, point the host at the server script with absolute paths:
{
"mcpServers": {
"openalex": {
"command": "/ABS/PATH/TO/python3",
"args": ["/ABS/PATH/TO/mcp_server/mcp_openalex_server.py"]
}
}
}The server resolves data/openalex.db relative to its own location, so it is independent of whatever working directory the host chooses.
The repository also ships a project-scoped .mcp.json, read by hosts such as Claude Code, that registers the server through mcp_openalex_adapter.py. Its command points at the author's virtual environment, so change it to the Python interpreter in your own .venv before using it.
LLMs/
├── 01 - Basic Agentic Harness.ipynb # Notebook 1: Fundamentals
├── 02 - Small Language Model.ipynb # Notebook 2: n-gram language model
├── 03 - Advanced Agentic Harness.ipynb # Notebook 3: Production-grade patterns
├── 04 - Evaluating Agentic Harnesses.ipynb # Notebook 4: Eval suite, cost, failure modes
├── 05 - MCP Server.ipynb # Notebook 5: MCP server from scratch
├── 06 - Local Inference with vLLM.ipynb # Notebook 6: Self-hosted inference with vLLM
├── d4sci_harness.py # The harness from notebook 3, importable
├── mcp_server/ # Notebook 5's executable pieces
│ ├── mcp_openalex_server.py # The from-scratch MCP stdio server
│ ├── create_openalex_db.py # OpenAlex → SQLite build pipeline
│ ├── mcp_openalex_adapter.py # Dual-era protocol adapter
│ └── test_mcp_openalex_server.py # Wire-level server tests
├── .mcp.json # Project MCP config that launches the adapter
├── data/
│ ├── D4Sci_logo_ball.png
│ ├── D4Sci_logo_full.png # Logos and assets
│ ├── bgoncalves.png # Author photo
│ ├── openalex_raw.jsonl.gz # Raw OpenAlex subset data
│ └── openalex.db # OpenAlex SQLite database
├── outputs/ # Everything the notebooks write, one folder per notebook
│ ├── small_language_model/cache/ # Notebook 2's cached 4-gram model
│ ├── advanced_harness/ # Notebook 3's plan DAG figures
│ ├── evaluating_harnesses/ # Notebook 4's re-plan DAG figure
│ └── vllm/ # Notebook 6's extractions.jsonl
│ └── cache/ # Notebook 6's cached generations and measurements
├── d4sci.mplstyle # Custom matplotlib style
├── pyproject.toml # Dependency manifest (for `uv sync`)
├── uv.lock # Lock file for reproducible builds
└── LICENSE # MIT License
- Install
uv(if needed):
curl -LsSf https://astral.sh/uv/install.sh | sh- Create an environment and install dependencies:
git clone https://github.com/DataForScience/LLMs.git
cd LLMs
uv venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
uv syncThe project requires Python 3.13 or newer; uv venv picks a matching interpreter from pyproject.toml and downloads one if needed.
Notebook 1 is the only notebook that needs an API key: it calls Claude through AnthropicProvider and has no offline fallback. Notebooks 3 and 4 include a rule-based mock provider that is smart enough to drive the demos, notebooks 2 and 5 don't call an LLM at all, and notebook 6 runs its model locally (see Hardware for notebook 6).
Notebooks 3 and 4 ship with BACKEND = "anthropic", so to run them offline set:
BACKEND = "mock"For notebook 1, and for notebooks 3 and 4 on the real backend, export your Anthropic API key before launching Jupyter:
export ANTHROPIC_API_KEY=sk-ant-...Notebook 5 needs an OPENALEX_API_KEY only for a fresh data fetch — the cached raw
data lets the rest of the notebook run without credentials:
export OPENALEX_API_KEY=...Mock mode is deterministic, which makes it the right choice for CI and for following along. Real-backend runs are not — plans vary between runs, so expect the eval suite's pass rate to move around. That variability is the point of measuring it.
Keys aside, the first run of several notebooks downloads data: notebook 2 fetches WikiText-103
from Hugging Face, notebooks 3 and 4 fetch the all-MiniLM-L6-v2 embedding model, and notebook 6
fetches its model weights.
Notebook 6 loads a real model into GPU memory, so it needs a Linux machine with an NVIDIA GPU. uv sync installs vLLM and its GGUF plugin only on Linux
(both dependencies carry a sys_platform == 'linux' marker), so on macOS the rest of the
environment installs normally and notebook 6 is the one to skip. On Linux, the lock file pins
vLLM 0.31.0 on PyTorch 2.13 with CUDA 13 wheels, so your NVIDIA driver must support CUDA 13.
- Target machine. The notebook was written and measured on an NVIDIA DGX Spark (128 GB of unified memory at 273 GB/s). On other hardware, set
MEMORY_BANDWIDTH_GBPSandDEVICE_MEMORY_GBin the configuration cell to your device's figures and adjustGPU_MEMORY_UTILIZATIONto the memory you can spare; the model-choice arithmetic in Section 1 then tells you whether the default model still fits comfortably. - Model download.
Qwen/Qwen3.5-35B-A3B-FP8is a public, Apache-2.0 checkpoint, so no Hugging Face token is needed, but the first run downloads about 37 GB of weights, and the first engine start spends several minutes compiling the model and capturing CUDA graphs. - Data. The extraction runs over
data/openalex.db, the database built in notebook 5. - Result cache. Every generation and measurement is saved to
outputs/vllm/cache/, keyed on the vLLM version, the engine arguments, and the inputs. With the same vLLM version and settings, a re-run reloads those results instead of regenerating them (the engine itself still starts). The repository ships the cache from the original DGX Spark run, made with the vLLM version inuv.lock, so a fresh clone reloads those results too. SetREFRESH_CACHE = Trueto measure everything again, for instance on different hardware.
jupyter notebookReach out at info@data4sci.com or open an issue if something isn't working.
|
Web: www.data4sci.com |
