diff --git a/.github/workflows/check-docs.yml b/.github/workflows/check-docs.yml index 24e1f841..70c0eaf3 100644 --- a/.github/workflows/check-docs.yml +++ b/.github/workflows/check-docs.yml @@ -14,8 +14,25 @@ jobs: - name: Checkout uses: actions/checkout@v4 + # Jena 6 requires Java 21; the runner defaults to 17 + - name: Set up Java 21 + uses: actions/setup-java@v4 + with: + distribution: temurin + java-version: '21' + - name: RDF syntax check uses: AtomGraph/RDF-syntax-check@v1.0.5 + with: + jena-version: 6.1.0 + + - name: Install Linux packages + run: sudo apt-get update && sudo apt-get install -y libxml2-utils + + # riot checks rdf:XMLLiteral well-formedness; canonical form is only + # enforced up to Jena 4.7.0, so it is checked separately + - name: XMLLiteral canonical form check + run: ./check-xmlliterals.sh - name: Check links run: ./check-links.sh diff --git a/.gitignore b/.gitignore index f998a83d..51328d6c 100644 --- a/.gitignore +++ b/.gitignore @@ -4,3 +4,4 @@ linkeddatahub/.DS_Store docs/docs.trix docs/files.xml docs/html +docs/timestamps.xml diff --git a/CLAUDE.md b/CLAUDE.md index f9793eff..ec801e65 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2,80 +2,111 @@ This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. -## Repository Overview +## Repository overview -LinkedDataHub-Apps is a collection of data-driven applications built on LinkedDataHub. The repository contains system, demo, and user-submitted applications that are completely data-driven with no imperative code - only shell scripts for installation and SPARQL queries for data processing. +LinkedDataHub-Apps holds dataspaces built on LinkedDataHub. Every dataspace is a tree of RDF documents plus the files they carry, installed with the `ldh` CLI. There is no imperative code: behaviour lives in RDF (Turtle), SPARQL and XSLT, and shell only orchestrates CLI calls. -## Repository Structure +- `docs/` — the LinkedDataHub documentation, itself a LinkedDataHub app, published live at https://docs.linkeddatahub.com/ and as a static site under https://atomgraph.github.io/LinkedDataHub/linkeddatahub/docs/ +- `demo/northwind-traders/` — Northwind Traders as a knowledge graph: faceted search, parallax navigation, CSV import, a 969-line namespace ontology with constructors, constraints and views +- `demo/copenhagen/` — Copenhagen open data on a map, imported from CSV +- `demo/unesco-thesaurus/` — the UNESCO Thesaurus as a SKOS editor; its model and stylesheet come from the taxonomy editor package +- `packages/` — the package registry published at https://packages.linkeddatahub.com/; one package so far, `editor/taxonomy` +- `lib/` — the shared installer: `install.mk` (the interactive `make install`) and `ldh-app.sh` (functions every `install.sh` sources) +- `screencast/` — the Playwright rig that produces the docs' screenshots and clips, and the Northwind screencast tracks +- `check-links.sh`, `check-xmlliterals.sh` — repository-wide validators, run by CI -- `/linkeddatahub/docs/` - Documentation for LinkedDataHub open-source and Cloud versions -- `/demo/northwind-traders/` - Knowledge Graph representation of Northwind Traders sample database with faceted search -- `/demo/city-graph/` - Geospatial browser for Copenhagen open data with type-colored geospatial overview -- `/demo/skos/` - Basic SKOS editor with custom UI theme for concept and schema management +## Prerequisites -Each demo application includes: -- Installation shell scripts -- SPARQL queries for data import and processing -- CSV data files (for data import) -- Custom ontologies and constraints (stored as .ttl files) -- Optional custom XSLT stylesheets and CSS +The `ldh` CLI, built from a LinkedDataHub checkout and put on `PATH`: -## Common Commands +```bash +cd ../LinkedDataHub/cli && mvn package && export PATH="$PWD/bin:$PATH" +``` + +The docs build additionally needs `riot` (Apache Jena), Docker (for `atomgraph/saxon`), `xmlstarlet`, `xmllint` and `python3`. The screencast rig needs Node, `ffmpeg`, `cwebp`, `vhs` and GNU `sha1sum`. + +## Installing an app -### Installation -Applications use interactive Makefiles for installation: ```bash -make install # Interactive installation with prompts for BASE_URL, CERT_PATH, PASSWORD, PROXY_URL +make install # prompts for base URL, keystore path, keystore password and optional proxy URL ``` -### Documentation (linkeddatahub/docs/) +`make install` is `lib/install.mk`, included by every app Makefile. It exports the answers as `LDH_BASE`, `LDH_CERT_FILE`, `LDH_CERT_PASSWORD` and `LDH_PROXY` — the variables the CLI reads — and runs the app's `install.sh` in that environment. No credential is put on a command line, so the password needs no escaping. To install unattended, export the four variables and run `./install.sh`. + +Each `install.sh` sources `lib/ldh-app.sh` and is the app's own sequence of steps: + +- `ldh admin make-public` — an authorization that makes the app readable by anyone +- `ldh_app_import_ns admin/model` — resets the namespace ontology document with `patch-ontology.ru`, appends `ns.ttl` under `@base <{base}ns>`, clears the ontology cache +- `ldh push --dir . "$LDH_BASE"` — PUTs every RDF file to the document its path spells and uploads every other file into its folder's document +- `ldh_app_import_csv . imports.csv` — one `ldh import csv` per manifest row (`query_filename,csv_filename,target,title`) +- app-specific steps: `ldh admin create authorization`, `ldh packages add`, `ldh import rdf` + +**Re-runs converge but are not clean.** `ldh push` PUTs and the ontology is reset before re-import, so documents and the model end up the same. `make-public`, `create authorization` and every import are POSTs, so each run adds another authorization and another import record. Nothing probes for an earlier install. + +### App layout + +The file tree is the URL tree. `categories.ttl` installs to `{base}categories/`, the sibling folder `categories/` holds that document's uploads (and, by convention, its `.csv` and mapping `.rq`), and `root.ttl` is the base document itself. `.ldhignore` lists what is neither a document nor an upload — `admin/`, `Makefile`, `*.sh`, `*.csv`, repository screenshots — because the CLI skips nothing but hidden entries. + +Mapping queries are `CONSTRUCT { GRAPH ?g { … } }` over CSV rows addressed as `<#column>`, and build every URI from the `$base` constant the platform binds. **Do not pin a hostname in a query or document**: a class is `BIND(uri(concat(str($base), "ns#School")) AS ?type)`, a query on the `` endpoint may use relative IRIs (``, ``), and a query on the data endpoint cannot, since that endpoint parses with no base. + +Media is content-addressed. `ldh push` uploads a file to `{base}uploads/{sha1}`, and a CSV column or a document references it by that hash, so the bytes must be final before anything references them. + +## Documentation (`docs/`) + ```bash -make validate # Validate documents -make ttl-to-html # Convert Turtle files to HTML +make install # push the docs to a LinkedDataHub instance +make validate # Turtle syntax and canonical XMLLiteral check +make check-links # resolve every relative link in the whole repository +make ttl-to-html # static site into docs/html/ (gitignored) ``` -### Prerequisites -All installation scripts require LinkedDataHub CLI scripts in PATH: +Each `.ttl` is one page: a `dh:Container` or `dh:Item` with `rdf:_N`-ordered blocks, whose content is an XHTML `rdf:XMLLiteral`. The static build (`ttl-to-html.sh` + `ttl-to-html.xsl`) is a second renderer of the same sources and rewrites `uploads/{sha1}` to `files/{name}`, terminating on a hash it cannot find — a clean `make ttl-to-html` proves every baked hash resolves. + +CI (`.github/workflows/check-docs.yml`) runs RDF syntax, canonical-XML and link checks on every push. The static build and the GitHub Pages deploy (`publish-docs.yml`) run only on pushes to `master` that touch `docs/`. + +XHTML inside `rdf:XMLLiteral` must be exclusive canonical XML: attributes alphabetical, explicit end tags (`

`, ``), each start tag on one line. `check-xmlliterals.sh` diffs each literal against `xmllint --exc-c14n`. + +### Documentation media (`screencast/`) + +Screenshots and clips are produced by a scripted Playwright rig, never taken by hand. + ```bash -export PATH="$(find bin -type d -exec realpath {} \; | tr '\n' ':')$PATH" +cd screencast +node docs/shoot.mjs --base … --cert-file … --cert-password-file … [--only reference] +make docs-publish # optimise docs/out/ into ../docs/, printing sha1 per published file +make docs-fill # point the .ttl sources at the published files (idempotent) +``` + +- `docs/manifest.mjs` is the shot list — one entry per `div.screenshot-placeholder` in the `.ttl` sources, keyed by its caption. New slots go there, not in the runner. +- Each entry's `want` is asserted after `act` and before the capture, so a shot that never reached its state is reported missed rather than written. +- `blocked` slots stay as placeholders with their reason recorded. Do not fill one by hand. +- The shoot writes **masters** — 2880px lossless PNG, a `.webm` and an `.mp4` per clip — into `docs/out/`, plus an `index.json`. It never edits a `.ttl`; `docs/fill.mjs` does. +- `make docs-publish` derives the web assets into `docs/`: stills to WebP at 2240px, `.mp4`s copied (already CRF 20 / `yuv420p` / faststart), `.webm`s dropped. + +**Optimise before hashing.** Uploads are content-addressed at `{base}uploads/{sha1}`, so a reference is only valid for the bytes that ship — never hash a master. + +References are absolute, `/uploads/{sha1}`: the `uploads/` namespace hangs off the base URI, outside the document hierarchy, so a relative path would encode the referring document's depth and break when it moves. They live inside canonical `rdf:XMLLiteral` bodies: + +```xml +The document tree with a container expanded + ``` -## Application Architecture - -### Data-Driven Approach -- **Zero imperative code** - applications are entirely configuration-driven -- **SPARQL-based processing** - all data transformations use SPARQL queries -- **RDF/Turtle ontologies** - define application structure and constraints -- **Shell script orchestration** - only for installation and deployment tasks - -### Key Components per Application -1. **Installation scripts** (`install.sh`) - Deploy application to LinkedDataHub instance -2. **Data import scripts** (`import-csv.sh`, `import-rdf.sh`) - Load data from CSV/RDF sources -3. **SPARQL queries** (`queries/` directory) - Define data processing and views -4. **Ontology files** (`admin/model/*.ttl`) - Define classes, properties, and constraints -5. **Configuration scripts** - Create containers, charts, and authorizations - -### Application Flow -1. Install ontologies and model definitions -2. Create containers and authorization structures -3. Import data from CSV/RDF sources using SPARQL transformation queries -4. Create charts and views for data visualization -5. Configure access control and permissions - -### Custom Styling (SKOS demo) -Some applications like SKOS use custom XSLT stylesheets and CSS that need to be mounted in LinkedDataHub configuration and require increased payload size limits. - -## Important Notes - -- **Installation scripts are NOT idempotent** - subsequent runs may add data but aren't guaranteed to succeed -- **Special characters in passwords** need to be escaped in shell scripts -- **Custom stylesheets** require LinkedDataHub configuration changes and Docker volume mounts -- **Access control** - some applications require requesting append/write access to create/edit data - -## File Types and Conventions - -- `.ttl` - Turtle RDF files for ontologies and data -- `.rq` - SPARQL query files -- `.csv` - Data import files -- `.sh` - Shell scripts for installation and deployment -- `Makefile` - Interactive installation and build targets \ No newline at end of file +A clip is a `