Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
name: CI

on:
push:
branches: [main]
pull_request:

jobs:
check:
# windows-latest, not ubuntu: the app is Windows-first and several suites
# exercise Windows path semantics (drive letters, backslashes, reserved
# names) that a Linux runner would never catch.
runs-on: windows-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- run: npm ci
- run: npm run typecheck
- run: npm run lint
- run: npx prettier --check .
- run: npm test
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,3 +27,7 @@ build/logo-mockups/
# Channel signing key — never commit. See scripts/registry-channel.mjs
channel-private.pem
channel.json

# Playwright artifacts
test-results
playwright-report
4 changes: 3 additions & 1 deletion .prettierignore
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
node_modules
out
dist
resources/bin
resources
package-lock.json
# Design mockups are point-in-time artifacts, not living code.
docs/mockups
3 changes: 3 additions & 0 deletions .prettierrc.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,3 +2,6 @@ semi: false
singleQuote: true
printWidth: 100
trailingComma: none
# Match what git checks out on Windows (core.autocrlf) so `format` never
# rewrites the whole tree over line endings.
endOfLine: auto
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,8 @@ resources/bin/ bundled CLI binaries (gitignored; fetched by scripts, packed by e
- `npm test` — Vitest unit tests (arg-builders, format catalogs, output collision-safety).
- `npm run lint` / `npm run format` — eslint (flat config) / prettier.
- `npm run package` — electron-vite build + electron-builder NSIS installer to `dist/`.
- UI end-to-end: Playwright (scaffolded once the real UI exists — per the global testing rule).
- `npm run test:e2e` — Playwright end-to-end (launches the built app via `_electron`; run
`npm run build` first). Covers the preload/IPC/engine chain unit tests can't reach.

## Conventions

Expand Down
24 changes: 12 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,15 @@
<div align="center">
<img src="docs/logo.png" alt="Filesmith" width="116">

# Filesmith
# Filesmith

A local file toolkit for Windows. Drop files, pick a tool, get results.
A local file toolkit for Windows. Drop files, pick a tool, get results.

[![Latest release](https://img.shields.io/github/v/release/Maxaubert/Filesmith?style=flat-square&color=5b5bd6&cacheSeconds=1800)](https://github.com/Maxaubert/Filesmith/releases/latest)
[![Downloads](https://img.shields.io/github/downloads/Maxaubert/Filesmith/total?style=flat-square&color=5b5bd6&cacheSeconds=1800)](https://github.com/Maxaubert/Filesmith/releases)
[![Windows](https://img.shields.io/badge/Windows-10%20%7C%2011-0078D4?style=flat-square)](https://github.com/Maxaubert/Filesmith/releases/latest)
[![Built with](https://img.shields.io/badge/Electron%20·%20React%20·%20TypeScript-2b2e3a?style=flat-square)](#build-from-source)
[![License: MIT](https://img.shields.io/badge/License-MIT-22b364?style=flat-square)](LICENSE)
[![Latest release](https://img.shields.io/github/v/release/Maxaubert/Filesmith?style=flat-square&color=5b5bd6&cacheSeconds=1800)](https://github.com/Maxaubert/Filesmith/releases/latest)
[![Downloads](https://img.shields.io/github/downloads/Maxaubert/Filesmith/total?style=flat-square&color=5b5bd6&cacheSeconds=1800)](https://github.com/Maxaubert/Filesmith/releases)
[![Windows](https://img.shields.io/badge/Windows-10%20%7C%2011-0078D4?style=flat-square)](https://github.com/Maxaubert/Filesmith/releases/latest)
[![Built with](https://img.shields.io/badge/Electron%20·%20React%20·%20TypeScript-2b2e3a?style=flat-square)](#build-from-source)
[![License: MIT](https://img.shields.io/badge/License-MIT-22b364?style=flat-square)](LICENSE)
</div>

---
Expand Down Expand Up @@ -48,11 +48,11 @@ Plus batch queues, thumbnails for every kind (images, video frames, audio cover

The everyday tools (convert, compress, resize, and all PDF / video / audio / document operations) run **fully offline** from bundled binaries, with no AI and no downloads. The AI features are **opt-in**, and nothing AI runs unless you choose it:

| Feature | What it needs |
|---|---|
| **Remove Background** | The free [`uv`](https://docs.astral.sh/uv/) tool; a small AI model is downloaded on first use (once, then offline). The panel tells you before you commit files. |
| **Upscale** | An upscaling model, downloaded on first use. NVIDIA (PiD) mode needs an NVIDIA GPU. |
| **Generate** | An existing [ComfyUI](https://github.com/comfyanonymous/ComfyUI) install, which Filesmith drives headlessly. Supports SDXL, Flux 1, Flux 2 (klein), Z-Image, and Krea 2 models; missing text-encoders / VAEs can be downloaded from the panel. If something is missing, the panel says exactly what to do. |
| Feature | What it needs |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Remove Background** | The free [`uv`](https://docs.astral.sh/uv/) tool; a small AI model is downloaded on first use (once, then offline). The panel tells you before you commit files. |
| **Upscale** | An upscaling model, downloaded on first use. NVIDIA (PiD) mode needs an NVIDIA GPU. |
| **Generate** | An existing [ComfyUI](https://github.com/comfyanonymous/ComfyUI) install, which Filesmith drives headlessly. Supports SDXL, Flux 1, Flux 2 (klein), Z-Image, and Krea 2 models; missing text-encoders / VAEs can be downloaded from the panel. If something is missing, the panel says exactly what to do. |

Don't want AI? Just use the core tools, and the AI features stay out of your way.

Expand Down
18 changes: 9 additions & 9 deletions docs/adding-models.md
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
# Adding a model to Filesmith

Filesmith does not bake models in. What a model *is* — how to recognize it, what files it needs,
Filesmith does not bake models in. What a model _is_ — how to recognize it, what files it needs,
which ComfyUI graph runs it — is **data on disk**, so you can add a model family that did not exist
when your copy of Filesmith was built, without waiting for a release.

## Where the registry lives

Three layers, merged by `id`, later layers winning **field by field**:

| Layer | Path | Who writes it |
|---|---|---|
| 1. Built-in | `<install>/resources/registry/*.json` | ships in the installer, read-only |
| 2. Channel | `%APPDATA%/Filesmith/registry/channel/` | refreshed over the network (signed) |
| 3. **Yours** | `%APPDATA%/Filesmith/registry/user/` | you |
| Layer | Path | Who writes it |
| ------------ | --------------------------------------- | ----------------------------------- |
| 1. Built-in | `<install>/resources/registry/*.json` | ships in the installer, read-only |
| 2. Channel | `%APPDATA%/Filesmith/registry/channel/` | refreshed over the network (signed) |
| 3. **Yours** | `%APPDATA%/Filesmith/registry/user/` | you |

Two rules that will not change:

Expand Down Expand Up @@ -129,7 +129,7 @@ Available placeholders:
`${unet}` `${clip}` `${clip2}` `${vae}` `${model}` · `${prompt}` `${negative}` `${seed}` `${steps}`
`${cfg}` `${guidance}` `${sampler}` `${scheduler}` `${width}` `${height}` `${batch}` `${prefix}`

A value that is *exactly* one placeholder is replaced with the raw value, so `"seed": "${seed}"`
A value that is _exactly_ one placeholder is replaced with the raw value, so `"seed": "${seed}"`
yields a number, not a string. Use `workflow` for a bare diffusion model (UNETLoader + separate
encoders) and `checkpointWorkflow` for an all-in-one single-file checkpoint. An entry can have both.

Expand Down Expand Up @@ -165,7 +165,7 @@ always tell whether Filesmith saw the file at all.
Every download in `resources/registry/gen-archs.json` carries a real `sha256` and a URL pinned to
an immutable commit revision, with the `resolve/main` branch URL kept after it as a fallback mirror.

The hashes are not invented: Hugging Face stores large files in git-LFS, and an LFS object id *is*
The hashes are not invented: Hugging Face stores large files in git-LFS, and an LFS object id _is_
the sha256 of the content, exposed per file by the repo tree API.

```
Expand All @@ -183,7 +183,7 @@ trust-on-first-use, which is where everything was before.

## Publishing a channel update

Every companion URL points into someone else's repo. When one is reorganized, every *installed*
Every companion URL points into someone else's repo. When one is reorganized, every _installed_
copy of Filesmith gets a 404 and stays broken until a new release ships. The channel fixes that
without a release: publish a signed pack, and every install picks it up within a day.

Expand Down
4 changes: 2 additions & 2 deletions docs/comfyui-upscaler-import.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ uses) and run through our own tiled upscaler.
## Why this over bundling

- **Licensing**: most good community upscalers are non-commercial (CC-BY-NC-SA),
so we can't ship them. Running the user's *own* files is distribution-free.
so we can't ship them. Running the user's _own_ files is distribution-free.
- **VRAM**: tiled ESRGAN inference is bounded, so it avoids the large-output
blow-ups the diffusion path (PiD) hits.
- **Coverage**: spandrel auto-detects architecture, so one code path handles the
Expand All @@ -22,7 +22,7 @@ uses) and run through our own tiled upscaler.
1. **Import policy** — load anything spandrel accepts; badge known-good models
**Verified**, the rest **Experimental**; diffusion/unloadable files are shown
greyed as **Unsupported** with the reason.
2. **Placement** — imported models sit *alongside* the existing options
2. **Placement** — imported models sit _alongside_ the existing options
(Photo / Anime bundled Real-ESRGAN, PiD). Nothing is removed. PiD keeps its
own diffusion path (and its pending input-resize fix).
3. **Storage** — **reference in place**: read models straight from the ComfyUI
Expand Down
28 changes: 22 additions & 6 deletions docs/compression-upgrade-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,13 @@ A running design doc for expanding the **Compress** tab, decided category by
category with the user. Status per section: DECIDED / discussing / TODO.

**STATUS: IMPLEMENTED** (images, video, audio, PDF). Engine + per-kind options UI
+ live video-resolution preview + Ghostscript/ffprobe bundling all landed and
verified end-to-end against the real binaries. Possible follow-up: allow WAV/FLAC
into audio Compress (wav→opus is a big win), currently excluded by canCompress.

- live video-resolution preview + Ghostscript/ffprobe bundling all landed and
verified end-to-end against the real binaries. Possible follow-up: allow WAV/FLAC
into audio Compress (wav→opus is a big win), currently excluded by canCompress.

Current baseline (what the Compress tab does today):

- Images: CaesiumCLT (jpg/png/webp/gif/tiff) + ImageMagick fallback, re-encode at
a quality value, **same format in/out**.
- Video: ffmpeg CRF, x264 (VP9 for webm), fixed 128k AAC audio, one quality slider.
Expand All @@ -31,16 +33,18 @@ Add a **Format** choice to the image Compress options. Three options:
narrower compatibility. The "squeeze it as hard as possible" option.

Key decisions:

- **Both new formats run in LOSSY mode**, driven by the existing quality slider.
Rationale: a Compress button should reliably make files smaller; lossy is
where the big wins are for photos (the common case). Lossless WebP/AVIF only
helps graphics and can make photos *larger*, so it's not offered here.
helps graphics and can make photos _larger_, so it's not offered here.
- **No "→ JPEG" option in Compress.** Converting to JPEG is a footgun on
transparency / sharp edges; that belongs in the Convert tool.
- Perfect-quality (lossless) format conversion, if ever wanted, also belongs in
Convert, not Compress.

Implementation notes:

- ImageMagick (already bundled) can already encode WebP and AVIF at `-quality`
(verified in the conversion matrix), so an **MVP needs no new binary** — route
the WebP/AVIF options through magick with the mapped quality.
Expand All @@ -51,6 +55,7 @@ Implementation notes:
- Transparency is preserved in lossy WebP/AVIF (unlike JPEG).

Second-wave / not now (nice-to-have specialists):

- mozjpeg (better JPEG encoder, 10-30% smaller), oxipng / pngquant (PNG lossless
/ palette), gifsicle (GIF). Bigger, format-specific wins; revisit later.
- "Strip metadata" toggle (EXIF/ICC/XMP) — cheap, always helps a little.
Expand All @@ -62,6 +67,7 @@ Second-wave / not now (nice-to-have specialists):
Three independent controls in the video Compress options:

### a. Codec (the biggest lever) — same "compatible / smaller / smallest" framing as images

1. **H.264 (compatible)** — universal, plays everywhere, least efficient. Today's
default. ffmpeg `libx264`.
2. **H.265 / HEVC (smaller)** — ~30-50% smaller than H.264 at equal quality,
Expand All @@ -70,16 +76,19 @@ Three independent controls in the video Compress options:
3. **AV1 (smallest)** — another ~20-35% smaller than H.265, slowest encode,
narrower support. ffmpeg `libsvtav1` (fast AV1 encoder). The bundled gyan.dev
essentials ffmpeg includes SVT-AV1 (verified).

- **Output is always .mp4**, driven by the chosen codec (H.264/H.265/AV1 all mux
there). The codec picker replaces the old auto-VP9-for-WebM behavior: a `.webm`
source is re-encoded to `.mp4` with the selected codec. (No VP9 option is
offered; VP9/WebM output would be a Convert-tool concern if ever wanted.)

### b. Quality — the existing slider

- Maps to CRF per codec (the CRF scale differs by codec: ~18-32 for x264/x265,
higher numbers for AV1, so use a per-codec mapping, not one shared range).

### c. Resolution — presets + live output list

- Presets: **Original / 1440p / 1080p / 720p / 480p / 360p / 240p**.
- Each preset means "fit within that box, **preserve aspect ratio, never
upscale**" — so it works for any shape (landscape, portrait, ultrawide) and any
Expand All @@ -99,6 +108,7 @@ Three independent controls in the video Compress options:
- Optional later: a percentage / custom-dimensions mode for power users.

### Audio track (inside video)

- MVP: keep the current fixed AAC 128k. Optionally expose audio bitrate / "remove
audio" later, reusing the audio-category decisions below.

Expand All @@ -109,33 +119,38 @@ lossy formats — mp3/m4a/aac/ogg/opus/wma; lossless flac/wav are excluded by
`canCompress`, as today.) All via bundled ffmpeg, no new tools.

### a. Codec / format picker — parenthetical labels for the tradeoff

1. **MP3 (compatible)** — universal, least efficient. `libmp3lame`.
2. **AAC (balanced)** — better quality per byte, widely supported (`.m4a`). ffmpeg `aac`.
3. **Opus (smallest)** — most efficient (~40-60% smaller at equal quality),
especially at low bitrates; fine on modern devices. `libopus`.

- "Keep format" behavior can remain the default (re-encode in the source codec),
with MP3/AAC/Opus as explicit targets.

### b. Bitrate picker (replaces the quality slider for audio)

- User picks the **bitrate directly**, shown in kbps (not an abstract quality
slider). Preset values: **320 / 256 / 192 / 128 / 96 / 64 kbps** (default ~192).
- ffmpeg `-b:a <N>k`. Lower bitrate = smaller file; Opus stays clean lower than
MP3/AAC.

### Explicitly NOT doing (this round)

- **No "already-compressed" warning.** (User decision.)
- Sample-rate reduction / stereo→mono downmix / Voice-Podcast-Music presets:
deferred, revisit later if wanted.

## 4. PDF — DECIDED

Bundle **Ghostscript** (`gs`) — the only mature offline single-binary tool that
downsamples the *images inside* a PDF (mutool/qpdf only touch structure, which is
downsamples the _images inside_ a PDF (mutool/qpdf only touch structure, which is
why today's compress does almost nothing on scanned / image-heavy PDFs). Worth
the ~60 MB. License: AGPL, compatible since Filesmith is open-source and already
ships GPL ffmpeg (no network service, so AGPL's extra clause doesn't bite).

### a. Compression level picker — default Balanced

- **Lossless** — today's `mutool clean -gggg -z`. No image changes, safe, small
savings. Best for text-only PDFs / when you don't want to risk altering render.
- **High quality** (~300 dpi) — Ghostscript `-dPDFSETTINGS=/printer`. Downsamples
Expand All @@ -149,9 +164,10 @@ ships GPL ffmpeg (no network service, so AGPL's extra clause doesn't bite).
mutool is the conservative "shrink without altering" pass.

### b. Grayscale toggle — optional, OFF by default

- Converts color → grayscale for docs that don't need color (scanned documents to
email). ~30-50% extra. Ghostscript `-sColorConversionStrategy=Gray
-sProcessColorModel=DeviceGray`. Off by default so it never surprises anyone.
-sProcessColorModel=DeviceGray`. Off by default so it never surprises anyone.

Base gs invocation:
`gs -sDEVICE=pdfwrite -dPDFSETTINGS=/<tier> -dNOPAUSE -dBATCH -dQUIET -o out.pdf in.pdf`
Expand Down
Loading
Loading