Skip to content

Add opt-in glyph caching to vello_cpu and vello_hybrid backends - #116

Merged
nicoburns merged 2 commits into
mainfrom
devin/1791061248-glyph-caching
Oct 3, 2026
Merged

nicoburns merged 2 commits into
mainfrom
devin/1791061248-glyph-caching

Conversation

@nicoburns

@nicoburns nicoburns commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Summary

Exposes Vello's (experimental) glyph atlas cache in anyrender_vello_hybrid and anyrender_vello_cpu. Off by default.

  • Runtime flags only (no cargo feature):
    • hybrid: VelloHybridRendererOptions::glyph_caching(bool) (+ public field), VelloHybridScenePainter::with_glyph_caching(bool), WebGlScenePainter::with_glyph_caching(bool)
    • cpu: VelloCpuScenePainter::set_glyph_caching(bool), VelloCpuImageRenderer::set_glyph_caching(bool). VelloCpuWindowRenderer is a generic softbuffer/pixels wrapper with no options hook, so caching can't be enabled there.
  • Implementation is just .atlas_cache(enabled) on the glyph_run builder, plus a guard (below).

Note: VelloHybridRendererOptions gains a public field, which is breaking for anyone building it with a struct literal.

Skewed/rotated runs bypass the cache

glifo 0.4 only validates the transform when inserting a glyph into the atlas (supports_atlas_caching), not on lookup, and the cache key doesn't include skew/rotation. So once an upright glyph is cached, a skewed run (e.g. synthetic italics) with the same font/size is drawn from the upright entry. supports_glyph_caching therefore disables the atlas for any run whose transform * glyph_transform isn't a positive uniform scale + translation.

tests/glyph_caching.rs (vello_cpu) covers this: it fails (2525 px differing by >64) with the guard removed.

Performance

Blitz paint_bench on https://en.wikipedia.org/wiki/Barack_Obama, 1366x768 @2x, 100 frames per run, 3 interleaved runs per cell. Times are the median µs/frame. "scroll 40" moves the viewport 40 CSS px every frame.

Correction: an earlier version of this description had "scrolling" rows that were actually static frames (the benchmark's scroll range was computed as 0). The tables below are from reruns with that fixed.

macOS, Metal (macOS 26.5.2 VM, "Apple M4 Pro (Virtual)", 12 cores, wgpu adapter Apple Paravirtual device). This is a GPU-backed Metal device but paravirtualised, not bare metal. Values are run1 / run2 / run3.

backend scroll cache encode (paint_scene) rasterize total first frame
hybrid 0 off 3201 / 3197 / 3234 3373 / 3654 / 3496 6593 / 6918 / 6756 219–239ms
hybrid 0 on 1143 / 1149 / 1123 3421 / 3447 / 3000 4546 / 4629 / 4121 223–272ms
hybrid 40 off 3671 / 3755 / 3621 2391 / 2514 / 2169 6091 / 6309 / 5808 214–241ms
hybrid 40 on 1106 / 1144 / 1117 2282 / 2283 / 2155 3527 / 3545 / 3367 218–229ms
cpu 0 off 3503 / 3383 / 3140 1124 / 1188 / 1052 4865 / 4902 / 4463 6–11ms
cpu 0 on 1179 / 1227 / 1201 3728 / 3801 / 3630 4942 / 5053 / 4855 31–42ms
cpu 40 off 5867 / 5490 / 4977 907 / 808 / 915 7345 / 7310 / 6417 6–8ms
cpu 40 on 1540 / 1581 / 1458 7201 / 7716 / 7619 8731 / 8899 / 9084 18–19ms

Linux, no GPU (8 cores; "hybrid" runs on Mesa software Vulkan, so its rasterize column is CPU-emulated and not representative). Values are min–max over the 3 runs.

backend scroll cache encode (paint_scene) rasterize total first frame
cpu 0 off 2149–2218 1370–1389 3504–3592 6–7ms
cpu 0 on 1474–1503 2542–3759 4053–5250 42–43ms
cpu 40 off 3427–3453 801–897 4239–4544 7–8ms
cpu 40 on 1819–1865 6060–7611 8145–9821 43–60ms
hybrid 0 off 4678–4729 38481–38883 43206–43681 111–115ms
hybrid 0 on 1936–1957 39149–40626 41207–42564 114–124ms
hybrid 40 off 5460–5490 32205–32620 37936–38325 112–113ms
hybrid 40 on 1692–1727 30731–31282 32475–32830 111–116ms
  • vello_gpu (hybrid): encode time drops ~60–70% on both machines, rasterize is unchanged within noise. On Metal that is a net ~-33% per frame static and ~-42% scrolling. No consistent change to the first frame.
  • vello_cpu: not a win. The work moves from encode to rasterize and grows: static is a wash on macOS and slower on Linux; scrolling is ~20% slower on macOS and ~2x slower on Linux. The first frame is several times slower (atlas population).
  • Output with cache on vs off: ~3–4% of pixels differ, all antialiasing-level (max channel delta 59–60).

Link to Devin session: https://dioxus.staging.devinenterprise.com/sessions/9fbb960ba27249bca490d138236e0bd1
Open in Devin Desktop: https://dioxus.staging.devinenterprise.com/desktop/session/9fbb960ba27249bca490d138236e0bd1?variant=devin-insiders
Requested by: @nicoburns

Adds a `glyph_caching` cargo feature (sets the default) and runtime
setters to enable Vello's experimental glyph atlas cache.

Runs with a skewed, rotated, mirrored or non-uniformly scaled transform
always bypass the cache: glifo 0.4 only validates the transform when
inserting into the atlas, not when looking up, so such runs would
otherwise be drawn from upright cache entries.
@staging-devin-ai-integration

Copy link
Copy Markdown

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

@nicoburns
nicoburns merged commit 870407d into main Oct 3, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant