Make memcached opt-in, default to a capped in-process cache - #4814
Merged
Merged
Conversation
The free Memcachier dev plan is timing out on every call in production (58 IO::TimeoutError in a 14 minute log window). Each request pays the 0.5s socket timeout and then renders uncached anyway, so the cache is currently pure overhead on a dyno that is already being OOM-killed. Memcached is now used only when USE_MEMCACHED=true and MEMCACHIER_SERVERS is present. Otherwise production uses a memory store capped at 32MB via RAILS_CACHE_SIZE_MB, so deploying restores real caching without the timeouts and without an addon upgrade. The memory store is per process, so homepage_stats gets an expires_in to bound staleness. It is the only fragment that relies on expire_fragment rather than a versioned cache key. Sessions are cookie-based and every other cached value is either keyed by cache_key_with_version or carries a TTL, so nothing depends on the cache being shared between processes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CloCkWeRX
reviewed
Sep 21, 2026
CloCkWeRX
approved these changes
Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The free Memcachier
devplan is failing in production. In a 14 minute window ofgrowstuff-prodlogs there were 58IO::TimeoutError: Blocking operation timed out!againstmc2.dev.ec2.memcachier.com:11211.With
socket_timeout: 0.5, every one of those is a half-second stall followed by a full uncached render. So right now the cache costs us latency and buys us nothing — on a dyno that is simultaneously being OOM-killed every ~9 minutes (R14→R15→ SIGKILL 137).What this does
Memcached becomes opt-in. Production uses it only when
USE_MEMCACHED=trueandMEMCACHIER_SERVERSis set. Otherwise it falls back to:memory_store, capped at 32 MB and tunable withRAILS_CACHE_SIZE_MB.Deploying this restores working caching immediately, with no addon upgrade and no spend. If we later move to a paid Memcachier plan, setting
USE_MEMCACHED=trueswitches back with a config var and a restart — no deploy.USE_MEMCACHEDMEMCACHIER_SERVERSmemory_store(32 MB)memory_store(32 MB)truemem_cache_storetruememory_store(32 MB)That last row matters: the old code did
(ENV["MEMCACHIER_SERVERS"] || "").split(",")and handed Dalli an empty server list if the var ever went missing. Now an empty list can't reach Dalli.Is a per-process cache safe here?
Yes, as far as I can tell:
config/initializers/session_store.rbuses:cookie_store, not the cache.Rails.cache.fetchis regenerable and either keyed bycache_key_with_version(so writes invalidate naturally, across processes) or carries anexpires_in.MemoryStore, deliberately, per the comment added in Stop Rack::Attack depending on memcached, and cut the memcached timeout #4800.The one exception is
homepage_stats, the only fragment cached under a static key with no TTL, relying solely onexpire_fragment. With a per-process store that only clears the worker that handled the write, so this PR gives itexpires_in: 1.hourto bound how long another worker can serve a stale copy. The counts in it are member/crop/planting/garden totals, so an hour is plenty fresh.Memory impact
Worth being explicit, since the dyno is memory-constrained: this store is per Puma worker, so the cost is
RAILS_CACHE_SIZE_MB × WEB_CONCURRENCY. At the currentWEB_CONCURRENCY=2that's up to 64 MB; atWEB_CONCURRENCY=1it's up to 32 MB. That's a real cost, but it replaces re-rendering every page on every request, which is the more expensive of the two.Testing
rubocop config/environments/production.rb— clean.haml-lint app/views/home/_stats.html.haml— 8 lints, byte-identical to the 8 already ondev(same offences, shifted 2 lines by the added comment). No new lints.rspec spec/views/home— 12 examples, 0 failures.rspec spec/requests/rack_attack_spec.rb— 2 failures, but these are pre-existing: I confirmed they fail identically on a clean tree with this branch stashed. They're the honeypot and excessive-crawling ban tests. Unrelated to this change, but see below.Not in this PR
This is one of several things making production unstable and it is not the root cause. The dyno is being OOM-killed because
config/puma.rbwas switched to clustered mode with 2 workers +preload_app!in dc8503c (Sep 17), tripling the process count on a 512 MB Basic dyno.WEB_CONCURRENCY=1is the immediate mitigation there.Also flagging that the two failing Rack::Attack ban specs are worth a look on their own — crawler bans are held in per-process memory and the dyno currently restarts every ~9 minutes, so long-window bans never accumulate in production either.
BUT ALSO:
i did this tonight, which sseems to have help:
🤖 Generated with Claude Code