Skip to content

Make memcached opt-in, default to a capped in-process cache - #4814

Merged
CloCkWeRX merged 2 commits into
devfrom
use-in-process-cache-store
Sep 21, 2026
Merged

CloCkWeRX merged 2 commits into
devfrom
use-in-process-cache-store

Conversation

@Br3nda

@Br3nda Br3nda commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Why

The free Memcachier dev plan is failing in production. In a 14 minute window of growstuff-prod logs there were 58 IO::TimeoutError: Blocking operation timed out! against mc2.dev.ec2.memcachier.com:11211.

With socket_timeout: 0.5, every one of those is a half-second stall followed by a full uncached render. So right now the cache costs us latency and buys us nothing — on a dyno that is simultaneously being OOM-killed every ~9 minutes (R14 → R15 → SIGKILL 137).

What this does

Memcached becomes opt-in. Production uses it only when USE_MEMCACHED=true and MEMCACHIER_SERVERS is set. Otherwise it falls back to :memory_store, capped at 32 MB and tunable with RAILS_CACHE_SIZE_MB.

Deploying this restores working caching immediately, with no addon upgrade and no spend. If we later move to a paid Memcachier plan, setting USE_MEMCACHED=true switches back with a config var and a restart — no deploy.

USE_MEMCACHED MEMCACHIER_SERVERS Store
unset set memory_store (32 MB)
unset unset memory_store (32 MB)
true set mem_cache_store
true unset memory_store (32 MB)

That last row matters: the old code did (ENV["MEMCACHIER_SERVERS"] || "").split(",") and handed Dalli an empty server list if the var ever went missing. Now an empty list can't reach Dalli.

Is a per-process cache safe here?

Yes, as far as I can tell:

  • Sessions are unaffected — config/initializers/session_store.rb uses :cookie_store, not the cache.
  • Every Rails.cache.fetch is regenerable and either keyed by cache_key_with_version (so writes invalidate naturally, across processes) or carries an expires_in.
  • Rack::Attack is unaffected — it already uses its own in-process MemoryStore, deliberately, per the comment added in Stop Rack::Attack depending on memcached, and cut the memcached timeout #4800.

The one exception is homepage_stats, the only fragment cached under a static key with no TTL, relying solely on expire_fragment. With a per-process store that only clears the worker that handled the write, so this PR gives it expires_in: 1.hour to bound how long another worker can serve a stale copy. The counts in it are member/crop/planting/garden totals, so an hour is plenty fresh.

Memory impact

Worth being explicit, since the dyno is memory-constrained: this store is per Puma worker, so the cost is RAILS_CACHE_SIZE_MB × WEB_CONCURRENCY. At the current WEB_CONCURRENCY=2 that's up to 64 MB; at WEB_CONCURRENCY=1 it's up to 32 MB. That's a real cost, but it replaces re-rendering every page on every request, which is the more expensive of the two.

Testing

  • rubocop config/environments/production.rb — clean.
  • haml-lint app/views/home/_stats.html.haml — 8 lints, byte-identical to the 8 already on dev (same offences, shifted 2 lines by the added comment). No new lints.
  • rspec spec/views/home — 12 examples, 0 failures.
  • Verified the store-selection logic against all five env-var combinations in the table above.
  • rspec spec/requests/rack_attack_spec.rb — 2 failures, but these are pre-existing: I confirmed they fail identically on a clean tree with this branch stashed. They're the honeypot and excessive-crawling ban tests. Unrelated to this change, but see below.

Not in this PR

This is one of several things making production unstable and it is not the root cause. The dyno is being OOM-killed because config/puma.rb was switched to clustered mode with 2 workers + preload_app! in dc8503c (Sep 17), tripling the process count on a 512 MB Basic dyno. WEB_CONCURRENCY=1 is the immediate mitigation there.

Also flagging that the two failing Rack::Attack ban specs are worth a look on their own — crawler bans are held in per-process memory and the dyno currently restarts every ~9 minutes, so long-window bans never accumulate in production either.

BUT ALSO:

i did this tonight, which sseems to have help:

$ heroku config:set WEB_CONCURRENCY=1 -a growstuff-prod
Setting WEB_CONCURRENCY and restarting ⬢ growstuff-prod... done, v511
WEB_CONCURRENCY: 1

$ heroku config:set RAILS_MAX_THREADS=3 -a growstuff-prod
Setting RAILS_MAX_THREADS and restarting ⬢ growstuff-prod... done, v512
RAILS_MAX_THREADS: 3

$ heroku config:set MALLOC_ARENA_MAX=2 -a growstuff-prod
Setting MALLOC_ARENA_MAX and restarting ⬢ growstuff-prod... done, v513
MALLOC_ARENA_MAX: 2

🤖 Generated with Claude Code

The free Memcachier dev plan is timing out on every call in production
(58 IO::TimeoutError in a 14 minute log window). Each request pays the
0.5s socket timeout and then renders uncached anyway, so the cache is
currently pure overhead on a dyno that is already being OOM-killed.

Memcached is now used only when USE_MEMCACHED=true and MEMCACHIER_SERVERS
is present. Otherwise production uses a memory store capped at 32MB via
RAILS_CACHE_SIZE_MB, so deploying restores real caching without the
timeouts and without an addon upgrade.

The memory store is per process, so homepage_stats gets an expires_in to
bound staleness. It is the only fragment that relies on expire_fragment
rather than a versioned cache key.

Sessions are cookie-based and every other cached value is either keyed by
cache_key_with_version or carries a TTL, so nothing depends on the cache
being shared between processes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Br3nda
Br3nda requested a review from CloCkWeRX September 20, 2026 09:42
Comment thread app/views/home/_stats.html.haml Outdated
@CloCkWeRX
CloCkWeRX merged commit 58fb44f into dev Sep 21, 2026
16 checks passed
@CloCkWeRX
CloCkWeRX deleted the use-in-process-cache-store branch September 21, 2026 08:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants