Skip to content

fix: release monitors on PostgreSQL database drops so that they don't interrupt the drop - #195

Merged
aaron-congo merged 8 commits into
mainfrom
fix/ar-db-drop-with-topology-monitor
Oct 6, 2026
Merged

aaron-congo merged 8 commits into
mainfrom
fix/ar-db-drop-with-topology-monitor

Conversation

@aaron-congo

@aaron-congo aaron-congo commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

fix: release monitors on PostgreSQL database drops so that they don't interrupt the drop

Description

With an Aurora dialect (aurora-pg, auto-detected or set via wrapper_dialect), every ActiveRecord workflow that drops a PostgreSQL database fails on aws_postgresql with:

PG::ObjectInUse: ERROR:  database "<name>" is being accessed by other users
DETAIL:  There is 1 other session using the database.

That covers db:drop, db:reset, db:purge, db:test:prepare, and bin/rails test after a schema change. The stock postgresql adapter, and aws_postgresql with wrapper_dialect: pg, both work.

Root cause: PostgreSQL refuses to drop a database that other sessions are connected to. ActiveRecord disconnects its own connections first, but the wrapper's topology monitor keeps its own connection to the database it was started for. The monitor is shared by every connection with the same cluster_id. Confirmed against Aurora by counting pg_stat_activity sessions on the target database across all instances: 1 session remained after ActiveRecord disconnected, and it was the monitor's.

The monitor can be in one of two places:

  1. The process running the drop, e.g. bin/rails db:drop. ActiveRecord drops through a maintenance connection that shares the monitor.
  2. Another process. bin/rails test checks the schema from the test process, which starts a monitor there. It then calls clear_all_connections! and runs bin/rails db:test:prepare in a child process to purge and reload the test database. The parent's monitor blocks the child's drop.

Fix:

  • Case 1: AwsPostgreSQLAdapter#drop_database stops the topology monitor before running DROP DATABASE, through a new WrapperPgConnection#stop_topology_monitor. That delegates to the host list provider's existing stop_monitor (a no-op for non-topology providers).
  • Case 2: both adapters prepend a small ConnectionHandler extension (AwsConnectionHandler). After clear_all_connections!, it stops the wrapper's monitors and clears the cached topology, so the wrapper is back in the state of a fresh process and the next connection starts a new monitor straight away. clear_all_connections! is ActiveRecord's "close everything" call, used by test tooling and database tasks; flush_idle_connections! and normal pool reaping are not affected. The extension skips the topology cache via the new StorageService#registered? when no wrapper connection has registered it.
  • aws_mysql2 needs no adapter change: MySQL drops a database with other sessions connected (confirmed live). It gets the clear_all_connections! extension because both adapters load it.
  • The Blue/Green status providers (bg plugin) also keep their own monitoring connections to the database, outside the monitor service, so drop_database and the clear_all_connections! extension now also stop them via BlueGreenPlugin.clean_up_providers, and the extension clears the cached Blue/Green status. The next initial connection starts a new provider. If an app calls clear_all_connections! during a switchover, monitoring restarts on the next connection, but traffic is not held in between; this is acceptable for ActiveRecord's "close everything" call, which is normally used by test tooling and database tasks.

Tests:

  • New unit examples, all of which failed before the change:
    • WrapperPgConnection#stop_topology_monitor stops the provider's monitor without going through the plugins.
    • AwsPostgreSQLAdapter#drop_database stops the monitor before executing DROP DATABASE.
    • clear_all_connections! stops the monitors after disconnecting the pools, clears the cached topology, passes the role through, and does not fail when the topology cache was never registered. flush_idle_connections! does not stop the monitors.
    • StorageService#registered?.
  • Unit suite passes on the ActiveRecord 7.2, 8.0, and 8.1 gemfiles; RuboCop is clean.
  • Manual check with a Rails 8.1.4 app against Aurora PostgreSQL 17 (3 instances), aurora-pg dialect, development and test databases:
    • bin/rails test after adding a migration passes (1 run, 0 failures). It fails on main, and fails with only the drop_database change.
    • db:drop, db:prepare, db:reset, and db:test:prepare all succeed.
    • After clear_all_connections!, the next query opens a new connection and a new topology monitor starts.
  • Repeated the manual check with wrapper_plugins: bg,failover,initial_connection (no Blue/Green deployment needed; the status monitors connect regardless). Without this change, bin/rails test, db:drop, and db:reset failed with being accessed by other users; with it, all pass, and a new Blue/Green provider starts on the next connection after clear_all_connections!.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

…gresql

PostgreSQL refuses to drop a database that other sessions are connected to. ActiveRecord disconnects its own connections before dropping, but the wrapper's topology monitor, which is shared by the cluster's connections, kept its own connections to the database it was started for. db:drop, db:reset, db:test:prepare, and db:purge therefore failed with "database is being accessed by other users" whenever an Aurora dialect was in use.

AwsPostgreSQLAdapter#drop_database now stops the topology monitor through the new WrapperPgConnection#stop_topology_monitor before dropping. The next connection that needs topology starts a new monitor. MySQL does not refuse to drop a database in use, so aws_mysql2 is unchanged.
bin/rails test checks the schema from the test process, which starts a topology monitor there, then clears all ActiveRecord connections and runs db:test:prepare in a child process to drop and reload the test database. The monitor in the parent process kept its connection to the test database, so the child's DROP DATABASE failed with "database is being accessed by other users". Stopping the monitor in drop_database cannot help, because that runs in the child.

The ActiveRecord adapters now prepend a ConnectionHandler extension that, after clear_all_connections!, stops the wrapper's monitors and clears the cached topology, so the next connection starts a new monitor straight away. StorageService#registered? lets it skip the cache when no wrapper connection has registered it.
@aaron-congo
aaron-congo requested a review from a team as a code owner October 6, 2026 01:49
…atabase

The Blue/Green status providers keep their own monitoring connections to the database, outside ActiveRecord's pools and outside the monitor service, so with the bg plugin enabled they blocked DROP DATABASE the same way the topology monitor did.

drop_database and the clear_all_connections! extension now also stop them through BlueGreenPlugin.clean_up_providers, and the extension clears the cached Blue/Green status along with the cached topology. The next initial connection starts a new provider.
@aaron-congo aaron-congo changed the title fix: release the topology monitor when needed so PostgreSQL databases can be dropped fix: release monitors on PostgreSQL database drops so that they don't interrupt the drop Oct 6, 2026
…ailover spec

The spec warmed the topology cache before establish_connection, but
clear_all_connections! now stops the monitors and clears that cache, so the
warm-up was lost. If the new monitor's first connection raced with the
reader outage, it never learned the writer and failover timed out (seen on
MySQL in CI). Wait for the connection's own monitor to cache every instance
before cutting connectivity instead.
sophia-bq
sophia-bq previously approved these changes Oct 6, 2026
Comment thread CHANGELOG.md Outdated
@aaron-congo
aaron-congo merged commit 68d0bd8 into main Oct 6, 2026
10 checks passed
@aaron-congo
aaron-congo deleted the fix/ar-db-drop-with-topology-monitor branch October 6, 2026 17:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants