fix: release monitors on PostgreSQL database drops so that they don't interrupt the drop - #195
Merged
Merged
Conversation
…gresql PostgreSQL refuses to drop a database that other sessions are connected to. ActiveRecord disconnects its own connections before dropping, but the wrapper's topology monitor, which is shared by the cluster's connections, kept its own connections to the database it was started for. db:drop, db:reset, db:test:prepare, and db:purge therefore failed with "database is being accessed by other users" whenever an Aurora dialect was in use. AwsPostgreSQLAdapter#drop_database now stops the topology monitor through the new WrapperPgConnection#stop_topology_monitor before dropping. The next connection that needs topology starts a new monitor. MySQL does not refuse to drop a database in use, so aws_mysql2 is unchanged.
bin/rails test checks the schema from the test process, which starts a topology monitor there, then clears all ActiveRecord connections and runs db:test:prepare in a child process to drop and reload the test database. The monitor in the parent process kept its connection to the test database, so the child's DROP DATABASE failed with "database is being accessed by other users". Stopping the monitor in drop_database cannot help, because that runs in the child. The ActiveRecord adapters now prepend a ConnectionHandler extension that, after clear_all_connections!, stops the wrapper's monitors and clears the cached topology, so the next connection starts a new monitor straight away. StorageService#registered? lets it skip the cache when no wrapper connection has registered it.
…opology-monitor # Conflicts: # CHANGELOG.md
…atabase The Blue/Green status providers keep their own monitoring connections to the database, outside ActiveRecord's pools and outside the monitor service, so with the bg plugin enabled they blocked DROP DATABASE the same way the topology monitor did. drop_database and the clear_all_connections! extension now also stop them through BlueGreenPlugin.clean_up_providers, and the extension clears the cached Blue/Green status along with the cached topology. The next initial connection starts a new provider.
This was referenced Oct 6, 2026
Open
…ailover spec The spec warmed the topology cache before establish_connection, but clear_all_connections! now stops the monitors and clears that cache, so the warm-up was lost. If the new monitor's first connection raced with the reader outage, it never learned the writer and failover timed out (seen on MySQL in CI). Wait for the connection's own monitor to cache every instance before cutting connectivity instead.
sophia-bq
previously approved these changes
Oct 6, 2026
sophia-bq
approved these changes
Oct 6, 2026
JuanLeee
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
fix: release monitors on PostgreSQL database drops so that they don't interrupt the drop
Description
With an Aurora dialect (
aurora-pg, auto-detected or set viawrapper_dialect), every ActiveRecord workflow that drops a PostgreSQL database fails onaws_postgresqlwith:That covers
db:drop,db:reset,db:purge,db:test:prepare, andbin/rails testafter a schema change. The stockpostgresqladapter, andaws_postgresqlwithwrapper_dialect: pg, both work.Root cause: PostgreSQL refuses to drop a database that other sessions are connected to. ActiveRecord disconnects its own connections first, but the wrapper's topology monitor keeps its own connection to the database it was started for. The monitor is shared by every connection with the same
cluster_id. Confirmed against Aurora by countingpg_stat_activitysessions on the target database across all instances: 1 session remained after ActiveRecord disconnected, and it was the monitor's.The monitor can be in one of two places:
bin/rails db:drop. ActiveRecord drops through a maintenance connection that shares the monitor.bin/rails testchecks the schema from the test process, which starts a monitor there. It then callsclear_all_connections!and runsbin/rails db:test:preparein a child process to purge and reload the test database. The parent's monitor blocks the child's drop.Fix:
AwsPostgreSQLAdapter#drop_databasestops the topology monitor before runningDROP DATABASE, through a newWrapperPgConnection#stop_topology_monitor. That delegates to the host list provider's existingstop_monitor(a no-op for non-topology providers).ConnectionHandlerextension (AwsConnectionHandler). Afterclear_all_connections!, it stops the wrapper's monitors and clears the cached topology, so the wrapper is back in the state of a fresh process and the next connection starts a new monitor straight away.clear_all_connections!is ActiveRecord's "close everything" call, used by test tooling and database tasks;flush_idle_connections!and normal pool reaping are not affected. The extension skips the topology cache via the newStorageService#registered?when no wrapper connection has registered it.aws_mysql2needs no adapter change: MySQL drops a database with other sessions connected (confirmed live). It gets theclear_all_connections!extension because both adapters load it.bgplugin) also keep their own monitoring connections to the database, outside the monitor service, sodrop_databaseand theclear_all_connections!extension now also stop them viaBlueGreenPlugin.clean_up_providers, and the extension clears the cached Blue/Green status. The next initial connection starts a new provider. If an app callsclear_all_connections!during a switchover, monitoring restarts on the next connection, but traffic is not held in between; this is acceptable for ActiveRecord's "close everything" call, which is normally used by test tooling and database tasks.Tests:
WrapperPgConnection#stop_topology_monitorstops the provider's monitor without going through the plugins.AwsPostgreSQLAdapter#drop_databasestops the monitor before executingDROP DATABASE.clear_all_connections!stops the monitors after disconnecting the pools, clears the cached topology, passes the role through, and does not fail when the topology cache was never registered.flush_idle_connections!does not stop the monitors.StorageService#registered?.aurora-pgdialect, development and test databases:bin/rails testafter adding a migration passes (1 run, 0 failures). It fails onmain, and fails with only thedrop_databasechange.db:drop,db:prepare,db:reset, anddb:test:prepareall succeed.clear_all_connections!, the next query opens a new connection and a new topology monitor starts.wrapper_plugins: bg,failover,initial_connection(no Blue/Green deployment needed; the status monitors connect regardless). Without this change,bin/rails test,db:drop, anddb:resetfailed withbeing accessed by other users; with it, all pass, and a new Blue/Green provider starts on the next connection afterclear_all_connections!.By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.