Skip to content

fix(agents): reload the launchd agents after an upgrade; tests stop writing the real log - #15

Merged
mbeczynski merged 1 commit into
mainfrom
fix/agents-after-upgrade
Oct 6, 2026
Merged

mbeczynski merged 1 commit into
mainfrom
fix/agents-after-upgrade

Conversation

@mbeczynski

Copy link
Copy Markdown
Contributor

The problem

After brew upgrade from 1.3.0 to 1.3.1, launchd stopped starting the agents from the replaced bundle (job state = spawn failed, last exit reason = OS_REASON_CODESIGNING, nothing in any log). The watchdog and image attach went silent. buffer-guard kept running the old code.

The fix: AgentRepair

Reloading is done with bootout + bootstrap, retried.

  • After an upgrade: when the app starts on a new version (Homebrew quits it and reopens it after the upgrade), it reloads the agents that are safe to restart: backup-health, gdrive-attach, buffer-guard.
  • When an agent cannot start: it is reloaded by the app every few minutes and by the watchdog every 30 minutes. This covers gdrive-buffer without restarting a working mount: it is touched only once it has died.
  • drive-status gets an Agents: line.

Also

  • Tests no longer write to the real cloudmachine.log. Before this, a BACKUP FAILURE: … test canary line, which reads like an alarm, and lock messages landed there. Checked: a full swift test adds 0 lines.
  • Build number: the release passes CM_BUILD_NUMBER. v1.3.1 got a date because git failed inside build-app, which now says why when that happens.
  • Docs: the cask and docs no longer claim agents start by themselves after an upgrade.

Checks

  • swift test 329/0, lint passes, brew style on the cask passes.
  • The same bootout + bootstrap repair was done by hand on the production Mac: the agents start again and the mount stayed up.

Merging releases 1.3.2 automatically.

🤖 Generated with Claude Code

…riting the real log

After `brew upgrade` (1.3.0 -> 1.3.1) launchd refused to start the agents
from the replaced bundle: `job state = spawn failed`, `OS_REASON_CODESIGNING`,
no log line. The backup watchdog and the image attach went silent, and
buffer-guard kept running the old code. The cask said this could not happen.

AgentRepair reloads them (`bootout` + `bootstrap`, retried):
- when the app starts on a new version - Homebrew quits and reopens it for
  an upgrade - the agents that are safe to restart are reloaded;
- any agent launchd cannot start is reloaded, by the app every few minutes
  and by the backup watchdog every 30. That covers gdrive-buffer without
  restarting a working mount: it is touched only once it has died.
`drive-status` shows an `Agents:` line.

Also:
- Under XCTest the logger writes to a temporary directory. `swift test`
  appended lock messages and a "BACKUP FAILURE: ... test canary" line to the
  real cloudmachine.log, where it reads like an alarm.
- The release passes the commit count as CM_BUILD_NUMBER; v1.3.1 got a date
  because git failed inside build-app, which now says why when it happens.
- Cask, packaging README and README no longer claim agents start by
  themselves after an upgrade.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mbeczynski
mbeczynski merged commit 1354f34 into main Oct 6, 2026
4 checks passed
@mbeczynski
mbeczynski deleted the fix/agents-after-upgrade branch October 6, 2026 08:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant