Skip to content

Re-read pulls updated in the last 20 minutes - #503

Open
probablyian wants to merge 3 commits into
masterfrom
periodic-org-search-refresh
Open

probablyian wants to merge 3 commits into
masterfrom
periodic-org-search-refresh

Conversation

@probablyian

@probablyian probablyian commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Connects to #502

Some QA and CR stamps never show up on the board. The webhook carrying them never arrived, and GitHub doesn't resend it. Until now only a restart or the card's refresh button fixed that, and a restart skips repos missing from the config. This is option 6 from #502, run every 20 minutes instead of 10.

Summary

  • Sweep: every 20 minutes the server searches for pulls updated since an hour before the last sweep, open or closed. The hour covers search finding an update 15 or more minutes late. It re-reads each hit the way the card's refresh button does, skipping ones it already re-read at the same updated_at.
  • Repos missing from the config: the search goes by owner (user:ifixit), so fixbot, expo, DocHarvestor and the rest get re-read too.
  • Dropped closes: a pull whose close webhook was lost leaves the board within one sweep instead of at the next restart.
  • Restarts: the first sweep after a restart reaches back to an hour before the newest pull update in the DB, up to a day. That covers events dropped while the server was down (#502 has the outages).

It costs 3 search calls an hour plus 8 to 15 REST calls per changed pull. 234 pulls changed from 2026-09-26 00:00Z through 2026-09-29. No config change is needed.

pulldasher-dev and the shared quota

pulldasher-dev will run its own sweep once someone deploys this there with update-pulldasher-dev. Its config has a different GitHub token from production's, but GitHub counts the quota per account, not per token. If both tokens belong to the same account, dev's re-reads come out of production's 5,000 calls an hour. I haven't checked who owns each token.

QA

  • Run npm test on Node 24 and confirm test/recent-pulls-sweep.test.js passes. CI only runs lint and build.
  • Before the restart, note SELECT MAX(date_updated) FROM pulls (epoch seconds).
  • Deploy this branch and restart pulldasher.
  • 20 minutes after the restart, find the log line refreshing N of M pulls updated since <time>.
  • <time> is an hour before that value, or 80 minutes before the restart if that is earlier.
  • 20 minutes after the restart, confirm the container logs have no Failed to search for pulls updated since line.
  • Push a commit to an open pull in a repo missing from the config.
  • Confirm the board shows that pull's new head within 25 minutes:
curl -s -H "Authorization: Bearer $(gh auth token)" https://pulldasher.cominor.com/api/v1/pulls \
  | jq '.pulls[] | select(.repo == "iFixit/<repo>" and .number == <N>) | .head_sha'

https://claude.ai/code/session_01Loy5Y4CXz9jacGkTPFENdw

QA and CR stamps go missing from the board when a webhook never
arrives, and GitHub doesn't resend it. Until now only a restart or the
card's refresh button re-read the pull, and a restart only covers repos
listed in config.repos. #502 has the pulls the board missed on
2026-09-29, including 8 in repos the config leaves out.

Every 20 minutes the server now searches each owner in config.repos
for pulls updated since the previous sweep started, open or closed:

  user:ifixit is:pr updated:>=2026-09-29T22:37:07Z

and refreshes each hit through refresh.pull, one at a time. Searching
by owner reaches every repo that owner holds, configured or not. Search
returns closed pulls too, so a dropped close webhook heals within one
sweep.

Each window starts 5 minutes before the previous sweep started, since
GitHub doesn't document how long search indexing takes. The first
sweep runs 20 minutes after startup and reaches back to 25 minutes
before startup, so it also covers a restart shorter than that. A failed search keeps its
window for the next sweep. A pull that fails to refresh is logged by
refresh.pull and waits for its next update.

The issue suggested 10 minutes; this uses 20. That's 3 search calls an
hour plus the re-reads, 8 to 15 REST calls per pull by #502's count.

It runs one search per owner instead of joining the qualifiers. With
advanced_search=true, "user:iFixit user:octocat" is an AND: it found 0
pulls where the default search found 38.

Verified against the live API with the new searchUpdatedPulls: a search
since 2026-09-26 00:00Z paged through all 234 pulls GitHub counted, and
a first sweep refreshed the same 6 pulls gh's search returned.

Note: every hit is re-read, including old closed pulls that get a late
comment or label. In that 234-pull sample, 9 were closed more than 25
minutes before their last update. A bulk edit of many old pulls would
re-read every one of them in a single sweep.

Connects to #502

Claude-Session: https://claude.ai/code/session_01Loy5Y4CXz9jacGkTPFENdw
@probablyian
probablyian marked this pull request as ready for review September 29, 2026 23:20
The first sweep after a restart only reached back to 25 minutes before
startup. Events dropped during a longer outage stayed lost in repos the
startup refresh skips. Closes from the morning before
cominor was rebuilt on 2026-08-28 (ops#765 at 03:38Z, DocHarvestor#242
at 03:53Z, host booted at 16:46Z) are still open on the board.

At boot the server now reads MAX(pulls.date_updated) before any webhook
or the startup refresh saves a pull. That is about when the previous run
last saved a pull. The first sweep's window starts 5 minutes before it
instead of 25 minutes before startup, but never more than 24 hours back.
When the newest update is less than 20 minutes old, nothing changes.

#502 counted 111 org pulls changed in 24 hours, so a full day of
catch-up is 900 to 1,700 REST calls at 8 to 15 per pull, inside the
5,000 an hour quota.

An outage the process stays up through was already covered. A failed
search keeps its window, so the first sweep once the network is back
reaches back to the last sweep that worked. That was the 2026-09-29
case, when cominor lost its network from about 10:50Z to 14:55Z.

Note: date_updated is only written when pulldasher saves the pull row,
not for a comment, so MAX can be older than the last event handled.
That makes the window longer, never shorter.

Connects to #502

Claude-Session: https://claude.ai/code/session_016ZaR6H662c518wQxrq7kBC
GitHub search can find an update late. Its results show a pull's current
updated_at, but the updated: filter runs on an index that can lag behind
it. At 00:36Z on 2026-09-30, 2 of 17 org pulls updated in the previous
30 minutes weren't findable by their own updated_at. Both had a comment
as their newest event, which is how a QA or CR stamp arrives:

  iFixit/ops#1119      updated 00:22:23Z, findable about 15 min later
  iFixit/ifixit#64954  updated 00:07:59Z, not findable 33 min later

With a 20-minute interval and a 5-minute overlap, an update that shows
up L minutes late is caught only if the next sweep starts within 25
minutes of it. L = 15 is a coin flip, and L >= 25 is never caught. A
dropped comment webhook in that gap stayed dropped.

The overlap is now 60 minutes. Each sweep remembers the updated_at every
pull had when it re-read it, and skips a hit with the same updated_at.
The wider window costs search results, not REST calls. A 25-minute
window at 00:31Z held 14 pulls, so an 80-minute one should still fit in
one page of 100.

A pull whose refresh fails isn't marked as re-read, so the next sweep
that finds it tries again.

Solutions considered:

- Widen the overlap alone: every changed pull would be re-read up to 4
  times, at 8 to 15 REST calls each.
- Skip search and list recently updated pulls per repo: #502 option 4,
  which misses repos not in config.repos.

Connects to #502

Claude-Session: https://claude.ai/code/session_016ZaR6H662c518wQxrq7kBC
@probablyian

Copy link
Copy Markdown
Member Author

QA 🟢

The tests pass, and live sweeps found every pull a per-repo scan found. The deploy steps haven't run: update-pulldasher runs docker, and I don't have docker access on the host. A deploy leaves 68 stale rows from repos missing from the config.

Check at ad07848 Result
npm test on Node 24.21.0 71 of 71 pass, recent-pulls-sweep.test.js 7 of 7
CI Build passes
First sweep after a restart, seeded with 2026-09-29 10:28:56Z, the board's last update before the outage 111 pulls, including all 13 the outage dropped
Default first sweep, window 23:22:41Z to 00:42:41Z 26 pulls
Newest 30 pulls of all 307 unarchived iFixit repos, same window 25, all in the sweep
Second sweep right after the first 20 hits, 0 re-read
Pulls search indexed late iFixit/ops#1119 and iFixit/ifixit#64954, both in the sweep
How the sweeps ran

createRecentPullsSweep ran with a personal token and refreshApi.pull swapped for a recorder, so nothing was written to a DB. Whether the production token can search every private repo is unchecked.

iFixit/ops#1119 was updated at 00:22:23Z and became findable by that time about 15 minutes later. iFixit/ifixit#64954 was updated at 00:07:59Z and still wasn't findable by it at 00:47Z. The sweep found it anyway, because its window also covers the pull's previous indexed time, 23:56Z.

The sweep's 26th pull, iFixit/ops#1120, changed again after it ran, so the scan dated it outside the window.

Stale rows the deploy won't clear

70 of the 316 open rows on the board are closed on GitHub or show an older head. A restart fixes the 2 in configured repos, iFixit/ifixit#64895 and iFixit/ifixit#64885. The other 68 stay, because findStaleOpenPulls only checks configured repos. The first sweep reaches back an hour before the newest update in the DB, which will be minutes old at deploy.

The 68 rows by repo, and clearing them

Redelivering the failed 2026-09-29 deliveries would clear 10 of them, and bin/refresh-pull on each clears the rest. #502 has when each batch was dropped.

Repo Closed Open, older head
ops 41 2 (#114, #322)
DocHarvestor 12 0
ifixit-schooner-fw 4 0
fixbot 3 0
ifixit-cli 2 0
fixhub-app 2 0
valkyrie-next 1 0
expo 1 0

Post-deploy QA Steps

  • Before the restart, note SELECT MAX(date_updated) FROM pulls (epoch seconds).
  • Deploy with update-pulldasher --run and restart.
  • 20 minutes later, the log has refreshing N of M pulls updated since <time> and no Failed to search for pulls updated since line.
  • <time> is an hour before the noted value, or 80 minutes before the restart if that is earlier.
  • iFixit/ifixit#64895 is gone from the board, and iFixit/ifixit#64885 shows GitHub's current head.
  • Push to an open pull in an unconfigured repo. Within 25 minutes, the PR body's curl shows the new head.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant