Repository navigation
[Ulmo] Library v2 commit endpoint timeout with large Meilisearch indexes #38993
Description
Activity
I’d appreciate your input here @ormsbee
I think the Batch document submission as described in the proposed short-term fix would be a huge win.
But also I think reviewing the arch as described in option A for the proposed long-term fix is needed, although I don't know what was the rationale behind the original decision.
On a 500K-document index, each document takes 5-7 seconds to index (inherent to Meilisearch's inverted index rebuild for large indexes). This is not resource-bound — we tested with 30Gi RAM / 6 CPU and observed the same 7s per document. Three blocks = 21 seconds, exceeding the 15-second timeout.
This seems crazy. It's expected behavior for Meilisearch to take 5-7 seconds to add a single document? Are we changing something in the index configuration itself that would cause it to force a rebuild of the entire thing?
FYI @bradenmacdonald, @kdmccormick, @openedx/interest-performance
The library commit endpoint
/api/libraries/v2/<lib_key>/commit/returns a 500 error when publishing 3+ blocks on instances with large Meilisearch indexes (500K+ documents). The publish itself succeeds (data is committed in Learning Core), but the post-publish search index update exceeds the hardcoded timeout in dispatch_and_wait.OK, something weird is going on here.
First of all,
dispatch_and_waitis explicitly designed not to throw 500 errors when the indexing takes longer than the timeout. The entire point ofdispatch_and_waitis that if indexing takes longer than the timeout, the indexing will still get done in the background, and a 200 response will be returned to the user anyways, before the indexing actually completes.What exactly is the 500 error you're seeing?
A minor second point is that the
dispatch_and_waittimeout is 10 seconds, and has never been 15 (I checked the git history). So: which version of the code are you using? This timeout should have absolutely no effect on whether or not things end up getting indexed or not. It just affects the user experience, and whether or not the frontend knows when to refresh the results (once the 200 response is returned, the frontend will refresh, so we try but don't insist on waiting to send the 200 only after the index is updated).On a 500K-document index, each document takes 5-7 seconds to index (inherent to Meilisearch's inverted index rebuild for large indexes). This is not resource-bound — we tested with 30Gi RAM / 6 CPU and observed the same 7s per document. Three blocks = 21 seconds, exceeding the 15-second timeout.
I agree that this sounds problematic. I will note that in the past we have had performance issues that looked like Meilisearch issues but turned out to be bugs on the
edxappside. Have you measured this using the pure Meilisearch API directly?And again, the timeout should not cause the indexing to fail.
Meilisearch v1.8 with ~500K documents in tutor_studio_content index
Meilisearch has released a lot of improvements since version 1.8 came out on May 7, 2024. In particular, version 1.12 introduced a next-generation indexer that's supposed to be 4x faster. So the very first thing I would do is measure indexing speed directly (using plain Meilisearch APIs) on a recent version of Meilisearch.
The fix: use Meilisearch only for search functionality (full-text search, filtering, facets). For the library component listing (the default view showing all components), fetch directly from the Learning Core relational database via a dedicated REST endpoint.
This is making a big assumption that library authors are mostly using "the default view" without any full-text search, filtering, or facets applied. I'm not sure that's true, especially for relatively large libraries. I would assume that it's common to have at least some basic filter or facet applied, like showing only components, or viewing the contents of a single collection. As soon as that's the case, your suggested Option A no longer really helps (because then the results would still be using Meilisearch), and it just makes the code more complicated by having two different retrieval paths.
Batch document submissions: In send_events_after_publish (or send_change_events_for_modified_entities on master), collect all documents from the publish operation and submit them to Meilisearch in a single update_documents call instead of one document at a time.
This would be nice, and I'd like to do it, but we're a bit architecturally constrained at the moment since libraries and search are separate apps which rely on events to keep in sync. Achieving this would either require replacing the
LIBRARY_BLOCK_CREATED,LIBRARY_BLOCK_UPDATED, etc. events with aLIBRARY_BLOCKS_CHANGEDbatch event, or update the low-levelENTITIES_DRAFT_CHANGEDevent to include data about context (library ID / course ID) and opaque keys. Either way is a fairly significant breaking change of all the related events.Option B: Fully async indexing with per-component loading state in the frontend
This approach I really like, and I don't think there's any technical reason not to do it. It just needs the investment to make it happen. Although I would say that it probably shouldn't be done as a one-off thing and should be approached thoughtfully so it becomes a pattern that we can use throughout other APIs and other parts of the CMS/LMS UI. (Instructor tasks, import/export, etc.)
- changed the title
[-]Library v2 commit endpoint timeout with large Meilisearch indexes[/-][+][Ulmo] Library v2 commit endpoint timeout with large Meilisearch indexes[/+]on Aug 17, 2026 The library commit endpoint
/api/libraries/v2/<lib_key>/commit/returns a 500 error when publishing 3+ blocks on instances with large Meilisearch indexes (500K+ documents). The publish itself succeeds (data is committed in Learning Core), but the post-publish search index update exceeds the hardcoded timeout in dispatch_and_wait.OK, something weird is going on here.
First of all,
dispatch_and_waitis explicitly designed not to throw 500 errors when the indexing takes longer than the timeout. The entire point ofdispatch_and_waitis that if indexing takes longer than the timeout, the indexing will still get done in the background, and a 200 response will be returned to the user anyways, before the indexing actually completes.What exactly is the 500 error you're seeing?
A minor second point is that the
dispatch_and_waittimeout is 10 seconds, and has never been 15 (I checked the git history). So: which version of the code are you using? This timeout should have absolutely no effect on whether or not things end up getting indexed or not. It just affects the user experience, and whether or not the frontend knows when to refresh the results (once the 200 response is returned, the frontend will refresh, so we try but don't insist on waiting to send the 200 only after the index is updated).Sorry, I mixed the code between master and the Ulmo release. All of this issue is happening on Ulmo. I've updated the ticket to reference the correct location where the error is thrown — tasks.py on release/ulmo. The except TimeoutError there doesn't catch celery.exceptions.TimeoutError, which is why it surfaces as a 500 (I exactly see this error: celery.exceptions.TimeoutError: The operation timed out).
I see that on master this was refactored into dispatch_and_wait which correctly catches CeleryTimeout and passes gracefully — so the 500 is indeed an Ulmo-only bug. I understand that the indexing itself is not affected and it will eventually complete in the background. That said, from a UX perspective, when the response returns before indexing finishes, the frontend refreshes and shows stale data from Meilisearch — blocks appear stuck in an intermediate state until things catch up. This can be confusing for users since it looks like the publish didn't work.
Reacted by Braden MacDonaldOn a 500K-document index, each document takes 5-7 seconds to index (inherent to Meilisearch's inverted index rebuild for large indexes). This is not resource-bound — we tested with 30Gi RAM / 6 CPU and observed the same 7s per document. Three blocks = 21 seconds, exceeding the 15-second timeout.
I agree that this sounds problematic. I will note that in the past we have had performance issues that looked like Meilisearch issues but turned out to be bugs on the
edxappside. Have you measured this using the pure Meilisearch API directly?And again, the timeout should not cause the indexing to fail.
Meilisearch v1.8 with ~500K documents in tutor_studio_content index
Meilisearch has released a lot of improvements since version 1.8 came out on May 7, 2024. In particular, version 1.12 introduced a next-generation indexer that's supposed to be 4x faster. So the very first thing I would do is measure indexing speed directly (using plain Meilisearch APIs) on a recent version of Meilisearch.
Good points — I'll measure directly via the Meilisearch API and also test with v1.12+ to isolate. Will report back with results.
Reacted by Braden MacDonaldHello, I bring the results. I tested directly via the Meilisearch API (single document PUT to our tutor_studio_content index with ~500K documents, on v1.8):
{ "uid": 60950, "indexUid": "tutor_studio_content", "status": "succeeded", "type": "documentAdditionOrUpdate", "canceledBy": null, "details": { "receivedDocuments": 1, "indexedDocuments": 1 }, "error": null, "duration": "PT8.302285333S", "enqueuedAt": "2026-08-18T19:46:00.618974855Z", "startedAt": "2026-08-18T19:46:00.621689551Z", "finishedAt": "2026-08-18T19:46:08.923974884Z" }That confirms it takes ~8.3 seconds for a single document on the Meilisearch side alone, with no edxapp code involved.
Reacted by Braden MacDonaldI also tested batching — submitting 5 documents in a single API call to the same index:
{ "uid": 60952, "indexUid": "tutor_studio_content", "status": "succeeded", "type": "documentAdditionOrUpdate", "details": { "receivedDocuments": 5, "indexedDocuments": 5 }, "duration": "PT8.130113417S", "enqueuedAt": "2026-08-18T19:58:07.56261503Z", "startedAt": "2026-08-18T19:58:07.564130698Z", "finishedAt": "2026-08-18T19:58:15.694244115Z" }5 documents in a batch take essentially the same time (~8.1s) as a single document (~8.3s). This suggests that batching document submissions on the edxapp side would be a significant improvement — a 5-block publish would go from 5×8s = 40s (current serial approach) down to ~8s total.
Reacted by Braden MacDonaldDo you know if the insert performance degraded linearly over time, or if there was a sudden spike? Is it possible that it's hitting some memory limit?
Do you know if the insert performance degraded linearly over time, or if there was a sudden spike? Is it possible that it's hitting some memory limit?
We don't have historical data to say whether it degraded linearly over time — the only data we have is from the tests I performed. I ran more than 10 tests and the results are consistent: most of them are above 6 seconds of execution time per document.
Regarding memory limits: we tested extensively. The Meilisearch process itself peaks at ~116Mi RSS during indexing (measured via kubectl top). We tested with memory allocations ranging from 2Gi up to 30Gi — the indexing time remained the same (~8s per document) regardless. The pod has never been OOM-killed (zero restarts). Storage is gp3 with 3000 baseline IOPS, so disk I/O shouldn't be a constraint either.
Based on all this, it seems inherent to the index rebuild cost at this document count rather than hitting any resource limit. That said, we haven't tested with a newer Meilisearch version yet — that's our next step.
Follow-up: Meilisearch v1.8 vs v1.12 benchmark
Following the suggestion to test a newer Meilisearch version, I benchmarked single-document indexing directly via the Meilisearch API (no edx-platform in the path) on our tutor_studio_content index (~500K documents).
Method: PUT
/indexes/tutor_studio_content/documentswith a single fresh document (create → delete → create each run), then poll GET/tasks/<uid>and read the duration field. Median of 10 runs.Results:
Version Per-document indexing time (median of 10) v1.8 ~5.7s (range 5.58-6.10s) v1.12.8 ~4.2s (range 4.11-4.40s) v1.12 gives roughly a 25-30% improvement for single-document inserts. Not the 4x advertised for bulk incremental imports — that speedup doesn't apply to this single-document Studio-content workload (many filterable/sortable attributes, rebuilt per document).
Environment parity: The two versions were measured on different clusters (v1.8 on stage, v1.12 on SIT), but the relevant hardware is identical:
Same EBS storage: both gp3, 20GB, 3000 IOPS, 125 MB/s throughput
Same node instance family: m5a.2xlargeSo the ~25-30% difference is attributable to the Meilisearch version, not infrastructure.
Based on these benchmark results, my takeaway is that the version upgrade alone (v1.8 → v1.12) only gives ~25-30% improvement on single-document indexing (~5.7s → ~4.2s), which isn't enough to reliably solve the multi-block commit timeout — indexing cost still scales linearly per document.
From my point of view, the "Proposed Short-Term Fix" — batching document submissions — is the most promising path. Instead of indexing each published block one at a time (N × per-document cost), we'd collect all documents from a publish and send them in a single update_documents call, so the per-batch index rebuild overhead is paid once instead of N times.
I ran a clean benchmark to confirm this (v1.8, fresh inserts, ~500K-document index, median of 10 runs):
Approach 5-document total One-at-a-time ~28.5s (5 × 5.7s) Single batch ~7.42s The batch of 5 fresh documents (7.42s) is only marginally slower than a single document (5.7s), which confirms the per-operation index rebuild is paid mostly once per batch. That's ~3.8x faster for a 5-block publish, and it works on the current version with no upgrade needed — it brings a 5-block commit from ~28s (times out) down to ~7.4s (within even the original 15s window).
Braden pointed out what the effort would be here. Do you agree that batching is the right direction here, and is that effort estimate accurate from your perspective? @ormsbee @bradenmacdonald
8 remaining items
Quick thought: Are there incompatibilities that would prevent us from jumping all the way up to the current version (1.53.1)? Even 1.12 came out a year and a half ago.
Longer thought: This doesn't seem tenable in the long run. If indexing is becoming a multi-second operation at 500K items, pushing it async or even batching it doesn't seem like a sustainable long term solution. Even if we rearrange things to batch, are we willing to have an 8 second delay on any changes to the search index we build so much of the UX around? Will that grow to 20 when the index doubles in size? 40 when we add some new feature?
I've been operating on the assumption that we're doing something wrong somewhere, and that simply adding a document should not cause so much work to need to be done. I honestly still hope that's the case. I would expect something like this to happen in tens or low hundreds of milliseconds.
@dwong2708: Thank you so much for doing this investigation. Would it be possible for you to show the exact payload that's being sent (minus any problematic auth stuff)?
no problem, here the payload:
[ { "id": "test-batch-doc-001", "type": "library_block", "block_type": "html", "display_name": "Batch Test Block 1", "content": {"html_content": "Batch test document one for measuring batch indexing speed."}, "context_key": "lib:WGU:DW0001", "org": "WGU", "access_id": 1 } ]
I ran some benchmarks on my Macbook today to get a bit more insight into what's going on.
TL;DR: the issue is unfortunately quite real, and in fact worse than the 3-block case in the report suggests. Publishing ~30 components exceeds the timeout even on the newest Meilisearch. Both proposed short-term fixes work, and batching is the most important: a 100-block publish goes from 33s to 0.7s (Finding 8). Upgrading Meilisearch is also worth 4.3× on its own, with an in-place migration that takes 2.5 seconds. For the 3-block case in the report, batching plus an upgrade compound to roughly a 13× reduction.
Method
Claude generated a dependency-free benchmark suite (Docker + stdlib Python, no
openedx-platformor Tutor at all, not even the Meilisearch python lib). The tooling reproduces the Studio index configuration and document shapes, seeds an index to a target size, and measures the write operations the platform actually performs. Harness and raw results: https://github.com/open-craft/meili-bench- Index config transcribed from
index_config.py— the same 20 filterable attributes, 15 searchable attributes, distinct attribute and ranking rules. - Corpus: 500k synthetic documents matching the four shapes in
documents.py(274k library blocks, 176k course blocks, 35k containers, 15k collections) across 1,200 libraries and 800 courses. Text is drawn from a Zipf-distributed vocabulary so term frequencies resemble natural language — duplicating one sample document 500k times would produce pathological posting lists and would not measure realistic index maintenance. - Environment: Docker on an M1 Macbook, container limited to 4 CPU / 8 GB. One version at a time, using named volumes (not bind mounts).
- Measurements: p50/p95 over 20 reps after 3 warm-up reps, at index sizes of 50k / 100k / 250k / 500k documents. Latency is decomposed into Meilisearch's own
startedAt→finishedAt, queue wait, and client polling.
publish_3_sequentialreproduces today's publish path exactly: three separate documents, each sent in its ownupdate_documents()call and fully awaited before the next begins, because each is a separateLIBRARY_BLOCK_PUBLISHEDevent handled by its own synchronous task.Results at 500k documents (p50)
Operation 1.12.8 1.36.0 (Tutor today) 1.53.1 (latest) update_1— one block edited1.22s 417ms 294ms update_3— 3 blocks, one batched call1.38s 423ms 301ms update_100— 100 blocks, one call1.65s 805ms 671ms publish_3_sequential— today's path3.98s 1.32s 920ms add_1_new— new block created4.13s 3.48s 3.63s delete_filter_1— block deleted4.24s 3.32s 3.55s search23ms 28ms 23ms Scaling of the publish path across index sizes:
Index size 1.12.8 1.36.0 1.53.1 50k 465ms 239ms 193ms 100k 830ms 344ms 259ms 250k 1.97s 659ms 473ms 500k 3.98s 1.32s 920ms Finding 1 — the Meilisearch version matters, and most of the win is already shipped
1.12 → 1.53 is a 4.3× improvement on the publish path at 500k, and the gap widens with index size (2.4× at 50k, 4.3× at 500k) — so it helps most exactly where the problem bites.
The practically important part: Tutor already ships 1.36.0, which captures 2.66s of the 3.06s improvement. Anyone still on 1.8 or 1.12 should upgrade before anything else is considered. Moving from 1.36 to 1.53 is worth a further 400ms — real, but not the main event.
Upgrading seems to be purely beneficial with no regressions: search latency is unchanged (23-28ms), bulk ingest throughput is within noise, and the on-disk index is the same size (~8.6 GB at 500k) on all three versions.
Finding 2 — batching works, and is nearly free
Batch size barely affects cost. On 1.53.1 at 500k: one document costs 294ms, three cost 301ms, one hundred cost 671ms. The marginal cost of an additional document in a batch is ~4ms, against ~300ms for a separate awaited call.
So the issue's proposed fix #2 is worth 2.9× on 1.12 and 3.1× on 1.53, on its own, at 500k. Combined with an upgrade: 3.98s → 301ms, ~13×.
Worth noting where this fix belongs.
_update_index_docs()already accepts a list of documents, andupsert_xblock_index_doc()already batches a block and all its children into one call. The reason a 3-block publish makes three calls is upstream: each block emits its ownLIBRARY_BLOCK_PUBLISHEDevent, handled by its own synchronous task. Batching means collecting documents across the publish incontent_libraries, not changingcontent/search/api.py.Finding 3 — a share of the wait is our own polling, not Meilisearch
_wait_for_meili_task()polls with exponential backoff capped at 2 seconds:sleep_delay = 0.020 time.sleep(sleep_delay) current_status = client.get_task(info.task_uid) while current_status.status in ("enqueued", "processing"): time.sleep(sleep_delay) sleep_delay = min(sleep_delay * 1.5, 2.0) # <-- up to 2s asleep past completion current_status = client.get_task(info.task_uid)
Measuring with a tight 10ms poller and separately simulating the schedule above, the overhead for a 3-block publish ranges from 8% to 31% of total observed latency depending on where completion happens to fall in the backoff schedule (387ms of 1.31s on 1.53.1 at 500k; 578ms of 4.55s on 1.12.8). It averages around 20% and it is pure sleep past an already-finished task.
This gets worse as Meilisearch gets faster, because the backoff schedule is unchanged while the task shrinks. Capping the delay at ~250ms is a one-line change worth roughly a fifth of the remaining latency, and it composes with both other fixes.
Finding 4 — creates and deletes are the slow path
Creating a new document (3.63s) and deleting one (3.55s) are ~12× the cost of updating an existing one (294ms) on 1.53.1 at 500k — and unlike updates, they barely improved across versions (1.12×, 1.19×).
This does not affect publish. I checked the code: publish and edit paths call
update_documentson an existing document, and the publish handlers explicitly early-return when the published version is gone, leaving removal to the_DELETEDhandler. So publish is the fast case.It does mean that after the fixes above, creating and deleting blocks become the dominant remaining latency. One concrete inefficiency there:
delete_index_doc(..., delete_children=True)awaits a delete-by-id and then a
delete-by-filter — two sequential round trips, ~7s at 500k. Those could be a single filter expression.The slowness of creates has been partially explained: facet insertion is ~76% of the cost (2,758ms of 3,626ms at 500k on 1.53.1), it appears as soon as any filterable attribute exists rather than scaling with how many are declared, and it is insensitive to the number of facet values per document (Finding 10). A residual ~6× gap over updates persists even with zero filterable attributes, and that part is still unexplained.
Deletes remain unattributed. The ablation could not measure them without filterable attributes — delete-by-filter needs something to filter on — so there is no facet-free delete measurement to compare against. Ruled out so far: document type, and the delete-by-id vs delete-by-filter distinction (3.341s vs 3.352s, no difference).
Finding 5 — the index config matters
I ablated the index settings on both versions at 500k. The result was a bit surprising.
update_1, 500k documents:profile searchable filterable 1.12.8 1.53.1 full— whatindex_config.pyapplies15 named 20 1.221s 294ms no-filters15 named 0 1.070s 142ms minimal— stock Meilisearch defaults["*"]0 – 1.608s Two things fall out:
The 20 filterable attributes cost ~150ms per document — on both versions. 151ms on 1.12.8, 152ms on 1.53.1. (This is a 500k figure; at 100k the same comparison shows no measurable cost at all — see the caveat in Finding 10.) The absolute cost of facet maintenance did not change between versions at all; what got faster is everything else (1.070s → 142ms, a 7.5× speedup). That's why the facets look like 12% of the cost on 1.12 and 52% on 1.53 despite being the same work.
Restricting
INDEX_SEARCHABLE_ATTRIBUTESsaves ~1.5s per document. Stock Meilisearch defaults make every field searchable, and that is 11× more expensive than the 15-attribute list we specify — far outweighing what the facets cost. The current configuration is doing considerably more good than harm.So the searchable list is actively protective and should not be broadened casually. As for the filterable list — Finding 10 shows that trimming it would gain nothing, because the cost comes from having facets at all, not from how many attributes are declared.
Side note: dropping the filterable attributes also made bulk ingest 2.2× faster (487 → 1,087 docs/s) and search 2.9× faster (23ms → 8ms), so a full re-index gets cheaper too — but obviously at the cost of the filtering the Authoring MFE needs.
Finding 6 — upgrading is nearly free
Migrating the seeded 500k index in place from 1.12.8 to 1.53.1:
upgrade task duration: 2.47s boot to healthy: 1.2s documents after: 500,0002.5 seconds. No dump/restore, no re-index, no data loss. On this evidence the upgrade recommendation carries almost no operational risk.
One trap: the flag has been renamed. Older documentation says
--experimental-dumpless-upgrade; current builds want--upgrade-db/MEILI_UPGRADE_DB. Starting a new Meilisearch on an old database without it exits immediately with an incompatibility error rather than migrating.I also re-ran the whole matrix against the upgraded index. It performs comparably to a natively-seeded 1.53.1 index — some operations modestly faster, some modestly slower, within about ±25% — so there is no evidence an upgraded index needs to be rebuilt to get the benefit.
Finding 7 — follow-up experiments
All on 1.53.1 at 500k, on the upgraded index. Compare against
update_1= 335ms in the same run (run-to-run variance is roughly ±25%, so within-run ratios are more trustworthy than cross-run ones).Experiment Result Reading update_1_rewrite— replace all text1.274s vs 335ms Updates are cheap mainly because little changes. See caveat below. add_10_new— 10 new docs, one call3.964s vs 3.524s for one Creates are batchable too — ~49ms marginal per document publish_3_pipelined— enqueue 3, await once607ms vs 709ms sequential Only a 15% win; explicit batching (345ms) is 1.8× better again update_10001.933s Marginal cost holds at ~1.6ms/doc from 100 → 1000 delete_id_1vsdelete_filter_13.341s vs 3.352s No difference — delete-by-id is not cheaper update_1_block/_container/_collection348 / 331 / 286ms Document shape barely matters Three of these are worth spelling out.
Updates are cheap because the document barely changes. A full text rewrite costs 1.274s — 3.8× a typical small edit. The
update_1numbers throughout this write-up therefore represent a metadata-ish edit (title, timestamps, a modified line of text). A publish that copies draft content into the published fields sits somewhere between the two. Treat 335ms as a floor for the publish path and 1.27s as the ceiling for a substantial content edit.Pipelining is not a substitute for batching. Claude suggested that simply not awaiting between the three tasks might let Meilisearch auto-batch them and recover most of the win. It does not: across all 20 reps the three tasks landed in 2 distinct batches, not 1 — the first task starts processing immediately and only the second and third merge. Result: 607ms, against 345ms for the same three documents in a single
update_documents()call. Collecting the documents is worth doing properly.Two of Claude's other hypotheses that the data did not support: Delete-by-id is not faster than delete-by-filter, so merging the two round trips in
delete_index_doc(..., delete_children=True)is worth ~2× only because it is two awaits instead of one, not because of the filter. And Claude had hypothesized that container documents were expensive becausecontent.child_usage_keysis indexed as searchable text; containers are actually the cheapest of the three types measured. That hypothesis was wrong.Finding 8 — "Publish All" blows the timeout, on the fastest version tested
I had been measuring 3-block publishes because that is what the issue reports. That turns out to be a lucky small case. Publishing a whole library, on 1.53.1 at 500k, sequentially (today's path) versus the same documents batched:
blocks sequential (today) batched ratio 3 788ms 342ms 2.3× 25 7.91s 493ms 16× 50 16.81s 633ms 27× 100 32.98s 709ms 46× Sequential cost is dead linear at ~330ms per block. Batched cost is nearly flat: 100 blocks cost barely more than 3.
Publishing about 30 components exceeds the 10s timeout on the newest Meilisearch, on an idle 500k index. On 1.12 (~1.3s per block) the limit is roughly 8 components. Under concurrent load (Finding 9) it tightens further, to about 12 on 1.53.
This reframes the problem. The timeout is not really bounded by index size — it is bounded by how many blocks you publish at once, and any library large enough to be worth a "Publish All" button will exceed it. Batching does not merely improve this; it removes the scaling entirely.
Finding 9 — concurrency is fine, and search does not suffer
N writer threads each running the loop a Celery worker runs (mutate, PUT, await its own task), with a reader searching throughout. 1.53.1, 500k, 45s per configuration:
writers writes/s write p50 write p95 search p50 search p95 0 – – – 18ms 52ms 1 3.03 333ms 365ms 5ms 7ms 2 3.33 598ms 669ms 5ms 8ms 4 6.06 653ms 761ms 6ms 10ms 8 9.83 812ms 888ms 6ms 9ms Search latency is unaffected by write load — 5-6ms p50, under 10ms p95, whether one writer or eight. No errors at any level. (The writers=0 row ran first, immediately after seeding, so its 18ms/52ms is almost certainly a cold-cache artefact; every warm configuration sits at 5-6ms.)
Write latency degrades gracefully. 8× the concurrency costs 2.4× the latency while throughput rises 3.2×, so Meilisearch is merging concurrent writes rather than serializing them. But it does mean a block costs ~812ms rather than ~333ms on a busy instance, which is what tightens the publish limit to ~12 blocks.
Finding 10 — facet cost is not about how many attributes you declare
Separating the two possible causes, at 100k documents on 1.53.1:
Axis 1 — number of declared filterable attributes (identical documents):
filterable attributes update_1add_1_new0 89ms 239ms 5 90ms 727ms 10 90ms 730ms 20 89ms 715ms Axis 2 — facet values carried per document (all 20 attributes declared):
document facet values update_1add_1_newbare 0 118ms 703ms typical 9 76ms 711ms rich 98 286ms 781ms Two clean conclusions:
Trimming
INDEX_FILTERABLE_ATTRIBUTESwould gain nothing. Creates show a step change from 0 to 5 attributes (239ms → 727ms) and then nothing from 5 to 20; update cost is flat across the whole range at this index size. The cost is in having facets at all, not in how many attributes are declared. Since the platform needs facets, this is not a lever — which retracts the "be deliberate about the list" framing in Finding 5.One caveat on that update row, because it appears to contradict Finding 5. At 100k the facet cost on updates is unmeasurable (89-90ms across 0 to 20 attributes), yet the same 0-vs-20 comparison at 500k in Finding 5 shows ~150ms. So facet maintenance on updates is size-dependent — it emerges somewhere between 100k and 500k documents. It does not change the conclusion: the step is between 0 and 5 attributes either way, and dropping to zero is not an option.
Heavily-tagged content is ~4× slower to update. A document carrying 98 facet values costs 286ms against 76ms for a typical 9-value document. So instances with heavy taxonomy and collection use will see materially worse publish latency than these averages suggest — a plausible part of why the original reporter saw worse numbers than I could reproduce.
Creates, by contrast, are insensitive to value count (703/711/781ms): they pay a flat facet penalty regardless.
Caveat: 10 reps at 100k, single run. The bare-vs-typical inversion on updates (118ms vs 76ms) is within noise at this scale, so read Axis 2 as "rich is much more expensive", not as a precise curve.
Recommendations, mapped to the issue
Proposal in the issue Verdict Configurable CONTENT_LIBRARIES_PUBLISH_TIMEOUTUnlikely to be useful IMHO. Batch document submission Confirmed, and it scales: ~3× at 3 blocks, 16× at 25, 46× at 100 (Finding 8). Highest value-per-line change available — it removes the per-block scaling entirely. Read component lists from Learning Core Could be worthwhile, but Meilisearch is fast enough if we batch it, and I'm still opposed for the reasons already mentioned Fully async + frontend loading states Worth doing in any case Upgrade Meilisearch 4.3× from 1.12; most of it already in Tutor 1.36. In-place migration takes 2.5s. First thing to check on any affected instance. (new) Cap the poll backoff at ~250ms ~20% of remaining latency, one line. (new) Batch the create path too add_10_newcosts the same asadd_1_new; creating blocks is as batchable as publishing them.(new) Merge the two delete round trips ~2× on block deletion — from halving the awaits, not from the filter form. (not) Trim INDEX_FILTERABLE_ATTRIBUTESNo gain. Cost is flat from 5 to 20 attributes (Finding 10). Do not broaden INDEX_SEARCHABLE_ATTRIBUTES— that list is saving ~1.5s/doc.(new) Treat "Publish All" as the real case 100 blocks = 33s sequential vs 0.7s batched. Any library big enough for the button exceeds the timeout today. (not) Rely on Meilisearch auto-batching Queued tasks only partially merge (2 batches out of 3 tasks). Explicit batching is 1.8× better. Caveats
- Synthetic corpus. Shapes and field cardinalities are modelled on the real index, but this is not real course content. Absolute numbers will differ.
- I only partially reproduced the reported 5-7s per document. For updates I measured 1.22s at 500k on 1.12.8, well short of it. But creates at the same version and size are 4.13s, and heavily-tagged documents are ~4× slower to update than typical ones (Finding 10) — so an instance creating blocks, or working with richly tagged content, on the older 1.8 could plausibly land in the reported range. The original report was on Meilisearch 1.8, which I did not test, with real data on different hardware. The ratios between versions, batch sizes and configurations should transfer better than the absolute times.
- Container-limited hardware (4 CPU / 8 GB, Apple Silicon). The original reporter noted that 30Gi/6CPU did not change their results, which is consistent with this being algorithmic rather than resource-bound.
- p50 of 20 reps, 3 warm-up reps discarded, one Meilisearch instance at a time. Run-to-run variance on the same configuration is roughly ±25%, so ratios within a run are more trustworthy than absolute times across runs.
update_*numbers assume a small edit. A full content rewrite is 3.8× more expensive (Finding 7). The true publish cost sits between the two.- Not tested: Meilisearch 1.8, and the real Studio stack end-to-end (this measures Meilisearch alone, not the Django/Celery/event overhead on top).
- The concurrent-load baseline (writers=0) ran first and was cold; treat its search figures as an artefact rather than a comparison point.
- Index config transcribed from
There is a new discussion about updating the Studio search backend to support Typesense, so I just want to point out that if anyone is going to take on the work of improving batching, it should be coordinated with any decision taken in that other ADR, to avoid wasting work, or perhaps to simplify any new abstraction layer we might build.
It's also worth asking if Typesense could eliminate these "slow index speed" concerns without a major architectural change.
CC @blarghmatey
We hit the same class of problem on mitxonline, so here are our numbers. Three things that may be useful: an answer to @ormsbee's question about jumping to 1.53.1, a same-machine measurement of what that upgrade is worth, and — since @bradenmacdonald raised it — a measurement of Typesense on the same box.
Yes, 1.53.1 works. We run it in production.
We are on Meilisearch v1.53.1, with a single-replica
studio_contentindex measured on 2026-08-28 at 1,319,937 documents / 11.91 GiB. So the compatibility question has a practical answer: no blockers, but two operational details are worth knowing before anyone upgrades a populated volume.MEILI_UPGRADE_DB=trueis required. Meilisearch refuses to start when the on-disk database was written by a different engine version. It migrates in place on startup and is a no-op once versions match. It is one-way — there is no downgrade path, so rolling the image tag back means restoring the volume from a snapshot. Take one first.- Give the startup probe far more room. The migration runs synchronously before the HTTP server binds. The Helm chart's default startup budget is 60s, which can kill the pod mid-migration and leave the database partially converted. We use
periodSeconds: 5, failureThreshold: 180— fifteen minutes.
What the upgrade is worth, measured on one machine
@dwong2708's 1.8 → 1.12.8 comparison showed ~25-30%. I ran 1.12.8 against 1.53.1 on the same host, same document shapes, same index configuration and same index size, so the version effect is isolated:
Meilisearch single document, 500,000-doc index 1.12.8 1499 ms 1.53.1 1142 ms That is ~24% (1.3x) — the same order as the 1.8 → 1.12.8 step, not a step change. Worth having, and safe, but it does not solve this: indexing cost still scales with index size, which matches where @dwong2708 and @bradenmacdonald already landed. Anyone hoping 1.53.1 is the fix should not.
Method: index configured exactly as
index_config.pydoes it (distinct attribute, 20 filterable, 15 searchable, 4 sortable attributes,sort-first ranking rules), seeded with documents shaped likesearchable_doc_for_library_block()output, then create → delete → create a single document and read the task's reportedduration. Median of 7.Typesense on the same box
@bradenmacdonald asked whether Typesense could remove these concerns without a major architectural change. I did not want to answer that from the docs, so I measured it identically — same host, same documents, same index size, an equivalent collection schema (same fields facetable, sortable and searchable):
index size Meilisearch 1.53.1 Typesense 30.2 ratio 50,000 727 ms 4 ms 183x 150,000 778 ms 3 ms 244x 300,000 1044 ms 6 ms 166x 500,000 1142 ms 6 ms 189x batch of 100 @ 500,000 1454 ms 16 ms 89x The shape matters more than the ratio. Meilisearch's per-document cost rises with index size (727 ms → 1142 ms from 50K to 500K); Typesense stays flat at 3-6 ms across the whole range. That speaks directly to @ormsbee's question of whether this becomes 20s when the index doubles — for one engine it trends that way and for the other it does not.
The comparison is like-for-like in the way that matters here. The Meilisearch figure is the queued task's duration, which is what
_wait_for_meili_taskactually blocks on. The Typesense figure is a synchronous write, and I verified the document is searchable immediately after the call returns — a 7ms write followed by a search that found it 4ms later, with no waiting. Typesense has no task queue, so_wait_for_meili_taskand everything built on it simply has no counterpart.Please treat the absolute numbers as machine-specific. My Meilisearch figures are much lower than those measured earlier in this thread, and I am not suggesting those measurements are wrong — different hardware, different document sizes. Only the ratios within my own runs are meaningful.
A different lever: index fewer documents
One angle I have not seen raised here. When we looked at what was actually in our 1.32M-document index, 99.93% of it was course XBlocks, not Libraries V2 — only 966 documents carried
Fields.created, whichdocuments.pysets exclusively in the library document builders.content/searchis gated solely onMEILISEARCH_ENABLEDwith no way to narrow what it indexes, so every course XBlock create/update/delete plusCOURSE_IMPORT_COMPLETEDandCOURSE_RERUN_COMPLETEDwrites intostudio_content. Our index tripled over two days in August, which we traced to a batch of course reruns and imports — Loki showed no settings re-index and noreindex_studiorebuild in that window, so it was in-place document writes.Since indexing cost scales with index size, and much of that content is there without being needed for library authoring, reducing what goes in is a lever alongside batching and async. #39041 proposes a
MEILISEARCH_COURSE_INDEXINGsetting for exactly this, with alibrary_downstream_onlymode that keeps the course-libraries Review tab working. It is not a substitute for the fixes discussed here — it narrows the problem rather than solving it, and it gives up the Studio course search modal — but for an operator in this position it is available now and it is a configuration change.On coordination
Agreed with @bradenmacdonald's point about coordinating with #39049. To be clear about what that proposal now is: after his review I have dropped the
content/search-private framing and said I would support taking it to the broader openedx/edx-search#245 re-architecture instead. That is my position rather than a settled decision — the venue is still open. Whichever way that lands, the batching work discussed here is worth doing on its own terms — it is the largest single win in the measurements above and it is independent of which engine is behind the interface.Happy to share the benchmark harness; it is dependency-free stdlib Python against two Docker containers.
- added a commit that references this issue
on Sep 1, 2026 In terms of a short-term fix for this, one idea came out of our discussion at the Core Arch Working Group today: currently the
studio_contentindex holds both course + library content. The slowness on the library side is typically caused by the huge volume of course content, which is why @blarghmatey has even proposed a fix of just disabling indexing of course content in Studio, at the cost of Studio course search not working at all.However, we don't actually need course and library content to be in the same index. If we split them into separate indexes, we'll solve most of this problem for now: the library search will be relatively fast because its index isn't so large, and the course index will continue to work just fine because its design doesn't require synchronous updates. I think that implementing the split on the backend would be relatively easy, but on the frontend a bit more work is required to preserve the "search all libraries and courses" use case using a multi-index search request.
This wouldn't solve the long-term problem for any instances that have huge library collections, so all the other proposed fixes are still worth considering. But it could be among the simplest things we can do now and backport to Verawood to solve most of the scaling pain?
@bradenmacdonald the index split is up for review:
- Backend: feat: split Studio search into separate course and library indexes #39103
- Frontend: feat: query separate course and library search indexes frontend-app-authoring#3248
Course blocks stay in
studio_content, and Libraries V2 blocks, containers and collections move to a newstudio_library_content. An existing install upgrades withcms migratethencms reindex_studio --libraries-only, which rebuilds only the library index and deletes the leftover library documents fromstudio_contentby filter. Courses are not reindexed.One correction to the frontend estimate: no multi-index search turned out to be needed. The Studio search modal has filtered to
type = "course_block"since openedx/frontend-app-authoring#1148, and nothing else searches courses and libraries together, so each surface just picks its index.Beyond the unit tests, which mock the client, I ran the Meilisearch calls this relies on against a real Meilisearch 1.36: the library temp-index swap, the delete-by-filter cleanup, and one tenant token covering both indexes. All behaved as expected. It has not yet been run on a full Studio deployment.
release/verawoodbackport branches for both repos are ready and I'll open them once the master PRs have had a look.Reacted by Braden MacDonald and Daniel Wong@blarghmatey Awesome! I'll take a look soon.
- added a commit that references this issue
on Sep 25, 2026 As a short-term fix, we have #39122 now merged for Verawood (and on master).
As a long-term fix, we will be implementing openedx/openedx-search#1 and giving operators the option to use Typesense, as well as refactoring the indexing code.
I believe that should be enough to close this issue, but feel free to re-open if you think more is necessary.
Problem
The library commit endpoint
/api/libraries/v2/<lib_key>/commit/returns a 500 error when publishing 3+ blocks on instances with large Meilisearch indexes (500K+ documents). The publish itself succeeds (data is committed in Learning Core), but the post-publish search index update exceeds the hardcoded 15-second timeout.This leaves the library MFE showing blocks in an intermediate state — marked as unpublished with no publish option visible — because the frontend relies on Meilisearch (/multi-search) for the component listing and the search index was never updated from the frontend's perspective.
Root Cause
wait_for_post_publish_eventsin tasks.py dispatches a Celery task and blocks with:The Celery task fires LIBRARY_BLOCK_PUBLISHED for each block, which triggers library_block_published_handler -> upsert_library_block_index_doc.apply(). The .apply() call runs synchronously and includes _wait_for_meili_task() — a polling loop that blocks until Meilisearch confirms indexing.
On a 500K-document index, each document takes 5-7 seconds to index (inherent to Meilisearch's inverted index rebuild for large indexes). This is not resource-bound — we tested with 30Gi RAM / 6 CPU and observed the same 7s per document. Three blocks = 21 seconds, exceeding the 15-second timeout.
Proposed Short-Term Fix
Make the timeout configurable: Replace the hardcoded timeout=15 in dispatch_and_wait with
getattr(settings, 'CONTENT_LIBRARIES_PUBLISH_TIMEOUT', 15)so operators can tune it for their index size without patching.Batch document submissions: In send_events_after_publish (or send_change_events_for_modified_entities on master), collect all documents from the publish operation and submit them to Meilisearch in a single update_documents call instead of one document at a time. Meilisearch processes a batch of N documents as a single indexing operation — the rebuild time is roughly the same whether it's 1 or 10 documents. This would reduce a 3-block commit from 3x7s = 21s down to ~7s total.
Proposed Long-Term Fix
There are two architectural options to eliminate this problem at the root:
Option A: Decouple the library component listing from Meilisearch
Currently, the library authoring MFE retrieves the list of components directly from Meilisearch (/multi-search). This creates a hard dependency: after any write operation, the API must wait for Meilisearch to be consistent before responding, otherwise the frontend shows stale data.
The fix: use Meilisearch only for search functionality (full-text search, filtering, facets). For the library component listing (the default view showing all components), fetch directly from the Learning Core relational database via a dedicated REST endpoint.
With this approach:
Option B: Fully async indexing with per-component loading state in the frontend
Instead of blocking the API response until all indexing is complete, make publish a fire-and-forget operation:
This moves the responsibility of handling eventual consistency to the frontend, which is better positioned to give users real-time feedback on the progress of bulk operations. The backend no longer needs to block a synchronous HTTP request waiting for an external service (Meilisearch) to finish.
Environment
edx-platform release/ulmo.3 (also affects master)
Meilisearch v1.8 with ~500K documents in tutor_studio_content index
Tutor 21.x k8s deployment
Steps to Reproduce
Evidence: