Hi Mike, love your project, used it ages ago and now playing with it again. On the v2.1 branch.
Quick summary:
This looks like #1446, but I have some ideas here that weren't discussed there.
Background:
I was messing around with an experimental compiler proof assistant that compiled to metaprogramming to Lua with the idea that it could speed up the very long tactics time with the tracing JIT. Mostly, prompting agents at this point. This was causing some rather weird trace patterns, probably, and there was a lot of flushing of traces causing a slowdown.
I got this minimal commandline reproducer
for i in $(seq 20); do
luajit -jv - 2>&1 <<'EOF' | grep -c 'failed to allocate'
local f = {}
for i = 1, 2000 do f[i] = load("local a = 0 for j = 1, 100 do a = a + j * " .. i .. " end") end
for r = 1, 30 do for i = 1, 2000 do f[i]() end end
EOF
done
I get 0-125 failures on Arm64 MacOS 26, none on (dockerized, arm64) Linux. With the below patch I get 0 failures on either.
Issue
What I gleaned (out of my depth here):
- The 64MiB memory region directly above the executable seems to be special in some way and macos is trying to allocate into it first
- When we flush traces because we are not able to allocate our traces in this region, we end up getting placed back into this region, causing degenerate repeated flushes and throwing away of optimized code
Possible solution
So opus 5.5 and astra argued for a while, with me trying to follow along, and settled on something that seemed pretty minimal to my naive eyes. I don't fully understand the repercussions of allocating the trace code in the out-of-range area the OS returns, though
src/lj_mcode.c | 8 ++++++--
1 file changed, 6 insertions(+), 2 deletions(-)
@@ -294,7 +294,7 @@ static void *mcode_alloc(jit_State *J, size_t sz)
- uintptr_t hint;
+ uintptr_t hint, bad = 0;
@@ -308,8 +308,10 @@
if (mcode_inrange(J, (uintptr_t)p, sz))
return p; /* Success. */
- else if (p)
+ else if (p) {
+ bad = (uintptr_t)p; /* The OS had space here. */
mcode_free(p, sz); /* Free badly placed area. */
+ }
@@ -327,6 +329,8 @@ fail:
+ } else if (bad) { /* Move the range to where the OS had space. */
+ mcode_setrange(J, bad + (sz >> 1)); /* Used after the flush. */
} else {
J->mcmax = 0; /* Switch to a new range after the flush. */
}
Hopefully something helpful
Thanks again for LuaJIT!
Adam
Hi Mike, love your project, used it ages ago and now playing with it again. On the v2.1 branch.
Quick summary:
This looks like #1446, but I have some ideas here that weren't discussed there.
Background:
I was messing around with an experimental compiler proof assistant that compiled to metaprogramming to Lua with the idea that it could speed up the very long tactics time with the tracing JIT. Mostly, prompting agents at this point. This was causing some rather weird trace patterns, probably, and there was a lot of flushing of traces causing a slowdown.
I got this minimal commandline reproducer
I get 0-125 failures on Arm64 MacOS 26, none on (dockerized, arm64) Linux. With the below patch I get 0 failures on either.
Issue
What I gleaned (out of my depth here):
Possible solution
So opus 5.5 and astra argued for a while, with me trying to follow along, and settled on something that seemed pretty minimal to my naive eyes. I don't fully understand the repercussions of allocating the trace code in the out-of-range area the OS returns, though
Hopefully something helpful
Thanks again for LuaJIT!
Adam