Conversation
The default of N_PROC follows the machine; -DN_PROC still overrides it.
OSX target when generating on another host, symbol lookups by bisection (pandas is no longer used), one preamble parse per process, smaller collections chunks.
Same recipe as Windows: the OSX headers come from the SDK the job already downloads, the symbols from the osx-64 occt package (the virtual __osx package has to be declared on a Linux host). The SDK is the sysroot, so that the host's glibc stays out of the include path; the include layout is chosen by PLATFORM, which the native macOS job still sets from the host. The occt version of the symbol dumps is read from environment.devenv.yml, the one place it is pinned; the Windows dump had stayed at 8.0.0, and members new in 8.0.1 were missing from the Windows bindings.
The generator walks sets in many places; with hash randomisation every run orders enums, members and collection chunks differently, and one V3d member came and went. Pinning the seed at the command that runs bindgen covers CI and local builds alike.
The five collections chunks of 100 were 5 GB translation units each and met at the end of the build; at 25 the largest unit is 3.2 GB and four jobs fit the runners. The Mac step had a space after its line continuation, which ended the cmake command before the build type and the SDK: OSX was compiled without optimisation.
sccache with the GitHub Actions cache backend. The generation is reproducible, so the translation units a commit does not touch are hits. SCCACHE_GHA_VERSION namespaces the keys: bumping it starts from an empty cache.
|
Thanks, but this has too much stuff at once. To begin with, do you know what is the impact of PCH? |
The most visible impact is roughly 2h back on every CI run ;) As for the PCH itself, it is on the pywrap side (CadQuery/pywrap#66). This PR only moves the submodule. Every header re-parsed the same preamble: parsing_header, the platform header, Standard_Handle.hxx and its includes. Risk-wise: the PCH is built per process in a temporary directory, keyed on the arguments and the preamble text, so nothing stale can be picked up between runs. The commits are the smallest set I could get, one idea each and in dependency order, to keep each step reviewable on its own. @adam-urbanczyk, would you rather have them as separate PRs? |
|
No need for separate PRs, I'm just curious how much gain comes from the PCH trick and how much from effectively increasing the N_PROC to 4. |
Roughly half each, on the Linux generation (4-core GitHub CI runner): N_PROC 2 -> 4: 49 -> 27 min. Not 2x, because the modules differ a lot in size and the transform step is single-threaded. The difference is in kind: N_PROC spreads the same work over the cores, the PCH removes about two thirds of the parsing work. |
While working on #226 I waited four hours for every CI run of the bindings, most of it on steps that repeat the same work. The same jobs now take about 1h50 from scratch and 37 minutes when the object cache hits, on the same free runners. Six commits, one idea each:
N_PROCdefaults to the machine's cores instead of 2.-DN_PROCstill overrides.PYTHONHASHSEED=0for bindgen: the generation is reproducible. Two runs give byte-identical sources, so builds can be diffed.SCCACHE_GHA_VERSIONto start empty.Also fixed:
Compile (Mac).V3d_ImageDumpOptionsmember that came and went.argslist ofparse_tugrew with every call.Depends on CadQuery/pywrap#66.