aurupdater

Cross-package tooling notes

2026-09-11: remote-board build times now label MAKEFLAGS/serial, not a bare “remote”

repo_record_build_time.sh’s “Build times” block in each package’s memory note already labeled local chroot builds with their actual parallelism (MAKEFLAGS=-jN or serial, derived from repo_build.sh’s own MAKEFLAGS env var), but remote ARM-board builds (eurobuild4/5/14) just got a fixed remote – no indication of whether that board built serially or with some -jN. User asked for parity with the local labels.

Fix: scripts/remote_build.sh (the agent script deployed to and run on the board) now resolves its own effective parallelism the same way makepkg itself would – MAKEFLAGS from the environment if exported (it currently isn’t, in the non-interactive ssh session this runs under), else parsed out of the board’s own /etc/makepkg.conf (last uncommented MAKEFLAGS= line, quotes stripped) – and echoes it as REMOTE_BUILD_PARALLELISM: MAKEFLAGS=-jN (or ...: serial) to stdout before starting the build. That line lands in the ssh output repo_build_remote.sh returns, which repo_build.sh already captures per-arch into ${_remote_tmpdir}/${arch}.log; on a successful build, repo_build.sh now greps that line back out and passes "remote, ${label}" (e.g. remote, MAKEFLAGS=-j2, remote, serial) to repo_record_build_time.sh instead of the bare remote. Falls back to plain remote if the line isn’t found (e.g. an older/unpatched remote_build.sh still cached on some board before its next self-deploy refresh).

No pkgdir files involved – the label rides through stdout/the existing log capture, not a synced-back sidecar file, so no new stray-file cleanup concern (cf. feedback_check_stale_artifacts_before_build.md).

2026-09-11: use MAKEFLAGS=-j6 for local euronuc builds going forward

Local chroot builds (x86_64/i686/pentium4/i486 on euronuc) have always been single-threaded by default (/etc/makepkg.conf’s MAKEFLAGS="-j1"), since repo_build.sh’s MAKEFLAGS passthrough (added 2026-09-04) is an opt-in the caller has to set, not a default. Came up when asked whether local archs could build in parallel with each other – they can’t, safely, as-is: every local arch’s chroot binds the same host pkgdir as both /startdir and /srcdest (no per-arch isolation the way the remote-board path already has, where each board gets its own rsynced copy), so two local archs building at once would race on the same src/ extraction. Cross-arch parallelism would need real staging work (isolate each local arch into its own copy, mirroring the remote pattern) – not done.

The safe, immediate win instead: set MAKEFLAGS=-j6 when invoking repo_release.sh/repo_build_staged.sh for local builds, so each (still-sequential) arch build at least uses multiple cores to compile. -j6 (not -j8, euronuc’s full core count) leaves headroom for the separate Archlinux32 builder daemon that also runs on this same host – same reasoning repo_build.sh’s own MAKEFLAGS mechanism was built around in the first place (never touches the real /etc/makepkg.conf, see the 2026-09-04 entry above). Adopted as a standing convention going forward, not a one-off.

2026-09-11 (same day, follow-up): MAKEFLAGS=-j6 baked in as an actual script default

Per direct follow-up request, the convention above is no longer just a “remember to pass it” habit – scripts/repo_build.sh now has : "${MAKEFLAGS:=-j6}" + export MAKEFLAGS right after its REPO_BUILD_ONLY_ARCH default. Still fully overridable by an explicit caller-set MAKEFLAGS (e.g. MAKEFLAGS="-j1" for the old serial behavior). The explicit export matters: if a caller never sets MAKEFLAGS at all (now the common case), := alone would only create a local shell variable that sudo --preserve-env=MAKEFLAGS further down couldn’t see.

Verified end-to-end against a real chroot build, not just an env-passthrough test (see this file’s 2026-09-04 MAKEFLAGS entry for why that alone isn’t trustworthy): ran scripts/repo_build_staged.sh adapted/gconf with no MAKEFLAGS set, confirmed MAKEFLAGS=-j6 + the scratch MAKEPKG_CONF in the actual running makechrootpkg/sudo processes’ environment via /proc/<pid>/environ, and confirmed the auto-recorded build time picked up the new label (memory/gconf.md: x86_64 41s, MAKEFLAGS=-j6, down from a prior serial run’s 55s).

Gotcha hit while verifying: repo_build_staged.sh copies built .pkg.tar.* back into the git-tracked arch/<category>/<pkg>/ dir by design (so repo_release.sh can find them to sign/publish) – calling it directly for a build-only test, with no follow-up repo_release.sh, leaves those files sitting in the git tree; cleaned them up manually afterward. Not a bug, just something to remember when testing the build step in isolation.

2026-09-12: per-board lock to stop concurrent remote builds colliding

Prompted by a real incident this session: git-wd40 and xxdiff-git ended up both compiling on eurobuild4 at the same time (a single-core, 423MB-RAM board already fragile running just one build) – missed because checking “is the board free” via a point-in-time ps/pgrep snapshot is inherently racy, and an earlier xxdiff-git attempt had also left an orphaned, un-tracked makepkg process tree running on that same board from way earlier in the session (parented to init, no longer attached to any local repo_release.sh this session could see) – a second, independent source of the same collision risk.

Fix: remote_build_lib.sh gained remote_build_lock <host> <label> / remote_build_unlock <host> – an atomic mkdir-based lock (portable, no flock dependency on the board) at REMOTE_BUILD_ROOT/.building.lock, holding a holder file with whatever label the caller gives it. repo_build_remote.sh now claims the lock right after its reachability ping, before touching the board at all, and releases it via an EXIT trap covering every path (success, failure, this script getting killed). A held lock fails the new build fast with a clear error naming the holder, instead of silently piling on.

Deliberately fails closed, not open: a lock left behind by a crashed/killed build stays locked until someone manually clears it (ssh <host> rm -rf <lockdir>, or remote_build_unlock <host>) – auto-expiring after some timeout was considered and rejected, since a still-legitimately-running multi-hour build (e.g. git-wd40’s armv7h, 1h46m) could otherwise have its lock stolen out from under it.

Bug caught during verification, not left in: first version of remote_build_unlock used rmdir, which fails silently on a non-empty directory (the lock dir isn’t empty – it holds the holder file) while the function unconditionally returned 0 regardless. A lock survived its own unlock call until this was caught by an actual end-to-end test (acquire -> concurrent-attempt-fails -> unlock -> acquire-again-succeeds) run against a real board (eurobuild14), not just a syntax check – fixed to rm -rf. Same lesson as [[makeflags_chroot_parallelism]]: verify remote-board plumbing against the real board, a shallow check isn’t enough.

2026-09-12: remote boards’ non-interactive PATH missing Perl’s core_perl/vendor_perl/site_perl dirs

git-wd40’s armv7h build failed generating a bundled Perl module’s man page: pod2man: command not found, Error 127. Not a missing dependency – perl (a real depends, correctly installed by makepkg’s own --syncdeps) was present, and pod2man genuinely existed on the board at /usr/bin/core_perl/pod2man. Root cause: remote_build.sh runs via non-interactive ssh -n, which skips the login-shell profile scripts (/etc/profile.d/perlbin.sh on Arch) that normally append /usr/bin/{core_perl,vendor_perl,site_perl} to PATH – confirmed directly that an interactive shell’s PATH includes these, a non-interactive ssh session’s doesn’t.

Fixed in remote_build.sh itself: appends the same three dirs to PATH (guarded by [ -d ... ], mirroring perlbin.sh’s own logic) right after set -u, before anything else runs. General fix, not git-wd40-specific – any future package needing a core-Perl tool on these remote boards would have hit the identical failure.

2026-09-13: local staging lock is per-package, not per-node (by design)

Added repo_build_staged.sh’s own lock (same mkdir-atomic pattern as remote_build_lock/unlock) after a second repo_build_staged.sh invocation for git-wd40 silently wiped an in-progress one’s not-yet-copied-back build output (i486/x86_64/aarch64 lost – see this file’s 2026-09-12 remote-lock entry for the matching incident on the board side). User asked why this lock is keyed by package (/data/INSTALL/<pkgname>/.building.lock) rather than by node (euronuc as a whole), given the remote board lock is host-keyed.

Deliberately different granularity for a reason: the remote lock is per-host because a single ARM board genuinely can’t run two builds at once regardless of which package (1 core, 423MB-1.8GB RAM) – that actually is a node lock, correctly scoped. Euronuc itself (8 cores, real headroom) is not in that position: different packages’ local chroot builds running concurrently is fine and has been relied on all session (pacman-static + libarchive-static together, git-wd40 + setserial together, etc.) – a node-wide local lock would needlessly serialize that. The actual bug was narrower: two invocations of the same package sharing the same staging dir. Per-package scoping fixes exactly that without losing the cross-package parallelism this project depends on. User confirmed this reasoning after asking about it directly.

Verified end-to-end against a real package (sc): manually planted a fake lock -> confirmed a real invocation gets rejected with the holder’s label -> cleared it -> confirmed a real build proceeds and the lock is released on exit (trap fires even though checkpkg legitimately errored at the very end, since the rebuild produced byte-identical output to what’s already published – unrelated to the lock itself).

2026-09-15/16: site “Currently Building” page – persistent build-status files + a fast site-regen path

User wanted a site page showing what’s building right now and where. Chose a static-status-file + republish-on-interval design over a truly live/dynamic one, to fit the site’s existing architecture (plain HTML, rsync-published, no backend) – repo_arch_status.sh already existed for a single package’s best-effort “is this building” (ps/ssh probes, no persistent state – see its own header comment), but nothing tracked it globally/cheaply across every package at once.

Mechanism: scripts/repo_lib.sh gained repo_build_status_start <pkgname> <arch> <host> / repo_build_status_end <pkgname> <arch>, writing/removing ../build_status/<pkgname>--<arch>.status (one file per running build, content "<host>\t<start-iso-utc>\t<pid>"). repo_build.sh calls these at its three actual build points: the arch=any case, the per-arch local-chroot loop, and inside the backgrounded subshell that waits on a remote board build. The <pid> recorded is always a local (euronuc) process – repo_build.sh itself for a local chroot build, or the backgrounded subshell for a remote one – deliberately never a remote PID, so a reader can check staleness with a plain local kill -0, never needing to ssh anywhere (if the local wrapper is gone, the entry is stale regardless of what the remote board thinks it’s doing).

scripts/generate_site.sh got a new --current-builds fast path: skips the full OVERVIEW/memory-notes rebuild entirely and only re-renders site/current-builds.html from build_status/*.status, self-healing (a kill -0-dead entry gets deleted, not shown). A full (or --only) run also calls the same render function and adds the page’s nav link on index.html – so a fast-path-only run never needs to touch index.html for the link to already be there. New scripts/watch_current_builds.sh [interval-seconds] (default 120s) loops that fast path + publish_site.sh.

Two bugs caught before/while verifying against a real build:

  1. The markdown table body rows were being assembled with a leading newline before each row ("${_cb_rows}\n| ... string growth), which put a blank line between the |---| header separator and the first data row – lowdown (correctly, per CommonMark) then treats that as “table has no body rows” plus a stray trailing paragraph, not a rendering bug in lowdown. Fixed by making each row string end with its own trailing newline instead, so rows concatenate with no gap. Caught by testing against synthetic build_status/*.status entries before ever wiring this into a real build – planting a fake in-progress entry (with a genuinely-alive PID, e.g. a sleep & job) and a genuinely-stale one (a PID that can’t exist) is the way to verify this end-to-end without waiting on a real multi-hour build; also confirmed the stale one gets deleted and the empty-state fallback text renders correctly.
  2. watch_current_builds.sh’s first draft self-terminated the moment build_status/ was empty at all, including on its very first check right after launch – fine for “watch an already-running build to completion”, wrong for “launch it slightly early, before the next arch has started yet” (exactly what happened launching it for linux-lts515’s remaining i686/i486/pentium4 while x86_64 was still mid-build under the pre-instrumentation script copy, see below). Fixed by only self-terminating on empty after having seen at least one non-empty check first (_wcb_seen_any guard) – it now waits quietly, polling, for the first sign of activity, then tracks it through to done.

Editing a script file while a build already has it open, live – confirmed the WHOLE invocation keeps running the pre-edit version, not just the phase already in progress: repo_build.sh was edited (this session) while linux-lts515’s x86_64 phase was already running under it (a long sudo ... | tee pipeline mid-flight). The edit itself didn’t crash anything – the build kept progressing normally the whole time, consistent with the classic “safe to upgrade a running script” property (an already-open file descriptor keeps reading its original inode’s bytes regardless of what happens to the path afterward, as long as the edit is a write-to-temp-then-rename rather than a true in-place truncate+overwrite of the same inode). But that same property means the entire invocation – not just the phase already in-flight – kept reading from that one original open fd for its whole lifetime: pentium4/i686/i486, reached hours later after x86_64 finished, are later reads from the same fd, not a fresh open(), so they got the pre-edit bytes too. Confirmed directly: build_status/ was never touched again after this session’s own manual test cleanup, across all four archs’ real builds (multi-hour each) – no repo_build_status_start/_end call ever actually ran in this invocation. A watch_current_builds.sh launched to track the remaining archs consequently ran for 8+ hours never seeing any activity (its _wcb_seen_any guard correctly never fired, so it also correctly never self-terminated on a false “done” – but it did have to be killed by hand once the real build finished, since nothing it was waiting for was ever going to happen). Lesson: a script edited while any invocation of it is currently running will not take effect for that invocation at all, no matter how many loop iterations/phases are still ahead of it – only a fresh invocation (started after the edit) sees the new code. Don’t edit a script while something has it open if avoidable; if it can’t be avoided, expect zero effect on the already-running invocation, not partial effect.

2026-09-16: remote board lock’s own mkdir ran before the /data mount-ensure step – an unmounted board looked exactly like real contention

oksh’s ARM rebuild (all 3 boards, armv6h/armv7h/aarch64) failed immediately on every attempt with ERROR: <host> is already building something else -- (empty holder label) – three separate retries, same result every time, on all three boards simultaneously. Initially read as real external contention (checked for other Claude sessions via ListAgents – all offline; checked cron/systemd timers – nothing relevant; asked the user directly rather than force-clearing a lock that might be protecting a genuine concurrent build). User confirmed nothing should be running there.

Actual cause, found by reproducing the exact lock command by hand: ssh <host> "mkdir '/data/INSTALL/.building.lock' ..." returned mkdir: cannot create directory ...: No such file or directory – not “File exists” (which is what a real held lock would produce). mountpoint -q /data confirmed all three boards currently have /data (the NBD block device) unmounted, likely from a reboot. repo_build_remote.sh already had a /data mount-ensure step (added after an earlier, different incident on 2026-08-27 – an unmounted /data failing the rsync step) – but it ran after the board-lock acquisition (remote_build_lock, in remote_build_lib.sh), whose own mkdir is also a write under /data/INSTALL/. remote_build_lock’s error path doesn’t distinguish “mkdir failed because the lock dir already exists” (real contention) from “mkdir failed because a whole parent directory is missing” (unmounted /data) – both just fail the mkdir, and the fallback cat .../holder fails too either way (empty output either because there’s no holder file, or because the whole path doesn’t exist), so both cases print the identical misleading “already building something else – “ message. This meant any board with /data unmounted would report 100% of its remote build attempts as “locked by someone else,” never as the real “mount needed” error, and never succeeding no matter how many times it was retried.

Fix: moved the mount-ensure block in repo_build_remote.sh to run before the lock acquisition, not after – confirmed directly: a manual ssh <host> mkdir '/data/INSTALL/.building.lock' ... reproduced the exact ENOENT before the fix, then a full retry after the fix built and published all three archs cleanly on the first pass (no more false locked errors). remote_build_lock itself wasn’t changed – it’s the call-site ordering in repo_build_remote.sh that was wrong, not the locking primitive.

Lesson for next time this pattern shows up: an “already building something else” error with a visibly empty holder label is a strong signal it isn’t real contention – a real lock’s holder file is never empty, remote_build_lock always writes a non-empty label ("<pkg> (<arch>), pid <n>, <timestamp>") when it actually succeeds in claiming the lock. Don’t stop at “it’s locked, try later” for an empty label specifically – reproduce the underlying mkdir/cat by hand against the board first.

2026-09-29: LAN mirror (archlinux.lan.brgn.ch) missing extra.db – x86_64 chroot syncs can fail

The archlinuxaba-x86_64-build devtools chroot’s base core/extra/ multilib repos (distinct from archlinuxaba itself, which points straight at archlinux32.andreasbaumann.cc via arch/config/pacman.conf.d/archlinuxaba.conf) come from this machine’s own /etc/pacman.d/mirrorlist, which has exactly one active Server line: https://archlinux.lan.brgn.ch/archlinux/$repo/os/$arch (a LAN mirror). Confirmed 2026-09-29 (check_ssl_cert build) that this mirror serves core.db fine but 404s on extra.db specifically – not transient, failed identically on a retry minutes later.

Fixed (user-approved – this is a host-level config change, not project-scoped) by adding a fallback line right after the broken one: Server = https://geo.mirror.pkgbuild.com/$repo/os/$arch. Had to edit two files, not one: the host’s own /etc/pacman.d/mirrorlist and /var/lib/archbuild/archlinuxaba-x86_64/root/etc/pacman.d/mirrorlist – the chroot keeps its own separate copy from whenever it was last created/synced, not bind-mounted or live-linked to the host’s file, so a host-only edit alone would not have fixed the chroot’s own sync. Needed sudo for both (plain Edit/Write gets EACCES).

If a local x86_64 build ever fails again with error: failed retrieving file 'extra.db' ... 404 (or any other base-repo db), check this mirror first before assuming it’s a packaging problem – and remember the chroot’s mirrorlist needs the same fix applied separately, the host-level one won’t propagate on its own. Not checked whether i486/i686/pentium4 chroots have their own separate mirrorlist copies with the same gap (they didn’t hit this error the one time it was observed, so either they’re pulling extra from a different working mirror already, or just got lucky) – worth checking proactively if one of those hits an extra.db 404 too.