local_logs added XCAT_HTTPD_ACCESS_LOG to the list of candidate logs and then
read the system paths as well. On a build host that runs apache, the check
counted the host's own /var/log/apache2/access.log beside the file the test
pointed at, so a test could not control what it measured. The empty-log case
read 64 lines it did not write.
An explicit XCAT_HTTPD_ACCESS_LOG now replaces the search instead of extending
it. Production behaviour is unchanged: nothing sets that variable there.
Without this change the bats case for an empty management-node log fails on a
host that serves apache, and passes on one that does not.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
(cherry picked from commit bcb4a77f1a)
check_provisioning_source.sh --baseline wrote one message for two conditions:
"no access log to baseline on ${MN_BASE:+$SN}${MN_BASE:-$MN}". When the management
node has a log and the service node does not, the second expansion returns the
management node's baseline spec, not a host name. On xcat42 the message read "no
access log to baseline on nosuchnode-xyz/var/log/httpd/access_log:881". That is the
message an operator reads when xdsh cannot reach the service node, so it has to name
the host that has no log.
Each condition now has its own test and its own message.
check_provisioning_source.bats covers the service-node condition. The case also
asserts the management node's log path is absent from the message, because the node
name alone matches the old text as a substring.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
(cherry picked from commit 89ac4558a0)
On the Ubuntu 24.04 hierarchy cell the check reported the right counts and then
failed: "xcat22-sn served xcat22-cn 277 boot-payload request(s)", "xcat22-mn-dhilst
served xcat22-cn 0", and then "no httpd access log with new entries could be read on
xcat22-mn-dhilst". The guard required the management node log to gain lines after the
baseline. It does not: the management node provisions the service node in
SN_setup_case, before the baseline the hierarchy case takes, and it is then idle. Its
silence is the result the check exists to find.
The guard now asks that the log hold at least one line, which is what makes a count of
0 mean something, and no longer asks for activity in the measured window.
The path filter also missed a doubled leading slash. The compute node fetches the root
image with wget as //install/netboot/<os>/<arch>/compute/rootimg.cpio.gz, which the
service node logged and the filter did not count.
check_provisioning_source.bats covers both. The four cases added in the commit before
this one fail without this change.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
check_provisioning_source.bats held a case that required the check to FAIL when the
management node logs nothing after the baseline. That is the hierarchical result, not
a broken log: the management node provisions the service node before the baseline and
then serves the compute node nothing. The suite passed 14 of 14 while the check could
not pass on a real cluster.
Two more cases cover the path filter. The compute node fetches the root image with
wget from a URL that carries a double slash, so the path reaches the log as
//install/netboot/..., and the filter matched neither direction: the request counted
as no boot payload on the service node, and a flat provision that served it went
unreported.
All four cases fail against the current script.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
check_provisioning_source.sh counted every request the compute node had ever
made, for any path. Both directions were wrong. A flat run left
management-node requests for the same address, so a later hierarchical run
read them as its own and failed. A service-node entry from an earlier run, or
a 404 for /favicon.ico, satisfied the positive check without any boot payload
being served.
The script now takes a baseline. --baseline records how many lines each access
log holds on the management node and on the service node, before provisioning,
and the check counts only lines after that point. It counts only requests
under /install or /tftpboot, which are the trees httpd serves a boot payload
from. A log shorter than its baseline was rotated, so it is read from its
first line. Without a baseline the check refuses to answer instead of reading
the whole log. The three hierarchy cases call --baseline before nodeset.
check_provisioning_source.bats covers both regressions: a management-node
request before the baseline no longer fails the run, a service-node request
before the baseline no longer satisfies it, and a request that carries no boot
payload does neither. Removing the baseline comparison fails the first two.
Removing the path filter fails the other two.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The comments added with the Ubuntu fixes carried the bug report and its
consequences: what a compute node stops at, what named answers to every query
once its working directory is unwritable, and that a failed export does not fail
the service node. The comment rules keep the invariant in the source and put the
symptom and the chain in the commit.
Each is cut to the fact the code cannot show: that the resolv.conf a
systemd-resolved host publishes holds a stub pointing back at this named, that
named must be able to write the directory it drops privileges into, that apt
refuses an unsigned repository, and that re-exporting a mount needs an fsid.
service.subiquity.tmpl keeps its header, which is identical to the sibling
compute template and should stay that way.
No executable line changes.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The GitHub check of master installs the xcat-dep packages of the latest
channel, which serves the stable release. A master change that needs a
new xcat-dep package then fails the check until that package is in a
stable release. The check now reads the devel channel, which carries
the xcat-dep packages of the next release. The 2.19 branch keeps the
latest channel.
The release information page lists every release up to 2.18.0. 2.18.2,
2.19.0 and 2.19.1 are published and absent from it, so a reader cannot
tell from the documentation which releases exist, when each one shipped,
or where its notes are.
docs/source/overview/_files/2.19.x.csv is new and holds the 2.19.0 and
2.19.1 rows, xcat2_release.rst gains its section above 2.18.x, and
2.18.x.csv gains the 2.18.2 row. Each date is the date of that release
on GitHub. 2.18.1 has no GitHub release of its own, so it has no row.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The release checklist said that the published .repo files point at
latest. latest is a link into the newest series, so a file that names
it gives the next series to users of this one as soon as that series
ships. Each file under repos/yum/X.Y/ now names repos/yum/X.Y/, and the
installation check uses the files as published.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Read the Docs takes the version in the title of each page from release
in docs/source/conf.py, which was last set for 2.17.0. Every build
since, 2.18.x and 2.19.0 included, is titled "xCAT 2.17.0
documentation". Set it to 2.20.0, the Version of master.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
genesis_ubuntu_build_root.t cut REQUIRED_PACKAGES out of
builddeb-genesis-base with a regex and evaluated it through bash -c. It
then matched the apt-cache fallback calls as text. Reformatting the
script broke it, and a wrong choice between renamed packages passed it.
genesis_payload_verification.t ran the verifier script and parsed its
stderr.
genesis_ubuntu_build_root.t now calls required_packages() with a chosen
set of carried packages and asserts the exact result: the amd64 extras,
the name picked for each renamed package, the order apt is asked,
tzdata-legacy, and the error for a release that carries neither name.
It reads the mandatory commands from XCAT::GenesisPayload, the code the
build uses. genesis_payload_verification.t calls the XCAT::GenesisPayload
functions with chosen payload trees and asserts the exact missing paths
and results.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
builddeb-genesis-base built its package list inline, between apt-get
update and apt-get install. A test could reach the list only by cutting
the REQUIRED_PACKAGES assignment out of the script and evaluating it.
It could not reach the choice between renamed packages (bind9-dnsutils
or dnsutils, util-linux-extra or util-linux) at all. The payload check
in verify-genesis-payload was bash, so a test could only run the script
and read its stderr.
XCAT::GenesisBuildRoot::required_packages() now returns the list for a
dpkg architecture. A code ref says which packages the release carries;
the default asks apt-cache. builddeb-genesis-base calls it at the same
point in the build and installs the same packages.
XCAT::GenesisPayload holds the mandatory-command list and the payload
check. verify-genesis-payload runs its main() and keeps the same
arguments, exit codes and messages. buildrpms.pl stages the module
beside the script. Both modules use core Perl only, because the build
root has perl-base and nothing more.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The error-command test ended with
[ ! -e /run/testnode-logs.tar ]
meant to show the archive had not gone to the default path. It cannot show that.
The file may exist for reasons that have nothing to do with this test, in which case
the assertion fails while nothing is wrong; and its absence would be equally true if
the override had never worked at all. It answers a question about the host, not
about the run.
The positive assertion above it already carries the proof: the tar shadow writes a
marker, and the test greps for that marker in the path it passed. The output being
there is what shows the redirection went there.
Removing it changes nothing about what the test catches. With the override taken out
of the template, so the archive path is hard-coded again, the remaining assertion
still fails.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Brings in #7866, which removes --database from the createrepo call. Without it the
branch cannot index a repository on the shared build tree: createrepo_c tries to
write primary.sqlite there and the NFS re-export answers
Cannot open .repodata/primary.sqlite: Can not create db_info table: disk I/O error
which failed every EL build in xcat-ci #44 while the Ubuntu builds, which do not use
createrepo, passed.
Killing the command is not killing the build. sh() ran the command through /bin/sh,
and cancellation signalled that shell alone -- but dpkg-buildpackage starts workers
of its own, and those survive their shell. The lock was then released while they
were still writing debian/changelog and debian/control, which is the state the lock
exists to prevent: the next build takes the checkout and the two rewrite it
together.
The command now runs in its own process group, so cancellation can take all of it.
Both sides call setpgid, so neither depends on which runs first, and INT and TERM
are blocked across the fork so cancellation cannot land in the window before the
group exists.
Cancellation escalates from the caught signal to KILL, and then CHECKS: a shell that
has exited is not a build that has stopped, so it waits for the whole group to
disappear rather than for the leader to be reaped. If the group is still there after
that, the locks are RETAINED and the process exits non-zero. Releasing a lock while
a worker may still be writing is worse than leaving a lock behind for a person to
clear -- the first corrupts a build, the second stops one.
cancel_build ignores INT and TERM while it runs, so a second Ctrl-C cannot interrupt
the cleanup half way and release the lock early.
sh() also reports a signalled command as 128+signal instead of 0. $? >> 8 is zero
for a child killed by a signal, so a build stopped mid-way looked to its caller like
one that had succeeded.
Two cases added to builddebs_lock_cancellation.t: a build whose worker is a
grandchild, and a command killed by a signal. Verified by signalling the pid instead
of the group, which leaves the worker running and turns the first red.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
mknb wrote BOOTIF=01-${netX/machyp} into the BIOS Genesis script.
machyp is an xNBA setting, and stock iPXE expands it to an empty string.
Without the boot MAC, the legacy Genesis stops with "Unable to find boot
device" after at least 10 minutes. The OpenEmbedded Genesis fails at once
with BOOT_INTERFACE_NOT_FOUND.
Use ${netX/mac:hexhyp}, as the UEFI Genesis script and xnba.pm do. xNBA
expands both settings to the same value, so boot with xNBA is unchanged.
The install monitor holds the later connections for a node whose handler
is still running, and forks the next one when it reaps that handler.
SIGCHLD is what brings the parent back to look: it interrupts the accept.
A handler can exit after the parent checks its queue and before the
accept begins. The signal is then handled where there is no accept to
interrupt, and the parent blocks in accept with a connection already
queued and ready to run. That connection waits until some other node
calls in. A single node retrying on its own waits until it times out.
do_installm_service now waits through wait_for_installm_connection, which
selects on the listening socket. The wait is bounded by
$installm_wakeup_seconds while connections are queued, so the parent
looks at its queue again instead of waiting for another client. An idle
monitor with an empty queue still waits without a bound, because a
handler that exits then leaves nothing to do.
The new case asserts the wait ends on its own bound with nothing to
accept, and ends at once when a connection is already there. It fails
when the bound is ignored.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The install monitor answers a node's requests in the order they arrived.
Two ways to lose that order had no test.
The first is a queued client that gives up. The parent holds the later
connections for a busy node unread, so a client that closes its socket
must not let the request behind it overtake the request that is running.
The new case runs one request, queues two more, drops the middle client,
and asserts the last request starts after the running one ends. It fails
when the per-node queue is removed.
The second is a fork that fails. The monitor then answers the node
itself, in line. The new case makes one fork fail and asserts the request
is answered and that the node's next request starts only after it. It
fails when the fallback drops the connection instead.
The file also asserts the lifted service holds no literal /var/run path.
Every access to the pid file goes through $installm_pidfile, so pointing
that variable at the scratch tree redirects all of them, and a path
written out again inside the routine would reach the host file.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xcatd_install_monitor_concurrency.t failed about one run in five. The
assertion "the node is released only after the destiny advance finished"
compares the time the test read "done" from the socket against the time
the plugin stand-in recorded when it finished. The stand-in wrote that
time with %.3f, which rounds up, so a recorded time can be later than the
moment it was taken. The test then failed on the rounding and not on the
order:
'1790220468.11892' >= '1790220468.119'
The events file now holds whole microseconds from gettimeofday, and the
comparisons read the same clock. gettimeofday rounds nothing.
Each case also writes an events file of its own. A handler forked by one
case outlives the monitor that forked it, so it could append to the case
that runs next.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The build lock is released in DESTROY, and perl does not run DESTROY when a signal
ends the process. A build stopped with SIGTERM or SIGINT therefore left its lock
directory behind, and the next build of that checkout died on
FATAL: another build of <path> already holds <dir> (held by [pid=NNNN])
naming a pid that had already exited. Nothing clears it but a person. One such
directory blocked an openSUSE target across three consecutive CI runs before anyone
looked at what the lock actually said.
buildrpms.pl has released its lock on cancellation for some time, through an END
block and an abort handler. This is the Debian builder catching up.
The order matters, and is the reason this is not simply an END block. The command in
flight is stopped BEFORE the lock is released: handing the checkout to a second build
while dpkg-buildpackage is still rewriting debian/changelog and debian/control in it
is worse than holding the lock a moment longer. The wait for that command is bounded,
so a subprocess that ignores the signal cannot hold the lock for ever either.
sh() now forks and execs rather than calling system(), because system() gives no pid
and a handler cannot stop what it cannot name. The child _exits rather than exits, so
it never runs the parent's END block and releases a lock the parent still holds.
The handler is installed by XCAT::BuildUtils::install_build_cancellation rather than
written inline in the builder, so a test can use the same wiring the builder uses. A
test that installs an equivalent handler of its own proves the helper works while
saying nothing about whether anything calls it -- the first version of this test did
exactly that, and passed with the wiring removed.
Release is idempotent: a signal handler and then DESTROY both reach it, and the
second must not remove a directory a LATER build has since taken.
builddebs_lock_cancellation.t terminates the holder, then takes the lock again, and
checks no build subprocess was orphaned. Verified by removing the wiring: assertions
9 through 12 fail, naming the leaked lock, the refused build and the stray process.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>