The comments added with the Ubuntu fixes carried the bug report and its
consequences: what a compute node stops at, what named answers to every query
once its working directory is unwritable, and that a failed export does not fail
the service node. The comment rules keep the invariant in the source and put the
symptom and the chain in the commit.
Each is cut to the fact the code cannot show: that the resolv.conf a
systemd-resolved host publishes holds a stub pointing back at this named, that
named must be able to write the directory it drops privileges into, that apt
refuses an unsigned repository, and that re-exporting a mount needs an fsid.
service.subiquity.tmpl keeps its header, which is identical to the sibling
compute template and should stay that way.
No executable line changes.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
nodepurge removes the autoinstall configuration of each node it deletes. The
cleanup loop was inline in the nodepurge sub of profilednodes.pm, which no test
can load, so the loop moves to xCAT::ProfiledNodeUtils->remove_node_config_files
with its behaviour unchanged.
nodepurge_autoinst_cleanup.t drives that routine against a scratch directory. It
fails here: the Subiquity node keeps its directory, and the preseed file and the
.pre and .post scripts are removed.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
A Windows UEFI install deferred to proxyDHCP gets boot-file-name "" in its
Kea reservation. Kea treats the empty string as not specified, so the
subnet classes still match on architecture, and xcat-uefi-x64 offers
xcat/xnba.efi. The firmware then boots xNBA instead of asking the
proxyDHCP daemon on port 4011.
kea_node_client_classes_for_nodes now also puts the MAC of a proxyDHCP
node in xcat-localboot, which every class that names a boot file already
excludes. The per-node class that sets option 60 to PXEClient names no
boot file and still matches.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
mknb wrote BOOTIF=01-${netX/machyp} into the BIOS Genesis script.
machyp is an xNBA setting, and stock iPXE expands it to an empty string.
Without the boot MAC, the legacy Genesis stops with "Unable to find boot
device" after at least 10 minutes. The OpenEmbedded Genesis fails at once
with BOOT_INTERFACE_NOT_FOUND.
Use ${netX/mac:hexhyp}, as the UEFI Genesis script and xnba.pm do. xNBA
expands both settings to the same value, so boot with xNBA is unchanged.
The install monitor holds the later connections for a node whose handler
is still running, and forks the next one when it reaps that handler.
SIGCHLD is what brings the parent back to look: it interrupts the accept.
A handler can exit after the parent checks its queue and before the
accept begins. The signal is then handled where there is no accept to
interrupt, and the parent blocks in accept with a connection already
queued and ready to run. That connection waits until some other node
calls in. A single node retrying on its own waits until it times out.
do_installm_service now waits through wait_for_installm_connection, which
selects on the listening socket. The wait is bounded by
$installm_wakeup_seconds while connections are queued, so the parent
looks at its queue again instead of waiting for another client. An idle
monitor with an empty queue still waits without a bound, because a
handler that exits then leaves nothing to do.
The new case asserts the wait ends on its own bound with nothing to
accept, and ends at once when a connection is already there. It fails
when the bound is ignored.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Merge current master into the openEuler support branch. Keep native
kernel, DHCP, locale and timezone requirements with the upstream payload
verifier and EL10 package paths. Preserve signed openEuler installation.
Update the native fixtures to check the merged payload contract.
The ordering between two requests for one node was a chain of pipes: each child
held the write end and the next child for that node read the previous one to end
of file. End of file there means the previous process is gone, not that its work
finished, so a child that died released the one behind it -- and released the
wrong one, because each child waits on its immediate predecessor rather than on
the request actually in flight. The fork-failure path closed the predecessor and
served the request in line without waiting at all.
The parent now owns the order. %installm_busy names the handler serving a node,
%installm_queue holds the connections accepted for that node meanwhile, and the
next one is forked when the handler ahead of it is reaped. A handler that dies
cannot release the one behind it, and the fork-failure path has nothing to fall
back over, because it is reached only when the node has no handler.
Three other changes the same design makes possible or necessary:
- The answer to a destiny advance now follows the advance. Every other request
is still answered before it runs, because its result does not change what
the node does next. Holding one node costs no other node anything now, and
it lets the node retry an advance whose handler died -- which the old order
could not, because "done" was already on the wire.
- The monitor drains its handlers before it exits. Without this a restart
orphans them into the systemd service cgroup, where anything still running
at TimeoutStopSec is killed after its answer was already sent.
- SIGCHLD is caught, so a handler exiting interrupts accept and the parent
comes back to look for a connection queued for that node.
The pid file path is a variable, so the test can point the lifted routine at a
scratch file instead of the one a restarting xcatd reads to tell the running
monitor to let go of the port.
xcatd_install_monitor_concurrency.t grew the cases for all of it. Four
mutations, each caught by one assertion: forking every connection at once turns
the ordering case red; answering a destiny advance before the plugin turns the
release case red; removing the drain turns the stand-down case red; and leaving
the emptied queue entry behind turns the leak case red. The full unit suite is
194 files, 5662 tests, green.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The step that fills lib/firmware in the Ubuntu netboot initrd asked modinfo
about $rootimg_dir/$module and looked a firmware name up under lib/firmware.
The copy step beside it takes a module from $customdir or $pathtofiles first,
and the kernel looks a firmware name up under updates/<kernel>, updates/,
<kernel>/ and lib/firmware. So a custom driver reached the initrd with no
firmware, even when the root image carried it, and a firmware override was left
out of the initrd altogether.
initrd_firmware_files now takes the module directories the copy step searches
and the kernel release. It resolves each module in that order before asking
modinfo, and keeps every firmware file that exists in the four directories the
kernel searches, so the override still wins on the node.
ubuntu_genimage_initrd_firmware.t covers both: its two new cases are red on the
commit before this one.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The first subiquity error command redirects into /run/#HOSTNAME#-logs.tar. A
test can shadow tar, but not the redirection: the shell opens that file
whether or not tar runs, so the test writes to the host it runs on. Unprivileged
the open fails and "exit 0" hides it; CI runs as root, where the same line
creates or truncates the file.
The path is read from XCAT_ERROR_ARCHIVE, falling back to the same default, the
way XCAT_ERROR_CONSOLE already does on the next line. An install sets neither
variable and writes where it always did.
The test points the variable inside BATS_TEST_TMPDIR, has the tar shadow emit a
marker so the output can be traced to the file it named, and asserts the default
path under /run was not touched. Reverting the template to the fixed path fails
that case.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
makedns reports "error was FORMERR" for an update that named rejected with
NOTAUTH. The FORMERR is the answer to the retry, not to the first attempt.
send_ddns_update in ddns.pm signs the packet the caller built, and signs that
same packet again on each attempt. Net::DNS::Packet::sign_tsig appends the TSIG
to the additional section, so the second attempt sends two TSIG records and
named answers FORMERR. FORMERR is neither NOTAUTH nor SERVFAIL, so the routine
stops and reports it. The NOTAUTH and SERVFAIL retry can never be accepted, on
any algorithm.
Each attempt now signs a request of its own. A packet cannot be unsigned again,
so ddns_update_request copies the prerequisite and update records into a new
Net::DNS::Update instead, and the caller keeps the unsigned original.
ddns_update_retry.t fails before this change: the second attempt carries two
TSIG records, and an update that the retry answers with NOERROR still reports
failure.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
makedns exits 1 on a management node that has Net::DNS below 1.36 and an
hmac-sha256 key, and reports "Failure encountered updating <zone> with entry
'', error was FORMERR".
update_namedconf in ddns.pm rewrites the named.conf key stanza to hmac-md5
whenever Net::DNS is below 1.36, and ddns_tsig_algorithm returns hmac-md5 for
the same reason. ddns_sign_update signs with site.dhcpomapialgorithm, which
xcatconfig sets to hmac-sha256 on EL9 and later. named matches a TSIG key by
name and by algorithm, so it answers NOTAUTH. The retry signs the same packet
a second time, and named answers FORMERR to the two signatures.
The version test protected the two-argument sign_tsig($name, $secret), which
produces an HMAC-MD5 signature only. ddns_sign_update signs every other
algorithm through a KEY RR, so the Net::DNS version no longer selects the
algorithm. This change deletes the rewrite and the version test, and signs with
the algorithm the key stanza declares. OmapiPolicy->algorithm_rr_type maps that
algorithm to its KEY RR number.
ddns_named_key_algorithm.t fails before this change: it reads the stanza as
hmac-md5 where the key was hmac-sha256.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The xcatd install monitor serves every installing node from one accept loop. It
accepts a connection, resolves the peer to a node and dispatches the request in
line, so nothing else is accepted until that request returns. On a management
node holding three connections that send nothing, a second node waited 5.7
seconds for the monitor's greeting.
do_installm_service in xCAT-server/sbin/xcatd calls plugin_command directly in
every branch. The per-request fork that once stood there is commented out with
the note that the node must be blocked, because 'nodeset next' and
'installstatus' for one node write the same chain row.
The monitor now gives each connection its own child and keeps that ordering per
node: the child for a node reads a pipe left by the previous child for the same
node, and starts when that pipe reaches end of file. Live children are capped at
64 and the rest wait in the listen backlog, the parent reaps them, and a child
that dies no longer takes the monitor with it. A per-node lock file was rejected
because it needs a new directory and gives no arrival order; letting the parent
wait for the busy node was rejected because it blocks the accept loop again.
xcatd_install_monitor_concurrency.t lifts do_installm_service out of the program
and drives it with real clients. Without this change the second node waits 5.7
seconds and the monitor does not survive a request that kills its handler.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Add native compute and service profiles, installer dispatch and partition
defaults. Use strict DNF image creation. Check stateless boot-file and
archive publication. Preserve native image customizations on reimport.
Restore missing image records during reimport.
Use ISC DHCP for openEuler release packages. Retain the EL and SUSE
dependency choices. Select Apache 2.4, the httpd service and native Perl
build generators.
Reuse NetworkManager for persistent routes, bond setup and connection
properties on openEuler. Preserve ifcfg handling where the connection still
uses it. Propagate route failures to local callers.
dhcp_ddns_policy.t and dhcp_isc_expression_grouping.t called BAIL_OUT
when a routine was missing or the plugin could not be read. prove stops
every remaining file on a bail-out, so one of them hides the results of
every test that would have run after it. die is just as loud and costs
only its own file.
Three facts in this branch were each written out in four places. That
SIGHUP reports success before Kea reads the file appears in Kea.pm twice
and in dhcp.pm again; that an installed node netboots when no boot file
reaches it appears in BootPolicy.pm and three times in dhcp.pm; the
dhcpd grouping rule appears in BootPolicy.pm and in the test header. Each
now stands once, without the symptom narration and the defect history
around it.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
ubuntu_subiquity_error_commands.t called BAIL_OUT at both places where
its extraction of the error-commands block stopped matching. prove stops
every remaining file on a bail-out, so a change to the template that
breaks the regex in this file also hides every test that would have run
after it. die is just as loud and costs only this file.
Three comments the branch added also carried the bug report: the test
header, the template comment and the Template.pm comment each traced the
failure from the error command to the provisioning timeout. Each now
states the constraint the reader cannot see in the code.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The ubuntu-26.04-ppc64el diskless node never starts its kernel. SLOF reports
W3411 and E3406 and falls through to disk. The node fetches the kernel and the
initrd in under a minute and fails six minutes later, with no network activity in
between: grub2 has the payload and cannot start it.
genimage copies the whole lib/firmware tree of the root image into the initrd.
That tree is 666 MB on Ubuntu 26.04, which takes the initrd to 719 MB. A 687 MB
initrd boots the same kernel on the same node; a 719 MB one does not.
The firmware copy now takes only the firmware that the drivers in the initrd ask
for, which modinfo reports for each module. The root image keeps its whole tree,
so the node that boots is unchanged.
ubuntu_genimage_initrd_firmware.t drives the firmware step over a root image whose
firmware tree holds files no driver asks for. It fails on the previous commit.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
genimage returns 2 with "Failed to find usr/bin/dig" for
ubuntu22.04 and ubuntu24.04 on ppc64el, so packimage writes no initrd.gz
and the diskless compute node never boots.
xCAT-server/share/xcat/netboot/ubuntu had a ppc64el package list for
20.04 and 26.04 only. imgutils::get_profile_def_filename then falls back
to compute.pkglist, which installs no dig, no cpio and no chrony.
Add the 22.04 and 24.04 ppc64el lists, and the ppc64le spelling each
release already carries. Both take the content of the 26.04 ppc64el list:
the ppc64el images build their initrd with mkinitrd, so they install
bind9-dnsutils and leave out the dracut packages the x86_64 lists take.
ubuntu_ppc64el_pkglists.t fails without these files.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
reg_linux_diskfull_installation_flat fails on ubuntu-22-ppc64le-devel and
ubuntu-24-ppc64le-devel in build #121 of xcat-core-devel-ubuntu-cd. The
Subiquity installer starts, errors in its early-commands, tars
/var/log/installer to port 8080 and reboots, nine times in 75 minutes. No
system is ever installed, so the address answers from the live installer and
the case ends on
"root@xcat25-cn: Permission denied (publickey,password)".
mkinstall in xCAT-server/lib/xcat/plugins/debian.pm selected
pre.ubuntu.subiquity and then replaced it with pre.ubuntu.ppc64 for every
ppc64 node, whichever installer was in use. pre.ubuntu.ppc64 writes a partman
recipe, and the early-commands append it to /autoinstall.yaml, which Subiquity
cannot parse.
pre.ubuntu.subiquity offered a UEFI branch and a BIOS branch, so a ppc64el node
took the BIOS branch and was given a bios_grub partition. POWER firmware loads
the boot loader from a PReP partition.
install_prescript now returns the script from the installer and the
architecture together, and the ppc64 script is reached only on the
debian-installer path. pre.ubuntu.subiquity gains a PReP branch, taken when
uname reports a POWER machine, which flags an 8M first partition prep and makes
that partition the grub device.
debian_install_prescript.t and the PReP case of ubuntu_subiquity_storage.t fail
without these changes.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
On Ubuntu 26.04 the installed compute node keeps systemd-resolved, which takes
the DNS server from DHCP and no search domain, so /etc/resolv.conf reads
"search .". mypostscript then runs "updateflag.awk $MASTER 3002" with the short
management node name, the flag update fails eleven times, and the node stays at
postbooting until retry_install.sh gives up. The management node does offer
domain-search; the node discards it.
compute.subiquity.tmpl writes /target/etc/netplan/00-xcat-install.yaml with
dhcp4: true alone, so systemd-networkd applies its UseDomains default of no.
Add dhcp4-overrides: use-domains: true to both branches, the one that renames
the interface and the one that matches by MAC alone.
ubuntu_subiquity_installnic.t runs the template's own late-command and asserts
the netplan it writes carries the setting. It fails without this change.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
When a Subiquity install fails, the node stays up with no disk or network
activity and reports nothing. The management node sees only the last progress
line the installer printed, and the case fails on the provisioning timeout
rather than on the error.
The error-commands entry in compute.subiquity.tmpl pipes the log archive into
"nc -l 8080". Subiquity waits for every error command to return, and that
listener returns only when a collector connects, which an unattended install
has none of.
The archive now goes to /run on the installer, and a second command prints the
end of the curtin log to the console, which the management node records with
the rest of the install. XCAT_ERROR_CONSOLE names that console, so the command
can be driven outside the installer.
ubuntu_subiquity_error_commands.t runs the template's own error commands with
nc, tar and tail replaced. Without this change the first assertion times out.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>