A node that has finished installing has chain.currstate "boot", and from
that point the server must stop naming it a boot file: a node given a
netboot script every time it powers on reinstalls itself forever, and does
so silently, because each individual boot looks like a successful one.
Neither backend did that. The wire suite found both.
ISC gated the whole boot-from-disk branch on $doiscsi, so the rule only
fired for a node with an iscsi table row. An ordinary installed node fell
through to the netboot branches below and its xNBA second stage was handed
its own install script. Both the xnba and the pxe branch had it; the pxe
one handed out pxelinux.0.
Kea reads "boot-file-name": "" as *unspecified*, not as "no boot file", so
the empty boot file on the node's reservation did not outrank the classes
every client shares. The always-evaluated xcat-bios class supplied
xcat/xnba.kpxe, the loader came back announcing user class xNBA, and the
per-network class then supplied the network's boot script -- so withholding
only the per-node script withheld nothing. The fix is the same mechanism the
DROP and NOIP classes already use: a xcat-localboot class holding the MACs,
and a "not member('xcat-localboot')" guard on every class that names a boot
file. Kea requires a class to be defined before it is referenced, so
xcat-localboot is written first.
iSCSI is the exception on both backends. Its root disk is on the network and
gPXE is what attaches it, so it still gets a loader: ISC keeps its $doiscsi
branch, and Kea leaves those MACs out of the class.
spec.md S-31 is rewritten to say what correct is -- no boot file at all, on
either request of an xNBA boot, whatever the netboot method -- rather than
the weaker "not its own script", which the second stage above shows is not
the same statement.
conf/localboot.conf asserts the stronger form and gains a netboot=pxe node,
since that branch is written separately and the xnba node passing says
nothing about it. dhcpfixture.sh defines it.
An assertion with != or not-in now holds against a target that is absent
from the reply: a reply naming no boot file has certainly not named the
node's install script. Every other operator still fails on absence, and
"absent" remains the way to assert absence itself.
S-35 asks both backends to tell a node its own name in option 12. The
previous attempt wrote the reservation's hostname with a trailing dot, on
the reading that a name already fully qualified is not qualified again.
That is true of Kea 3.0, which is what it was tested against, and false of
Kea 2.4, which is what Ubuntu 24.04 ships and what CI runs: it appends
ddns-qualifying-suffix regardless and the node was told
"node01.pok.stglabs.ibm.com" where ISC said "node01".
Confirmed against both versions with a reservation on a veth pair, DDNS
configured, asking for option 12:
reservation hostname "probenode." 2.4.1 -> probenode.example.test
3.0.3 -> probenode
hostname "probenode." and option-data 2.4.1 -> probenode.example.test
option-data host-name only 2.4.1 -> probenode
3.0.3 -> probenode
So a reservation's option-data cannot override the field either --
processHostnameOption adds option 12 before appendRequestedOptions runs,
and appendRequestedOptions only fills in what is not already there.
Leaving the hostname field out is the one form that answers with the
node's own name on both versions, and it is also the shape ISC has: an
option statement, separate from the ddns-hostname statement beside it.
kea_boot_for_node already emits the host-name option for every node, so
nothing needed to be added -- only the field removed.
Two consequences, both intended. Kea no longer has a per-reservation DDNS
name and will use whatever the client sends, which is what a cluster with
xCAT's own DNS already relies on; ddns-qualifying-suffix still qualifies
the dynamic clients that have no reservation. And `makedhcp -q` now reads
the name back out of the host-name option, because that is where it is.
The wire suite compared the two backends against xCAT-test/dhcptest/spec.md and
found four places where a node was told different things depending on which
daemon answered. Each is fixed at the side the spec calls correct.
Option 12 (S-35): Kea builds the host-name option out of a reservation's single
`hostname` field and runs it through ddns-qualifying-suffix on the way out, so a
node asking who it was got "node01.cluster.example.com" while ISC, which writes
option 12 and the DDNS name from separate statements, said "node01". A name that
already ends in a dot is fully qualified and is not qualified again, so the
reservation is now written "node01." -- the wire agrees with ISC, and the suffix
still qualifies the dynamic clients that have no reservation. A reservation's
option-data cannot be used for this: processHostnameOption adds option 12 before
appendRequestedOptions runs, and appendRequestedOptions only fills in options
that are not already present. Reservation lookups are made dot-insensitive so
`makedhcp -d` and `-q` keep matching a node by its bare name.
Option 43 / ISAN (S-20): Kea appends an encapsulated space to a reply only when
the option that carries it is configured too, and option 43 is `type: empty`
with `encapsulate: isan`. Naming only the isan-space sub-options left them with
nothing to travel in and the initiator was offered an address with no target, so
the empty container is now named alongside them. ISC needs no equivalent --
declaring `option isan.iqn` builds option 43 for it.
Boot file fallback (S-12): dropping an architecture branch from the ISC if/else
chain let that client fall through to the final `substring(filename,0,1) = null`
catch-all and be handed /yaboot, which is the loader substitution the spec
forbids. Suppressed architectures now get an explicit empty branch. Kea cannot
have this bug: its classes are independent and xcat-fallback excludes every
recognised arch.
Reload (S-40): `systemctl reload kea-dhcp4` returns 0 once the signal is
delivered, and Kea can then reject the config and keep serving the old one while
systemctl still reads active -- a node added by discovery was never adopted.
dhcp4 and dhcp6 intents therefore always carry a control-socket, and a reload
goes over it so the daemon reports whether it took, falling back to
reset-failed + restart when it cannot confirm.
The fixture now proves a daemon actually holds port 67 before it believes the
backend is serving, which is what turned the reload failure from a flake into a
reproducible case.
Verified on the wire against both backends on EL10 with dhcpd 4.4.3 and
kea 3.0.3: all ten cases rc=0 on each. Unit suite 848/848.
Three things the wire cases turned up, only the first of which is xCAT
answering a client wrongly -- the other two never got as far as an answer.
kea-dhcp4 would not start at all. The ONIE class named option "www-server",
which xCAT's dhcpd.conf declares as code 114 = string but which means the
standard option 72 to Kea -- a list of IPv4 addresses, so an installer URL
in one is a configuration error and the whole Kea pass of the wire cases
never ran. Naming the code instead says the same thing to both.
ISC answered a request for an address on a network it has never heard of
with silence, where the spec says DHCPNAK, because "authoritative" was
written into each subnet declaration and read from the subnet the
requested address belongs to -- which, for this case, is none of them. It
is now global as well, which is what Kea's authoritative:true already
covered.
The remaining two were the cases asking the wrong question. ISC sends an
option when the client asks for it, so an ONIE case that never put 114 in
its parameter request list could not see it however the class was
written; and the common-option case announced a PXEClient vendor class
while asserting the cluster's lease time, which is the ten-minute
firmware lease of S-51 rather than S-50. S-50 now says which clients it
is about.
siaddr comes from noderes, in an order neither backend had right. ISC
read xcatmaster only for petitboot and onie, so every other node in a
hierarchical cluster was sent to the management node rather than to its
service node -- and even for those two the address went into the URL
without going into siaddr, so one reply named two different machines.
Kea read xcatmaster for every node, but when nothing named a server it
used my_ip_facing rather than the subnet's value, which is a different
answer whenever networks.tftpserver names a third machine.
Both now read next_server_for_node: the node's tftpserver, then its
xcatmaster, then -- only for the methods that build a URL and so need an
address in hand -- the interface facing the node. A node that named none
of them inherits the subnet's value, which ISC states in the subnet and
Kea states by leaving next-server out of the reservation.
An xcatmaster that does not resolve is now an error on both rather than
silence on one: it is a misconfiguration, and sending the node somewhere
else instead hides it.
Appendix A rows 26 and 27. Unlike the rest, these two are not in
xcat-internal#175 -- the wire cases found them.
Appendix A rows 15 to 18, the remaining spec decisions where the two
backends answered the same client differently.
Row 15, a loader that is not on disk is not named. The ISC architecture
chain named every loader unconditionally, so a client whose loader was
never built spent a full TFTP timeout it could not diagnose; the Kea side
had always left the class out. The chain is now built from a list of
gated branches rather than a literal block, because dropping a branch
from a literal if/else if chain can leave a leading "} else if", which
dhcpd rejects outright.
Row 16, a *NOIP* interface draws no reply. Kea discards a packet assigned
to the class named DROP, and only one such class may exist, so every
dropped MAC in the cluster shares it and the user-context records which
node each belongs to. Syncing one node merges into that list and removing
one node prunes only its own entries, so makedhcp for a single node
cannot bring another node's interface back.
Row 17, a node that boots from disk, and row 18, a node deferring to
proxydhcp, are decided per node. Kea host reservations outrank every
client class, so the reservation has to fall silent -- an empty
boot-file-name -- and let the class carry the answer.
Rows 19 and 20 landed earlier and dictated the same shape: the
reservation names nothing and two mutually exclusive classes, one for the
vendor and one for "not the vendor", decide between them, because Kea
evaluates every class independently and has no else.
Appendix A decisions 19 and 20. ISC writes both as an if/else inside the
node's own host block:
if option vendor-class-identifier = "ScaleMP" { filename = "vsmp/pxelinux.0"; }
else { filename = "pxelinux.0"; }
Kea has no else, and a reservation outranks every client class, so
anything the reservation names cannot be overridden by the vendor class
that follows it. A ScaleMP hypervisor was therefore handed the
reservation's pxelinux.0 and booted the wrong binary, and an ISAN
initiator -- which reads its initiator name and root path out of option
43 and not out of option 17 -- was sent the standard form it ignores and
nothing it could use.
Both are now written as pairs of mutually exclusive per-node classes,
with the reservation deliberately naming neither the pxe boot file nor an
ISAN node's root path so that a class can decide. This is the same
mechanism the xNBA second stage already used, so the generator, the sync
and the removal are generalised from xnba to per-node classes and each
class carries the purpose that identifies it.
The isan option space ISC declares as "option space isan" is declared to
Kea as an option-def encapsulating option 43, with the initiator name in
sub-option 203 and the root path in 201.
Appendix A decisions 9, 21, 22 and 24. Each one is a difference an
operator never chose: the backend is picked by distribution version, so
whichever side is wrong is wrong on half the clusters.
9 -- Kea never loaded libdhcp_bootp.so, so a BOOTP-only client that ISC
answered got nothing. The hook now sits alongside host_cmds in one
hooks-libraries array, and when the hooks package is absent the
operator is told which clients that leaves unanswered rather than
being left to find out from a node that never boots.
21 -- The ISC host statements sent "send host-name", which is a dhcpd
*client* keyword: it never reached option 12 and the node was
handed no hostname at all. Kea has always sent one.
22 -- The PXE short lease existed on ISC in name only. dhcpd applies
min-lease-time after max-lease-time, so the cluster default won and
a pool address taken by a PXE ROM was held for half a day. The
class now names all three bounds, and Kea is given the same 600
seconds through xcat-pxe-lease.
24 -- ISC writes "authoritative;" into every generated subnet; Kea
defaults to the opposite. A node that moved rack asked to keep an
address from the network it had left and was answered with silence,
so it waited out a lease that would never be renewed instead of
being told to start over.
Also makes $xCAT_plugin::dhcp::callback a package variable. It was a file
lexical, so the tests' local() set something nothing read, and warnings
were only ever captured because an earlier process_request had left its
collector behind.
Eight of the parity decisions in the spec's Appendix A, where ISC dhcpd and
Kea answered the same frame differently and the operator never chose which
backend they got.
ISC gains two branches its if/else chain never had, so the client falls
through to /yaboot no longer:
- 0x000c ppc64 is given /boot/grub2/grub2.ppc. yaboot is not a UEFI
loader and cannot boot one of these machines (decision 2).
- 0x0010 is the same x86-64 UEFI firmware and the same loader as 0x0007,
announced by a machine set to fetch it over HTTP (decision 3).
Kea gains what ISC has always had:
- Etherboot, which predates option 93 and says what it is in option 60
alone, is recognised and given the BIOS loader (decision 10).
- onie_vendor is answered per subnet, not only per node: a switch
announces it on its first boot, before anyone has defined it as a node
(decision 11).
- A client nothing else recognises is given /yaboot. Kea has no else, so
the condition is the negation of every architecture and vendor class
another rule answers, rather than a dependence on class ordering
(decision 7).
- netboot=nimol is given /vios/nodes/<node> (decision 6).
and drops two answers ISC never gave:
- No BIOS loader on disk no longer means pxelinux.0 in its place. Naming
a file that is not there costs the client a timeout it cannot diagnose,
and a different loader boots something nobody asked for; the class is
simply not written (decision 8).
- netboot=petitboot sends the conf-file and nothing else. petitboot acts
on a boot file name when it sees one, so naming one as well sent the
machine after a TFTP fetch that never happens on ISC (decision 12).
The per-node statement for a UEFI node running netboot=xnba matched its two
architecture ids with a parenthesised alternation:
else if <user class> and (option client-architecture = 00:09
or option client-architecture = 00:07) { ... }
ISC dhcpd has no parenthesised grouping in its expression grammar, so that
condition cannot parse. The node-level second stage it guards -- the
per-node .uefi script that keeps two machines chainloading at the same
moment from running the same script -- has never been reachable.
Written as two branches instead, one per architecture id, which is how the
per-network chain in BootPolicy.pm already spells the same test.
dhcp_isc_expression_grouping.t is a guard rather than a test of this branch:
the per-node statements are built inline in addnode against a live database
and cannot be called from a unit test, so it scans the plugin for the one
construct that produces the failure and checks the rendered per-network
chain and the user class test for it as well. dhcpd's tokens only -- Perl's
own parenthesised `exists` and `not` are excluded, and a bareword `option`
after an opening paren cannot be Perl.
isc_xnba_user_class_test wrapped its two alternatives in parentheses so the
trailing `and option client-architecture = ...` would bind to the whole
alternation rather than to the second branch alone. ISC dhcpd has no
parenthesised grouping in its expression grammar, so every generated
dhcpd.conf became unparseable:
/etc/dhcp/.../dhcpd.conf line 11: left brace expected.
if (option user-class-identifier = "xNBA" or substring(option user-cl
^
followed by a cascade of "expecting a parameter or declaration" at each
`} else`, and a daemon that will not start at all.
suffix() gives the same coverage in a single expression that needs no
grouping: it returns the last four bytes, which are "xNBA" whether the
client sent the bare string or the RFC 3004 length-prefixed "\x04xNBA".
The unit test now pins the whole expression rather than each alternative,
and asserts the two properties that broke the daemon -- no leading paren,
no `or` -- so a future rewrite cannot reintroduce grouping unnoticed.
A chainloaded second stage announces itself in the user class, option 77, and
must be handed a script URL rather than the loader it just ran -- otherwise it
chainloads itself forever and the machine never finishes booting.
RFC 3004 length-prefixes each string in option 77. Plenty of firmware sends the
bare string instead, and the same loader sends either depending on how it was
built, so both encodings have to be recognised. The Kea policy has always
accepted both. The ISC side compared only `option user-class-identifier =
"xNBA"`, which is the bare form, so a conforming loader booting against an ISC
management node loops -- and does so silently, since from the server's side
every exchange looks like a normal first-stage boot.
Put the test in one place, isc_xnba_user_class_test, and use it from both the
per-network architecture chain and the three per-node statements. The per-node
ones reach dhcpd through omshell, so the helper takes the quoting its caller
needs. `substring(option user-class-identifier, 1, 4)` skips the length byte;
option 77 is declared as a plain string, and the same construct is already used
by the onie_vendor branch a few lines below.
debian_sysconfig_interface_keys fell back to querying the local dpkg when it
was handed no version. The caller always passes one, so the fallback bought
nothing -- and it made the "version could not be established" case answer
differently depending on whether the machine running the code happened to have
isc-dhcp-server installed. It passed on a developer's EL workstation and failed
on the Ubuntu CI runner, which is the failure it was always going to produce.
An undefined version now means what the caller means by it: could not be
established, so write every spelling. Covered by a test that stubs the lookup
to prove the answer no longer depends on the machine.
makedhcp wrote the interface list into /etc/default/isc-dhcp-server as
INTERFACES="..."
which is no longer always the variable the daemon is started with. The
systemd unit sources that file and expands exactly one variable onto dhcpd's
command line, and which one changed with the package:
14.04 4.2.4-7ubuntu12 sysvinit only $INTERFACES
16.04 4.3.3-5ubuntu12 unit $INTERFACES
18.04 4.3.5-3ubuntu7 unit $INTERFACES
20.04 4.4.1-2.1ubuntu5 unit $INTERFACES
22.04 4.4.1-2.3ubuntu2 unit $INTERFACESv4
24.04 4.4.3-P1-4ubuntu2 unit $INTERFACESv4
26.04 4.4.3-P1-4ubuntu2 unit $INTERFACESv4
with the matching v6 unit reading $INTERFACES before the change and
$INTERFACESv6 after it. Note the boundary is not an upstream ISC release:
20.04 and 22.04 both ship upstream 4.4.1 and are told apart only by the
Debian revision, so it has to be decided on the whole package version.
The sysvinit script does copy INTERFACES into INTERFACESv4 when the latter is
empty, but the unit is what starts the daemon on any of these releases and it
has no such bridge. So from 22.04 on, the unit ran
exec dhcpd -user dhcpd -group dhcpd -f -4 ... -cf $CONFIG_FILE $INTERFACESv4
against a variable xCAT never set. An unset variable expands to nothing, so
dhcpd was launched with no interface argument at all. It does not fail for
that: it binds every interface it can find and only warns about the ones with
no subnet declaration. site.dhcpinterfaces and servicenode.dhcpinterfaces
were therefore silently inert -- the provisioning NIC was served because
makedhcp had also emitted a subnet stanza for it, not because anything
honoured the setting, and any other interface on that subnet was served
alongside it.
The line match compounded it. m/^$dhcpd_key/ is not anchored on the
assignment, and INTERFACES is a prefix of both variables the package ships.
Since 18.04 the postinst seeds only INTERFACESv4="" and INTERFACESv6="" --
no INTERFACES line exists to match -- so the rewrite claimed those two lines
instead and overwrote both, discarding whatever debconf or the administrator
had put there and leaving a file with two INTERFACES assignments and nothing
either unit reads.
Pick the variables through a dispatcher keyed on the installed
isc-dhcp-server version, so both behaviours are served rather than one being
traded for the other, and anchor the match on the '=' so a key cannot claim a
line it is merely a prefix of. When the version cannot be established the
dispatcher writes every spelling, since an unset variable is the outcome that
leaves dhcpd bound to everything. Duplicate assignments left behind by the
old writer are collapsed to one, which repairs a file already damaged on an
upgraded management node.
The EL and SLES paths keep their single DHCPDARGS / DHCPD_INTERFACE /
DHCPD6_INTERFACE key and are unchanged.
On Debian and Ubuntu, makedhcp writes the list of interfaces it is serving
into /etc/default/isc-dhcp-server using the key set at dhcp.pm:2296:
$dhcpd_key = "INTERFACES";
That key stopped being the one the daemon is started with. The variable the
isc-dhcp-server systemd unit expands onto dhcpd's command line changed with
the package; verified by unpacking the archive's own debs:
trusty 4.2.4-7ubuntu12 sysvinit only INTERFACES
xenial 4.3.3-5ubuntu12 unit $INTERFACES
bionic 4.3.5-3ubuntu7 unit $INTERFACES
focal 4.4.1-2.1ubuntu5 unit $INTERFACES
jammy 4.4.1-2.3ubuntu2 unit $INTERFACESv4
noble 4.4.3-P1-4ubuntu2 unit $INTERFACESv4 (+ a v6 unit
reading $INTERFACESv6)
and every package from bionic onward ships a default file whose only
interface variables are INTERFACESv4 and INTERFACESv6, both seeded empty by
postinst. No INTERFACES line ships at all.
Two things go wrong on jammy and later.
The unit runs `exec dhcpd ... -cf $CONFIG_FILE $INTERFACESv4`. That variable
is never set, so it expands to nothing and dhcpd is launched with no
interface argument. dhcpd does not fail for this: it binds every interface
it can find and merely warns about the ones with no subnet declaration. So
site.dhcpinterfaces and servicenode.dhcpinterfaces are silently inert. The
provisioning NIC is served only because makedhcp also emitted a subnet
stanza for it, not because anything honoured the setting, and any other
interface on that same subnet -- a bridge port, a bond member, a second NIC
-- is served too.
The line match is m/^$dhcpd_key/, which is not anchored on the assignment.
"INTERFACES" is a prefix of both variables the package ships, so the rewrite
claims the INTERFACESv4 and INTERFACESv6 lines and overwrites both. The file
is left holding two INTERFACES lines and nothing the units read, discarding
whatever debconf or the administrator had set there.
Extract the rewrite into _sysconfig_interfaces_content() with no change in
behaviour, and add a unit test that asserts the effect rather than the
spelling: it writes the produced file to a temp path and sources it with sh
exactly as the unit does, then checks what would land on dhcpd's command
line. The test is red against the current key and stays red until the writer
sets the variables the daemon is actually started with.
The plugin is loaded and install_darch is called for the architectures
xCAT installs Ubuntu on and two it does not. Against the previous plugin the
test fails on the missing function: the mapping was inline in mkinstall.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The grub-common assertion fails against the previous metapackage. The rest pins
what must NOT change: a riscv64 management node still recommends the x86 boot
payload and the Genesis images of the other architectures, because it serves
them to the nodes it provisions.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The stubbed test proves the decisions copycd makes; this one proves the
artifact. It runs the real grub-mkimage over the grub2 package of copied media
and checks the image against the validation nodeset depends on, so a package
layout change or a grub-mkimage that stops accepting these inputs is caught
where it happens.
It needs media and the grub2 tools, so it skips without
XCAT_TEST_UBUNTU_RISCV64_MEDIA. Run on a management node against the 24.04 and
26.04 riscv64 trees.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The mirror assertions passed a Debian architecture straight in and so never
exercised the conversion genimage performs first. They now start from the xCAT
osarch value, which is what let the 32-bit x86 token reach the wrong archive.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Four assertions fail against the previous debian/control: the dependency has no
architecture restriction, and dpkg's own parser still reports it for riscv64.
The last assertion pins that the restriction drops nothing else.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Calls the selection the template performs. Covers the ports archive for
ppc64el and both spellings of ppc64, the main archive for every x86
spelling, an unknown architecture keeping the previous default, and
site.ubuntu_apt_mirror overriding both without an empty value blanking the
mirror.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Calls the escaping the plugin performs. Covers the Ubuntu installer seed
keeping the arguments after it, every separator grub2 recognizes, a command
line without one staying byte for byte the same, escaped and quoted values
surviving unchanged, a separator beside a quoted span in the same word, and
a variable reference being left alone so BOOTIF still expands.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Reads the list imgutils returns. Covers the drivers a node needs to reach
its root filesystem, the overlay module, the architectures that already
worked keeping theirs, and an unknown architecture still getting an empty
list.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Evaluates the selection genimage performs, so the test tracks the script.
Covers the ports archive for ppc64el and riscv64, the main archive for
amd64 and i386, and site.ubuntu_apt_mirror overriding both without an
empty value blanking the mirror.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Extends the builder tests with the per-package architecture lists and the
architectures a release declares, and adds the generated mklocalrepo.sh
mapping a riscv64 host to its own repository. Also checks that xcat and
xcatsn declare riscv64 without losing amd64 or ppc64el.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Drives the build with dpkg-deb and grub-mkimage shadowed by stubs, so the
assertions read the arguments that decide whether the image can boot over
the network: the firmware format, the prefix nodeset writes into, the
module directory, and the network modules. Also covers the warning a node
without a loader gets, an interrupted build leaving nothing behind, a real
riscv64 loader being kept, text, a truncated image, another architecture's
loader, images with nothing to execute and an image built for another boot
path all being replaced, and a rejected loader being gone when the media
cannot replace it.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Pin that both releases carry a riscv64 list, that it holds what the x86_64 list
holds, and that it keeps the kernel and the NFS client a diskless node needs.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Build a root filesystem for each architecture and run genimage's selection over
it, so the riscv64 libraries are taken from their own directory and the other
architectures keep the files they take today.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Pin the kernel and initrd the riscv64 live image carries, that the kernel name
the other live images use is not accepted for it, and that the architecture the
media reports maps to riscv64 in both directions.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Shell unit tests were introduced under xCAT-test/autotest/bats, but the existing source-tree unit suite already lives directly under xCAT-test/unit. Keeping the BATS suite under xCAT-test/bats makes the unit-test layout consistent and keeps autotest reserved for xcattest-driven functional cases.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The go-xcat shell behavior tests were written as Perl harnesses, which made the shell assertions harder to read and kept shell-specific setup outside a native shell test framework.
Add BATS to the GitHub Actions dependency set, run BATS tests from the same preserved source tree as the Perl unit suite, and move the go-xcat repository checks into xCAT-test/autotest/bats.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Both scripts run for real with PATH holding one directory that carries a
tcpdump stand-in, which records the path and arguments it was started with.
The plain run stops at the interface step, before any socket or capture. The
capture runs inside a private network and mount namespace, where the loopback
interface is the only one, its default route keeps the DHCP discover on the
host, a tmpfs over /tmp holds the dump file, and /usr/sbin/tcpdump is hidden
so the previous guard fails there on every host. Hosts that cannot create the
namespace skip that part.
The lists are resolved through get_pkglist_file_name, the resolver that picks
one for an osimage, and read through get_pkglist_tex, the parser that
produces OSPKGS. 59 of 94 assertions fail against the previous lists: the dead
names, the missing 26.04 lists, unixodbc, qemu-utils and libvirt packages,
and every kvm osimage resolving to the shared list with the old emulator.
The check functions and the installer functions that call them are taken
from the shipped script and run against a dnf stand-in that records its
probes. 27 of 40 assertions fail against the previous
script: EL10 and CentOS Stream were not checked, the probe went through dnf
list, a failed query read as a missing repository, and the CRB message
carried a repository file with signature checks disabled.