2
0
mirror of https://github.com/xcat2/xcat-core.git synced 2026-10-07 10:06:39 +00:00
Commit Graph

28127 Commits

Author SHA1 Message Date
Daniel Hilst 4d8d686747 fix(provtest): make the fixture's teardown finish what it started
Every loop over a list of names ran the xCAT clients on standard
input, so the first client read the rest of the file and the loop
ended after one name.  Teardown said it had removed the fixture while
five nodes, an osimage, a network and an unrestored site table were
still there, and the next setup recorded those leftovers as the state
to restore to.  The loops now read on descriptor 3.

`nodeset offline` is not enough either: it gives up when it cannot
reach the DHCP backend and never removes the staged kernel and initrd.
The artefacts the fixture's own nodes and images can produce are now
removed by name.

check refuses on any leftover node, not just the first.
2026-09-11 09:24:55 -03:00
Daniel Hilst a1bc164b5a ci: run the provtest checks after ci_test and before the DHCP wire test
The provisioning cases go where a node meets them: after the whole
ci_test set, so they are not reconfiguring the machine underneath it, and
before the DHCP phases.

run_provtest_unit_tests runs the offline half first -- the Python suite
and a parse of every scenario -- so a file that does not even load is
reported as that rather than as a case failure, and skips rather than
fails when python3 is absent. run_prov_wire_test then runs the prov_wire
cases once, not once per backend: nothing in them depends on which DHCP
server is installed, and what a node is told over DHCP is dhcptest's
subject.
2026-09-11 09:01:00 -03:00
Daniel Hilst 711c4870db test(provtest): add the autotest cases and ship the tree in the packages
Ten cases: two offline ones labelled ci_test, the unit suite and a parse
of every shipped scenario, and eight wire cases labelled prov_wire, one
per stage of the chain a node walks -- DNS, loader and config, install
tree, genesis and discovery artefacts, discovery requests, the xcatd
protocols, the install monitor, and the three failures that look alike
from a console.

The wire cases are deliberately not ci_test: while one runs the
management node has an extra veth pair and namespace, four nodes, a
rewritten zone and possibly a moved http port. Each restores all of it
from a trap, and each passes with a message when the host cannot host it.

provfixture.sh holds every xCAT command; the scenarios read none of it.
2026-09-11 09:00:54 -03:00
Daniel Hilst 2efcbcde08 test(provtest): assert what xCAT does rather than what the spec assumed
Running the scenarios against a live management node found nine places
where they asserted something xCAT does not do. Corrected here, each with
the reason in the file:

- install configs carry no xcatd= or destiny=; those are written for the
  genesis states only.
- #END OF SCRIPT is plain-transport framing; over TLS a reply ends with
  <serverdone>.
- getcredentials is named in <arg> and answered as <data><content/>
  <desc/></data>, so both fields are addressable now.
- currstate for an install is "install <osimage>", so the operator is
  starts-with.
- bootparams is empty for an ordinary install, so kernel and initrd
  cannot be asserted from getdestiny.
- a server never reached sends no <error>; assert the handshake instead.
- an NXDOMAIN answer is the expected result of a removal scenario.
- P-75 has to precede P-74: a withdrawn name cannot be nodeset again.

Also adds netutil.hex_net: a per-network config is named by as many hex
digits as the netmask covers, not by all eight.
2026-09-11 09:00:47 -03:00
Daniel Hilst 721bfe710e fix(dhcptest): find the dhcptest tree from a checkout three levels up
The fixture falls back to a checkout when the installed tree is absent,
and the relative path was one level short: from
autotest/testcase/dhcptest it has to climb to xCAT-test, not to
autotest. In a checkout the fallback therefore resolved to nothing and
every case skipped, silently and with a pass.
2026-09-11 09:00:40 -03:00
Daniel Hilst 5a969c8eb7 fix(credentials): refuse a getcredentials request naming no credential
Which credential is wanted is named in <arg>. A request without one
reached a dereference of an undefined array ref, so xcatd died inside the
plugin and shipped the Perl error back to the client instead of ignoring
the request.

Check for the element before the callback is made, and log and drop the
request the same way a bad callback port is dropped. A client that sends
a malformed request now learns nothing about the server, which is the
point: this path hands out cluster certificates to anything whose address
resolves.
2026-09-11 09:00:35 -03:00
Daniel Hilst 9f03656318 docs(test): move the wire specifications out of the test tree
Both spec.md files are design documents for this organisation: they argue
about what xCAT ought to put on the wire, cite the plugin lines that decide
it, and record where the original proposal was wrong. Shipped under
xCAT-test/ they would land in an RPM on every management node and be offered
upstream as part of a test directory, which is not what they are for.

They now live in the internal repository as specs/dhcp-wire.md and
specs/provision-chain.md. The suites keep their clause tags, so a failing
case still names the clause it belongs to.
2026-09-11 08:09:54 -03:00
Daniel Hilst 86533527e0 test(provtest): unit-test the offline half, and the rules it relies on
154 tests, standard library only, no network and no root: the config parser,
the assertion language, $step.field resolution, the offline machine, TAP
output, address arithmetic, and the decoding of captured dig, curl and tftp
output -- including an NXDOMAIN whose only record is an SOA, and a TFTP fetch
that created an empty file before finding the server had nothing to send.

test_boundaries.py asserts rules about this code rather than about xCAT: only
proc.py starts a process, no module opens a database or names an xCAT command,
and no step type can run one. A suite that asked xCAT what to expect would
pass on a cluster that could not boot a node.
2026-09-11 08:04:31 -03:00
Daniel Hilst 70a114e2c8 fix(provtest): assert against a whole list, not its first element
An assertion resolved its target through Reply.field, which collapses a list
to its first item. A name with two A records could therefore never satisfy
`data == <the second one>`, and _equal's list branch -- written for exactly
that case -- was unreachable.

Reply.whole returns the container intact and assertions use it. References
keep the old behaviour: $step.data substitutes one value into a command line,
so the first is the right answer there.
2026-09-11 08:04:31 -03:00
Daniel Hilst 1f06d22883 test(provtest): assert the 75 provision clauses as wire scenarios
Sixteen files, fifty scenarios, one per behaviour the specification names.
The cross-stage ones are the point: a config is fetched over TFTP, the
kernel path is read out of it, and that path is fetched. No file says what
the kernel will be called, which is what separates testing xCAT from
comparing xCAT's output against a copy of it.

The four tftp-*.conf files are alternatives, not a set -- a node has one
netboot method -- and the fixture selects the one it was defined with.

conf/tftp-grub2.conf @P-70 meets dhcptest at the only shared assertion:
the boot file DHCP named is fetchable under exactly that name.
2026-09-11 07:55:59 -03:00
Daniel Hilst 7ea4bbdfb4 test(provtest): add provtest, a wire-level provision chain client
provtest asks the management node the questions a booting node asks and
asserts on the answers. DNS, TFTP and HTTP are driven by dig, curl and
tftp, so the result is what a real client sees rather than what this
tool's own protocol parser believed. The xcatd stages use sockets: they
must choose their source address, because xcatd names a client by the
reverse lookup of the address a connection arrived from.

Scenarios are declarative .conf files sharing dhcptest's assertion
grammar, exit codes and TAP output. Nothing here reads the xCAT database
or runs an xCAT command; every expected value arrives via --set.

validate and list run offline, in a checkout, with no root.
2026-09-11 07:55:59 -03:00
Daniel Hilst 6f6acc1b08 docs(provtest): specify the provision chain a booting node observes
The provision chain has no specification anywhere: what a node fetches, in
what order, and what counts as a correct answer live only in the plugins
that emit the artefacts. A cluster can pass every ci_test case and still
boot nothing, because no test in the tree ever acts as a client.

spec.md states 75 clauses, @P-01 to @P-75, each citing the source that
decides it, covering DNS, TFTP, HTTP, discovery, findme, the xcatd request
verbs and the install monitor. Three clauses correct the proposal against
what xcatd actually does; Appendix A records which and why.

README.md documents the tool that runs them.
2026-09-11 07:55:46 -03:00
Daniel Hilst a64dfe149a fix(dhcp): stop handing an installed node a boot file on either backend
A node that has finished installing has chain.currstate "boot", and from
that point the server must stop naming it a boot file: a node given a
netboot script every time it powers on reinstalls itself forever, and does
so silently, because each individual boot looks like a successful one.

Neither backend did that. The wire suite found both.

ISC gated the whole boot-from-disk branch on $doiscsi, so the rule only
fired for a node with an iscsi table row. An ordinary installed node fell
through to the netboot branches below and its xNBA second stage was handed
its own install script. Both the xnba and the pxe branch had it; the pxe
one handed out pxelinux.0.

Kea reads "boot-file-name": "" as *unspecified*, not as "no boot file", so
the empty boot file on the node's reservation did not outrank the classes
every client shares. The always-evaluated xcat-bios class supplied
xcat/xnba.kpxe, the loader came back announcing user class xNBA, and the
per-network class then supplied the network's boot script -- so withholding
only the per-node script withheld nothing. The fix is the same mechanism the
DROP and NOIP classes already use: a xcat-localboot class holding the MACs,
and a "not member('xcat-localboot')" guard on every class that names a boot
file. Kea requires a class to be defined before it is referenced, so
xcat-localboot is written first.

iSCSI is the exception on both backends. Its root disk is on the network and
gPXE is what attaches it, so it still gets a loader: ISC keeps its $doiscsi
branch, and Kea leaves those MACs out of the class.

spec.md S-31 is rewritten to say what correct is -- no boot file at all, on
either request of an xNBA boot, whatever the netboot method -- rather than
the weaker "not its own script", which the second stage above shows is not
the same statement.

conf/localboot.conf asserts the stronger form and gains a netboot=pxe node,
since that branch is written separately and the xnba node passing says
nothing about it. dhcpfixture.sh defines it.

An assertion with != or not-in now holds against a target that is absent
from the reply: a reply naming no boot file has certainly not named the
node's install script. Every other operator still fails on absence, and
"absent" remains the way to assert absence itself.
2026-09-10 17:36:44 -03:00
Daniel Hilst cb345bba06 fix(dhcp): let a Kea reservation name a node without qualifying it
S-35 asks both backends to tell a node its own name in option 12. The
previous attempt wrote the reservation's hostname with a trailing dot, on
the reading that a name already fully qualified is not qualified again.
That is true of Kea 3.0, which is what it was tested against, and false of
Kea 2.4, which is what Ubuntu 24.04 ships and what CI runs: it appends
ddns-qualifying-suffix regardless and the node was told
"node01.pok.stglabs.ibm.com" where ISC said "node01".

Confirmed against both versions with a reservation on a veth pair, DDNS
configured, asking for option 12:

  reservation hostname "probenode."          2.4.1 -> probenode.example.test
                                             3.0.3 -> probenode
  hostname "probenode." and option-data      2.4.1 -> probenode.example.test
  option-data host-name only                 2.4.1 -> probenode
                                             3.0.3 -> probenode

So a reservation's option-data cannot override the field either --
processHostnameOption adds option 12 before appendRequestedOptions runs,
and appendRequestedOptions only fills in what is not already there.
Leaving the hostname field out is the one form that answers with the
node's own name on both versions, and it is also the shape ISC has: an
option statement, separate from the ddns-hostname statement beside it.

kea_boot_for_node already emits the host-name option for every node, so
nothing needed to be added -- only the field removed.

Two consequences, both intended. Kea no longer has a per-reservation DDNS
name and will use whatever the client sends, which is what a cluster with
xCAT's own DNS already relies on; ddns-qualifying-suffix still qualifies
the dynamic clients that have no reservation. And `makedhcp -q` now reads
the name back out of the host-name option, because that is where it is.
2026-09-10 16:44:42 -03:00
Daniel Hilst 8f90c68c82 test(dhcptest): assert an installed node is sent no script, and cite the spec
Two gaps in what the wire suite proves, both about traceability rather than
about a new behaviour.

S-31 had no wire case. A node whose chain.currstate is "boot" has an
operating system and must be left to start it; a node handed a netboot
script every time it powers on reinstalls itself forever, and does so
silently, because each individual boot looks like a successful one. The
Kea side of this is covered by dhcp_kea_plugin_intent.t, but the ISC side
is built inline in addnode against a live database and cannot be reached
by a unit test at all.

The new case asks twice with the same MAC: once as firmware with no user
class, once announcing the xNBA user class the first stage sets. Only the
second request can be answered with the node's script, so only the second
request can see the bug. What is asserted is S-31's own wording -- not
"no boot file", which would forbid the stage-1 binary a server is entitled
to send, but "not the node's xNBA script". The node is defined with
netboot=xnba because only that method generates a second stage.

The remaining ten scenarios already proven on the wire now say which
specification scenario they prove, in the same "S-nn:" form the rest of
the suite uses. That raises the scenarios cited by a wire case from 36 to
47 of the spec's 73, without running anything new -- the coverage was
there and was not traceable.
2026-09-10 16:26:18 -03:00
Daniel Hilst 05185b13db fix(dhcp): make Kea and ISC agree on the wire where the spec says they must
The wire suite compared the two backends against xCAT-test/dhcptest/spec.md and
found four places where a node was told different things depending on which
daemon answered. Each is fixed at the side the spec calls correct.

Option 12 (S-35): Kea builds the host-name option out of a reservation's single
`hostname` field and runs it through ddns-qualifying-suffix on the way out, so a
node asking who it was got "node01.cluster.example.com" while ISC, which writes
option 12 and the DDNS name from separate statements, said "node01". A name that
already ends in a dot is fully qualified and is not qualified again, so the
reservation is now written "node01." -- the wire agrees with ISC, and the suffix
still qualifies the dynamic clients that have no reservation. A reservation's
option-data cannot be used for this: processHostnameOption adds option 12 before
appendRequestedOptions runs, and appendRequestedOptions only fills in options
that are not already present. Reservation lookups are made dot-insensitive so
`makedhcp -d` and `-q` keep matching a node by its bare name.

Option 43 / ISAN (S-20): Kea appends an encapsulated space to a reply only when
the option that carries it is configured too, and option 43 is `type: empty`
with `encapsulate: isan`. Naming only the isan-space sub-options left them with
nothing to travel in and the initiator was offered an address with no target, so
the empty container is now named alongside them. ISC needs no equivalent --
declaring `option isan.iqn` builds option 43 for it.

Boot file fallback (S-12): dropping an architecture branch from the ISC if/else
chain let that client fall through to the final `substring(filename,0,1) = null`
catch-all and be handed /yaboot, which is the loader substitution the spec
forbids. Suppressed architectures now get an explicit empty branch. Kea cannot
have this bug: its classes are independent and xcat-fallback excludes every
recognised arch.

Reload (S-40): `systemctl reload kea-dhcp4` returns 0 once the signal is
delivered, and Kea can then reject the config and keep serving the old one while
systemctl still reads active -- a node added by discovery was never adopted.
dhcp4 and dhcp6 intents therefore always carry a control-socket, and a reload
goes over it so the daemon reports whether it took, falling back to
reset-failed + restart when it cannot confirm.

The fixture now proves a daemon actually holds port 67 before it believes the
backend is serving, which is what turned the reload failure from a flake into a
reproducible case.

Verified on the wire against both backends on EL10 with dhcpd 4.4.3 and
kea 3.0.3: all ten cases rc=0 on each. Unit suite 848/848.
2026-09-10 15:57:33 -03:00
Daniel Hilst 9da497cac4 fix(dhcp): let Kea start, let ISC refuse, and ask for what is asserted
Three things the wire cases turned up, only the first of which is xCAT
answering a client wrongly -- the other two never got as far as an answer.

kea-dhcp4 would not start at all. The ONIE class named option "www-server",
which xCAT's dhcpd.conf declares as code 114 = string but which means the
standard option 72 to Kea -- a list of IPv4 addresses, so an installer URL
in one is a configuration error and the whole Kea pass of the wire cases
never ran. Naming the code instead says the same thing to both.

ISC answered a request for an address on a network it has never heard of
with silence, where the spec says DHCPNAK, because "authoritative" was
written into each subnet declaration and read from the subnet the
requested address belongs to -- which, for this case, is none of them. It
is now global as well, which is what Kea's authoritative:true already
covered.

The remaining two were the cases asking the wrong question. ISC sends an
option when the client asks for it, so an ONIE case that never put 114 in
its parameter request list could not see it however the class was
written; and the common-option case announced a PXEClient vendor class
while asserting the cluster's lease time, which is the ten-minute
firmware lease of S-51 rather than S-50. S-50 now says which clients it
is about.
2026-09-10 14:18:41 -03:00
Daniel Hilst 26e8e277ef fix(dhcp): give both backends one answer for where a node is sent
siaddr comes from noderes, in an order neither backend had right. ISC
read xcatmaster only for petitboot and onie, so every other node in a
hierarchical cluster was sent to the management node rather than to its
service node -- and even for those two the address went into the URL
without going into siaddr, so one reply named two different machines.
Kea read xcatmaster for every node, but when nothing named a server it
used my_ip_facing rather than the subnet's value, which is a different
answer whenever networks.tftpserver names a third machine.

Both now read next_server_for_node: the node's tftpserver, then its
xcatmaster, then -- only for the methods that build a URL and so need an
address in hand -- the interface facing the node. A node that named none
of them inherits the subnet's value, which ISC states in the subnet and
Kea states by leaving next-server out of the reservation.

An xcatmaster that does not resolve is now an error on both rather than
silence on one: it is a misconfiguration, and sending the node somewhere
else instead hides it.

Appendix A rows 26 and 27. Unlike the rest, these two are not in
xcat-internal#175 -- the wire cases found them.
2026-09-10 14:01:22 -03:00
Daniel Hilst 980b7de169 fix(dhcp): close the last four ISC/Kea parity gaps
Appendix A rows 15 to 18, the remaining spec decisions where the two
backends answered the same client differently.

Row 15, a loader that is not on disk is not named. The ISC architecture
chain named every loader unconditionally, so a client whose loader was
never built spent a full TFTP timeout it could not diagnose; the Kea side
had always left the class out. The chain is now built from a list of
gated branches rather than a literal block, because dropping a branch
from a literal if/else if chain can leave a leading "} else if", which
dhcpd rejects outright.

Row 16, a *NOIP* interface draws no reply. Kea discards a packet assigned
to the class named DROP, and only one such class may exist, so every
dropped MAC in the cluster shares it and the user-context records which
node each belongs to. Syncing one node merges into that list and removing
one node prunes only its own entries, so makedhcp for a single node
cannot bring another node's interface back.

Row 17, a node that boots from disk, and row 18, a node deferring to
proxydhcp, are decided per node. Kea host reservations outrank every
client class, so the reservation has to fall silent -- an empty
boot-file-name -- and let the class carry the answer.

Rows 19 and 20 landed earlier and dictated the same shape: the
reservation names nothing and two mutually exclusive classes, one for the
vendor and one for "not the vendor", decide between them, because Kea
evaluates every class independently and has no else.
2026-09-10 13:56:30 -03:00
Daniel Hilst ae7a572132 fix(dhcp): give Kea the ScaleMP and ISAN vendor forms
Appendix A decisions 19 and 20. ISC writes both as an if/else inside the
node's own host block:

  if option vendor-class-identifier = "ScaleMP" { filename = "vsmp/pxelinux.0"; }
  else { filename = "pxelinux.0"; }

Kea has no else, and a reservation outranks every client class, so
anything the reservation names cannot be overridden by the vendor class
that follows it. A ScaleMP hypervisor was therefore handed the
reservation's pxelinux.0 and booted the wrong binary, and an ISAN
initiator -- which reads its initiator name and root path out of option
43 and not out of option 17 -- was sent the standard form it ignores and
nothing it could use.

Both are now written as pairs of mutually exclusive per-node classes,
with the reservation deliberately naming neither the pxe boot file nor an
ISAN node's root path so that a class can decide. This is the same
mechanism the xNBA second stage already used, so the generator, the sync
and the removal are generalised from xnba to per-node classes and each
class carries the purpose that identifies it.

The isan option space ISC declares as "option space isan" is declared to
Kea as an option-def encapsulating option 43, with the initiator name in
sub-option 203 and the root path in 201.
2026-09-10 13:43:42 -03:00
Daniel Hilst 1a531128de fix(dhcp): close four more ISC/Kea drifts from the parity spec
Appendix A decisions 9, 21, 22 and 24. Each one is a difference an
operator never chose: the backend is picked by distribution version, so
whichever side is wrong is wrong on half the clusters.

9  -- Kea never loaded libdhcp_bootp.so, so a BOOTP-only client that ISC
      answered got nothing. The hook now sits alongside host_cmds in one
      hooks-libraries array, and when the hooks package is absent the
      operator is told which clients that leaves unanswered rather than
      being left to find out from a node that never boots.

21 -- The ISC host statements sent "send host-name", which is a dhcpd
      *client* keyword: it never reached option 12 and the node was
      handed no hostname at all. Kea has always sent one.

22 -- The PXE short lease existed on ISC in name only. dhcpd applies
      min-lease-time after max-lease-time, so the cluster default won and
      a pool address taken by a PXE ROM was held for half a day. The
      class now names all three bounds, and Kea is given the same 600
      seconds through xcat-pxe-lease.

24 -- ISC writes "authoritative;" into every generated subnet; Kea
      defaults to the opposite. A node that moved rack asked to keep an
      address from the network it had left and was answered with silence,
      so it waited out a lease that would never be renewed instead of
      being told to start over.

Also makes $xCAT_plugin::dhcp::callback a package variable. It was a file
lexical, so the tests' local() set something nothing read, and warnings
were only ever captured because an earlier process_request had left its
collector behind.
2026-09-10 13:38:32 -03:00
Daniel Hilst e7f7717627 fix(dhcp): give both backends the same answer for eight kinds of client
Eight of the parity decisions in the spec's Appendix A, where ISC dhcpd and
Kea answered the same frame differently and the operator never chose which
backend they got.

ISC gains two branches its if/else chain never had, so the client falls
through to /yaboot no longer:

  - 0x000c ppc64 is given /boot/grub2/grub2.ppc.  yaboot is not a UEFI
    loader and cannot boot one of these machines (decision 2).
  - 0x0010 is the same x86-64 UEFI firmware and the same loader as 0x0007,
    announced by a machine set to fetch it over HTTP (decision 3).

Kea gains what ISC has always had:

  - Etherboot, which predates option 93 and says what it is in option 60
    alone, is recognised and given the BIOS loader (decision 10).
  - onie_vendor is answered per subnet, not only per node: a switch
    announces it on its first boot, before anyone has defined it as a node
    (decision 11).
  - A client nothing else recognises is given /yaboot.  Kea has no else, so
    the condition is the negation of every architecture and vendor class
    another rule answers, rather than a dependence on class ordering
    (decision 7).
  - netboot=nimol is given /vios/nodes/<node> (decision 6).

and drops two answers ISC never gave:

  - No BIOS loader on disk no longer means pxelinux.0 in its place.  Naming
    a file that is not there costs the client a timeout it cannot diagnose,
    and a different loader boots something nobody asked for; the class is
    simply not written (decision 8).
  - netboot=petitboot sends the conf-file and nothing else.  petitboot acts
    on a boot file name when it sees one, so naming one as well sent the
    machine after a TFTP fetch that never happens on ISC (decision 12).
2026-09-10 13:27:58 -03:00
Daniel Hilst 72bae42096 test(dhcptest): write "no boot file" as absent, which is what a reply carries
Four scenarios asserted `bootfile matches ^$` for the cases where the point
is that the client is handed nothing to fetch: OPAL-v3 and s390x, which get
a conf-file instead, a petitboot node, and the architecture whose loader was
taken off disk.

That assertion cannot pass. A server naming no boot file sends an empty
BOOTP file header and no option 67, and an empty value is absent -- so the
run reported "bootfile is not present in the reply" against both backends,
in every case, whatever the server had done. Four failures that said nothing
about either implementation.

`bootfile absent` is the assertion those scenarios wanted. A unit test pins
it, so the unsatisfiable form cannot come back unnoticed.
2026-09-10 13:16:13 -03:00
Daniel Hilst 2abcec617b test(dhcptest): assert the spec scenarios that had no wire case at all
The previous commit deleted the scenarios nothing ran. This one covers the
other half of the same review: places where spec.md states a behaviour and
nothing on the wire ever checked it.

Eight scenarios, each with the fixture work it needs:

  S-36/37/38  next-server follows a node's own tftpserver, then its
              xcatmaster, then the subnet -- three nodes, three sources,
              three different answers, which is the only arrangement in
              which a server that ignores the node attributes can be told
              apart from one that honours them
  S-08        a node whose mac attribute names two ports is reserved on
              both, at the two addresses its two hostnames resolve to
  S-33/34     a diskless node is told its iSCSI target: option 17 for a
              normal client, the ISAN vendor space for one announcing that
              vendor class
  S-12        an architecture whose loader is not on disk is served an
              address and no boot file, with a second architecture whose
              loader is still there as the control
  S-14        a URL names the port the web server actually listens on
  S-05        a dynamic range written as a CIDR block serves out of it
  S-59        a withdrawn reservation stops being handed out
  S-54        a client that speaks BOOTP and not DHCP is still served

Three of these needed the tool to grow: BOOTP has no option 53, so a
BOOTREQUEST step type and a BOOTREPLY message type; option 119 is a name
list that may be compressed against itself, so RFC 3397 decoding, with a
fall back to hex on anything malformed rather than a misreading to assert
against.

RELEASE exposed a state-machine bug on the way. A step that expects no
reply left the client where it was, which is right for a DISCOVER nobody
answered and wrong for a RELEASE, which is never answered and whose whole
point is giving the lease up. A scenario could not boot a node, release
and boot it again -- which is exactly how S-02 has to be asserted, since
a session lasts one scenario and a second boot cannot reference the first
across that boundary.

Every new conf is reachable: three cases0 entries drive the eight fixture
subcommands, so none of them is dead on arrival.
2026-09-10 13:11:51 -03:00
Daniel Hilst 4a8c2fb147 test(dhcptest): delete the scenarios nothing ran, and assert the ones that stayed
node-specific-second-stage.conf and no-reply.conf were never invoked by any
dhcpfixture.sh subcommand or any case in cases0. A scenario file that nothing
runs is not coverage; it is a claim about the server that no CI failure will
ever contradict. Both are gone, along with the README rows and the header
comments in static-vs-dynamic.conf that pointed readers at them.

What node-specific-second-stage.conf described is real behaviour, so it is
asserted where it actually runs: netboot-methods.conf already defines a node
whose netboot method is xnba, and now asserts that the same node announcing
user class xNBA is handed .../xcat/xnba/nodes/<node> -- in both encodings of
option 77. That is S-25, and it makes the neighbouring first-stage scenario
S-26 as well: the second-stage rule must not change what the firmware sees.

The chainload scenarios asserted only that the second stage differed from the
first. An empty boot file satisfies that, and so does a wrong URL, so both
backends could pass while handing the loader something it cannot run. They now
assert the exact per-network script URL, and a UEFI pair is added for S-23,
where the URL takes a .uefi suffix that a server keying on the user class
alone will not produce. hierarchy-dhcpserver.conf had `bootfile !=` with
nothing on the right-hand side; it asserts the node's own loader instead.

Three header comments still said the backends were free to differ over what an
unknown machine is told to boot. The specification settled that in S-56 -- the
answer has to come from the subnet, on both backends, or a machine cannot
reach the state where anyone could define it -- so the comments now say so.

Also drops the riscv64http branch of arch_loader, which had no caller left
once the HTTP-boot scenario began asserting on the TFTP loader's name as a
substring of the URL.
2026-09-10 12:56:35 -03:00
Daniel Hilst a467f01afb test(dhcptest): pin a PXE client's lease at the 600 seconds the spec names
Appendix A decision 22. xCAT writes `class "pxe" { max-lease-time 600; }`
into the ISC config and `min-lease-time` equal to the cluster default into the
subnet around it, so what firmware is actually granted is not stated anywhere
and has never been checked. Kea has no counterpart at all.

Asserting the number rather than "shorter than the default" is the point: a
lease time nobody wrote down is a lease time nobody can tell has regressed.
2026-09-10 12:32:38 -03:00
Daniel Hilst e3568cf061 test(dhcptest): state one behaviour per scenario and assert it on both backends
The specification had a scenario headed "Known asymmetries between the
backends" and several clauses of the form "on Kea ... on ISC ...". A node
does not choose its management node's DHCP backend -- Backend.pm picks kea
on Ubuntu >= 22.04 and EL >= 10 and isc below -- so every one of those
clauses describes a machine that boots on one distro and hangs on the next,
written down as though it were a requirement.

Each is now decided one way, per the review in VersatusHPC/xcat-internal#175,
and the losing side is recorded as work rather than as behaviour. Appendix A
lists all 25 decisions with the reason and the side that has to change;
appendix B summarises what moved in the document itself. Every scenario also
carries a review number, @S-nn, so an issue can name one behaviour instead of
a paragraph.

The suite follows the specification rather than either implementation:

  pxe-arch-matrix.conf grows ia64, ppc64 UEFI, x86-64 UEFI HTTP boot, OPAL-v3,
  QEMU s390x, Etherboot by vendor class alone, an ONIE switch that is not yet
  a node, and a client that announces nothing at all. It also asks for options
  12, 114 and 209, which ISC sends only to a client that asked.

  netboot-methods.conf is new: one defined node per netboot method, each with
  its own MAC, plus the ScaleMP vendor class, the node's own name in option 12,
  and an interface the operator marked *NOIP* that must not be answered.

  dhcpfixture.sh defines that node matrix, supplies the loaders and URLs, and
  no longer branches on the backend when deciding what an unknown machine is
  told to boot.

The cases are expected to fail until the implementations are brought into
line: the failures are the divergence list, on the wire.
2026-09-10 12:30:48 -03:00
Daniel Hilst f5e3f7775b test(xCAT-test): run the DHCP on-wire cases as their own phase after ci_test
The CI run is now three phases in order: the perl and bash unit tests,
the ci_test set, then the DHCP on-wire cases.

The wire cases carry dhcp_wire and deliberately not ci_test. They are the
only cases that do not leave the machine alone while they run -- a veth
pair appears, site.dhcpinterfaces and site.dhcpbackend change, the DHCP
daemon is stopped and started -- and although each one restores all of it
from a trap, a ci_test case running in between would be sharing a
management node that is mid-reconfiguration. Holding them out of the set
is also what stops them being run twice.

run_dhcp_wire_test runs the whole set once per backend installed: against
ISC, then against Kea, each pass bracketed by `dhcpfixture.sh
backend-setup` and `backend-teardown`. It reports separately from the
fast regression test and counts case runs rather than cases, since each
case is run once per backend and each failure carries the backend name.

The phase writes its own xcattest configuration without a dhcpbackend
key: xcattest applies [Table_site] before every case, so the regression
conf's `dhcpbackend=isc` would have reset the backend case by case and
left the run silently testing ISC twice.
2026-09-10 11:58:20 -03:00
Daniel Hilst 1df2861139 test(xCAT-test): assert one boot file per architecture, whatever the backend
arch_loader branched on the running backend, so each backend was measured
against its own behaviour and the two could disagree for ever with the
case still green. That is the drift the case exists to catch.

It now returns one loader per architecture. The backends do read one
thing off the machine before deciding -- dhcp.pm:3469 keys the BIOS and
UEFI classes on xnba.kpxe and xnba.efi being unpacked under tftpdir, and
BootPolicy.pm:84 keys the riscv64 HTTP class on grub2.riscv64 -- while
ISC names all three unconditionally. So setup places the three loaders it
finds missing and teardown removes exactly those, and both backends are
asked the same question. An empty file is enough: dhcptest asserts the
name in the reply and never opens a TFTP session.

Without that, a parity failure would only mean the runner had not
unpacked a loader, which is a fact about the runner and not about xCAT.

The chainload case loses a guard that can no longer fire, and run-arch
loses the branch that skipped an architecture the backend "does not
serve".
2026-09-10 11:58:12 -03:00
Daniel Hilst ba469ecfe7 test(xCAT-test): run the DHCP wire cases against both backends in CI
Which DHCP backend a management node runs is an implementation default of
its distro -- kea on the newer releases, isc on the older ones -- so a
node booting on the same network has to be told the same things either
way. Until now CI exercised whichever backend the runner happened to
configure, which proves half of that and hides every drift between the
two.

The wire cases now carry a dhcp_wire label, and the CI driver holds them
back from the main pass: it asks dhcpfixture.sh which backends are
installed, then runs the whole set against one and again against the
other. Each pass is bracketed by backend-setup, which points
site.dhcpbackend at the backend and stops the other daemon -- two servers
on one wire both answer the same DISCOVER -- and backend-teardown, which
restores the site table and restarts what was running before.

Running every case under one backend and then every case under the other,
rather than switching inside each case, reconfigures the daemon once per
pass instead of once per case, and gives each failure in the summary the
backend name it belongs to. The summary counts case runs rather than
cases, since a wire case is run once per backend.

The check gate is relaxed from [ -x ] to [ -f ]: the fixture invokes
dhcptest as `python3 src/dhcptest`, so the execute bit only matters to
someone running it directly, and testing it made every wire case skip on
Debian.

dhcptest_backend_switch is dropped, along with the fixture's backend and
alt-backend actions: with every case running under both backends there is
nothing left for a case that switches one.
2026-09-10 11:43:19 -03:00
Daniel Hilst 7dd3f418a2 fix(xCAT-test): ship dhcptest executable so the wire cases actually run
src/dhcptest was mode 100644 in git. xCAT-test.spec chmods it to 755 on
install, but the Debian package ships it through dh_install, which keeps
the source mode -- so on Ubuntu it arrived non-executable.

dhcpfixture.sh's check gate tested [ -x ], so every wire case skipped and
exited 0. The CI runner is Ubuntu, which means no dhcptest case has ever
exchanged a packet there: each one took about a second, which is less
than the fixture's setup alone, and passed.

Setting the bit in git makes both packages ship it the same way. The
gate itself is relaxed to [ -f ] in the following commit.
2026-09-10 11:43:09 -03:00
Daniel Hilst d77322289b fix(dhcp): split the per-node xNBA UEFI branch dhcpd could not parse
The per-node statement for a UEFI node running netboot=xnba matched its two
architecture ids with a parenthesised alternation:

    else if <user class> and (option client-architecture = 00:09
                              or option client-architecture = 00:07) { ... }

ISC dhcpd has no parenthesised grouping in its expression grammar, so that
condition cannot parse. The node-level second stage it guards -- the
per-node .uefi script that keeps two machines chainloading at the same
moment from running the same script -- has never been reachable.

Written as two branches instead, one per architecture id, which is how the
per-network chain in BootPolicy.pm already spells the same test.

dhcp_isc_expression_grouping.t is a guard rather than a test of this branch:
the per-node statements are built inline in addnode against a live database
and cannot be called from a unit test, so it scans the plugin for the one
construct that produces the failure and checks the rendered per-network
chain and the user class test for it as well. dhcpd's tokens only -- Perl's
own parenthesised `exists` and `not` are excluded, and a bareword `option`
after an opening paren cannot be Perl.
2026-09-10 11:13:27 -03:00
Daniel Hilst 6ecd1994fb fix(dhcp): match the xNBA user class without parenthesised grouping
isc_xnba_user_class_test wrapped its two alternatives in parentheses so the
trailing `and option client-architecture = ...` would bind to the whole
alternation rather than to the second branch alone. ISC dhcpd has no
parenthesised grouping in its expression grammar, so every generated
dhcpd.conf became unparseable:

    /etc/dhcp/.../dhcpd.conf line 11: left brace expected.
        if (option user-class-identifier = "xNBA" or substring(option user-cl
    ^

followed by a cascade of "expecting a parameter or declaration" at each
`} else`, and a daemon that will not start at all.

suffix() gives the same coverage in a single expression that needs no
grouping: it returns the last four bytes, which are "xNBA" whether the
client sent the bare string or the RFC 3004 length-prefixed "\x04xNBA".

The unit test now pins the whole expression rather than each alternative,
and asserts the two properties that broke the daemon -- no leading paren,
no `or` -- so a future rewrite cannot reintroduce grouping unnoticed.
2026-09-10 11:10:47 -03:00
Daniel Hilst 19b695b2ee docs(dhcptest): point the architecture table at the section it means
The examples table said "see backend parity below" and nothing below is called
that. The asymmetries are enumerated under "Feature: The two backends behave
the same on the wire", so name that heading and its scenario instead.
2026-09-10 10:16:11 -03:00
Daniel Hilst e57d7ae0b1 test(dhcptest): close the gaps between the spec and what CI actually runs
Measured against spec.md, the wire cases asserted one boot file, for one
architecture, on one backend. Seven of the eleven shipped .conf files were
executed by nothing at all -- only parsed by `dhcptest validate` -- and the CI
run pins the backend to isc, which skipped the one discovery assertion as well.
A green run therefore said very little about whether a node can boot.

What the fixture now drives:

- One DISCOVER per client architecture (option 93), asserting the loader each
  is handed. This is where the plugin branches most and where a mistake is
  most expensive. The backends genuinely differ in which architectures they
  will answer for at all -- Kea emits no UEFI or HTTP-boot class unless the
  loader is present -- so arch_loader holds that policy in one place and the
  run skips by name rather than asserting a filename nothing configured.
- Both encodings of the user class, against the fix in the previous commit.
- The lease itself: the four-way handshake, renewal, rebinding, and a DHCPNAK
  for an address off this network. A node moved between racks comes back
  asking for the address it still holds; a server that stays silent leaves it
  retrying forever, which looks exactly like a node that will not boot.
- Discovery end to end: a machine nobody has heard of takes a pool address,
  is defined, and from the next DISCOVER is served its own address -- with the
  daemon's pid asserted unchanged across the adoption. xCAT injects ISC
  reservations over OMAPI precisely so adopting one node does not drop every
  other node part-way through discovery, and a regression to rewriting
  dhcpd.conf and restarting would pass every config-level test in this tree.
- The options a deployment actually needs: gateway, resolver, domain, MTU,
  log server and lease time. An address on its own does not install an
  operating system; it fails later and less obviously. The fixture's network
  now sets an MTU so there is something to assert.
- The hierarchy case, which was mn_only, now runs in CI too.

Option 12 is deliberately not asserted: ISC emits `send host-name` through
OMAPI and only rewrites it to `option host-name` on the static host fallback
path, so it is not reliably on the wire. spec.md records that, along with what
this specification does not cover -- DHCPv6, relay agents -- so a green run is
not read as more than it is.
2026-09-10 10:03:46 -03:00
Daniel Hilst 50ebc0af43 fix(dhcp): accept both encodings of the xNBA user class on ISC
A chainloaded second stage announces itself in the user class, option 77, and
must be handed a script URL rather than the loader it just ran -- otherwise it
chainloads itself forever and the machine never finishes booting.

RFC 3004 length-prefixes each string in option 77. Plenty of firmware sends the
bare string instead, and the same loader sends either depending on how it was
built, so both encodings have to be recognised. The Kea policy has always
accepted both. The ISC side compared only `option user-class-identifier =
"xNBA"`, which is the bare form, so a conforming loader booting against an ISC
management node loops -- and does so silently, since from the server's side
every exchange looks like a normal first-stage boot.

Put the test in one place, isc_xnba_user_class_test, and use it from both the
per-network architecture chain and the three per-node statements. The per-node
ones reach dhcpd through omshell, so the helper takes the quoting its caller
needs. `substring(option user-class-identifier, 1, 4)` skips the length byte;
option 77 is declared as a plain string, and the same construct is already used
by the onie_vendor branch a few lines below.
2026-09-10 10:03:31 -03:00
Daniel Hilst fc5edde73c fix(dhcp): decide the Debian interface variables from the version given
debian_sysconfig_interface_keys fell back to querying the local dpkg when it
was handed no version. The caller always passes one, so the fallback bought
nothing -- and it made the "version could not be established" case answer
differently depending on whether the machine running the code happened to have
isc-dhcp-server installed. It passed on a developer's EL workstation and failed
on the Ubuntu CI runner, which is the failure it was always going to produce.

An undefined version now means what the caller means by it: could not be
established, so write every spelling. Covered by a test that stubs the lookup
to prove the answer no longer depends on the machine.
2026-09-10 09:46:34 -03:00
Daniel Hilst 58d8aa45eb docs(dhcptest): specify the DHCP behaviour a booting node observes
Write down, as scenarios, what xCAT's DHCP server is supposed to do from the
point of view of a machine booting on a provisioning network: which address it
gets, what it is told to boot, where it fetches the loader from, and what else
the reply has to carry for a deployment to complete.

Read out of the source rather than from the documentation, and cited back to it
line by line, so a scenario that stops matching xCAT is a bug in one of the two
and it is clear which.

Scenarios that cannot be seen from the wire -- a generated file, a daemon's
command line -- are marked [config] and belong to the Perl unit tests instead.
The rest defines the domain the dhcptest wire cases are measured against,
including the ISC/Kea asymmetries the source shows but nothing yet asserts.
2026-09-10 09:45:41 -03:00
Daniel Hilst 4bf66d7c0b fix(dhcp): start dhcpd on the interfaces xCAT serves on Debian
makedhcp wrote the interface list into /etc/default/isc-dhcp-server as

    INTERFACES="..."

which is no longer always the variable the daemon is started with. The
systemd unit sources that file and expands exactly one variable onto dhcpd's
command line, and which one changed with the package:

    14.04  4.2.4-7ubuntu12      sysvinit only   $INTERFACES
    16.04  4.3.3-5ubuntu12      unit            $INTERFACES
    18.04  4.3.5-3ubuntu7       unit            $INTERFACES
    20.04  4.4.1-2.1ubuntu5     unit            $INTERFACES
    22.04  4.4.1-2.3ubuntu2     unit            $INTERFACESv4
    24.04  4.4.3-P1-4ubuntu2    unit            $INTERFACESv4
    26.04  4.4.3-P1-4ubuntu2    unit            $INTERFACESv4

with the matching v6 unit reading $INTERFACES before the change and
$INTERFACESv6 after it. Note the boundary is not an upstream ISC release:
20.04 and 22.04 both ship upstream 4.4.1 and are told apart only by the
Debian revision, so it has to be decided on the whole package version.

The sysvinit script does copy INTERFACES into INTERFACESv4 when the latter is
empty, but the unit is what starts the daemon on any of these releases and it
has no such bridge. So from 22.04 on, the unit ran

    exec dhcpd -user dhcpd -group dhcpd -f -4 ... -cf $CONFIG_FILE $INTERFACESv4

against a variable xCAT never set. An unset variable expands to nothing, so
dhcpd was launched with no interface argument at all. It does not fail for
that: it binds every interface it can find and only warns about the ones with
no subnet declaration. site.dhcpinterfaces and servicenode.dhcpinterfaces
were therefore silently inert -- the provisioning NIC was served because
makedhcp had also emitted a subnet stanza for it, not because anything
honoured the setting, and any other interface on that subnet was served
alongside it.

The line match compounded it. m/^$dhcpd_key/ is not anchored on the
assignment, and INTERFACES is a prefix of both variables the package ships.
Since 18.04 the postinst seeds only INTERFACESv4="" and INTERFACESv6="" --
no INTERFACES line exists to match -- so the rewrite claimed those two lines
instead and overwrote both, discarding whatever debconf or the administrator
had put there and leaving a file with two INTERFACES assignments and nothing
either unit reads.

Pick the variables through a dispatcher keyed on the installed
isc-dhcp-server version, so both behaviours are served rather than one being
traded for the other, and anchor the match on the '=' so a key cannot claim a
line it is merely a prefix of. When the version cannot be established the
dispatcher writes every spelling, since an unset variable is the outcome that
leaves dhcpd bound to everything. Duplicate assignments left behind by the
old writer are collapsed to one, which repairs a file already damaged on an
upgraded management node.

The EL and SLES paths keep their single DHCPDARGS / DHCPD_INTERFACE /
DHCPD6_INTERFACE key and are unchanged.
2026-09-10 09:31:51 -03:00
Daniel Hilst 33bf641a72 test(dhcp): capture makedhcp leaving dhcpd unrestricted on Debian
On Debian and Ubuntu, makedhcp writes the list of interfaces it is serving
into /etc/default/isc-dhcp-server using the key set at dhcp.pm:2296:

    $dhcpd_key = "INTERFACES";

That key stopped being the one the daemon is started with. The variable the
isc-dhcp-server systemd unit expands onto dhcpd's command line changed with
the package; verified by unpacking the archive's own debs:

    trusty  4.2.4-7ubuntu12     sysvinit only   INTERFACES
    xenial  4.3.3-5ubuntu12     unit            $INTERFACES
    bionic  4.3.5-3ubuntu7      unit            $INTERFACES
    focal   4.4.1-2.1ubuntu5    unit            $INTERFACES
    jammy   4.4.1-2.3ubuntu2    unit            $INTERFACESv4
    noble   4.4.3-P1-4ubuntu2   unit            $INTERFACESv4 (+ a v6 unit
                                                reading $INTERFACESv6)

and every package from bionic onward ships a default file whose only
interface variables are INTERFACESv4 and INTERFACESv6, both seeded empty by
postinst. No INTERFACES line ships at all.

Two things go wrong on jammy and later.

The unit runs `exec dhcpd ... -cf $CONFIG_FILE $INTERFACESv4`. That variable
is never set, so it expands to nothing and dhcpd is launched with no
interface argument. dhcpd does not fail for this: it binds every interface
it can find and merely warns about the ones with no subnet declaration. So
site.dhcpinterfaces and servicenode.dhcpinterfaces are silently inert. The
provisioning NIC is served only because makedhcp also emitted a subnet
stanza for it, not because anything honoured the setting, and any other
interface on that same subnet -- a bridge port, a bond member, a second NIC
-- is served too.

The line match is m/^$dhcpd_key/, which is not anchored on the assignment.
"INTERFACES" is a prefix of both variables the package ships, so the rewrite
claims the INTERFACESv4 and INTERFACESv6 lines and overwrites both. The file
is left holding two INTERFACES lines and nothing the units read, discarding
whatever debconf or the administrator had set there.

Extract the rewrite into _sysconfig_interfaces_content() with no change in
behaviour, and add a unit test that asserts the effect rather than the
spelling: it writes the produced file to a temp path and sources it with sh
exactly as the unit does, then checks what would land on dhcpd's command
line. The test is red against the current key and stays red until the writer
sets the variables the daemon is actually started with.
2026-09-10 09:17:38 -03:00
Daniel Hilst 349204f783 fix(dhcptest): make the wire cases runnable from an installed package
Two things would have kept the wire cases from doing anything on a management
node built from the deb.

src/dhcptest was committed without its executable bit. The spec chmods it at
install time, but dh_install copies the source mode straight through, so on a
deb-installed node the fixture's `[ -x ]` check failed and every wire case
stood down with "dhcptest is not installed" -- passing while testing nothing,
which is the exact failure mode the fixture exists to avoid.

Teardown restarted the backend that was running before the fixture started but
left the one the backend-switch case switched to still running. Two servers on
one network answer the same DISCOVER, so the next case to look at that network
would see an offer from a daemon nobody asked for.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-10 08:54:14 -03:00
Daniel Hilst f2112dcafa test(dhcptest): drive xCAT-generated DHCP config from the wire
dhcptest could already assert on what a server puts on the wire, but on a
single-node management node -- the CI runner above all -- there was nothing
for it to talk to, so a wire case would have passed while proving nothing.

dhcpfixture.sh builds the missing network out of a veth pair: the server end
carries the management address and is the only interface named in
site.dhcpinterfaces, the client end has no address, which is the state a
provisioning NIC is in when a machine boots on it. It then defines a network
with a dynamic range and one node with a MAC, runs makedhcp so the
configuration comes from xCAT rather than from hand-editing, and restores
everything from a trap. All the xCAT knowledge lives there, next to the
xcattest cases, so dhcptest itself stays agnostic to it.

The new scenarios assert what a provisioning server has to get right:

  provision-vs-discovery.conf    a known machine gets its reservation and its
                                 own loader, an unknown one gets a pool address
  discovery-bootfile.conf        an unknown machine is told what to boot too
  node-specific-second-stage.conf  the chainloaded stage names its host
  hierarchy-dhcpserver.conf      a subnet whose pool was handed to another
                                 server ignores unknown MACs but still points
                                 known ones at it

Those are four files rather than four scenarios in one because they describe
server configurations that contradict each other, and because --set values are
resolved when a file is loaded, before -s selects anything: an unused
scenario's missing variable would otherwise abort a run that never intended to
use it. ISC dhcpd leaving the unknown-client boot file to the per-host blocks
while Kea puts it on the subnet is exactly such a contradiction, which is why
discovery-bootfile.conf is opt-in and the fixture only asks for it on Kea.

dhcptest_backend_switch reruns the same files after switching site.dhcpbackend
and regenerating, so the backend-agnosticism the tool exists to demonstrate is
itself under test rather than asserted in a comment.

The `in` operator now accepts an inclusive first-last range as well as a CIDR
and a list: a dynamic pool is written as two addresses because it rarely lines
up on a prefix boundary.

Verified against a real dnsmasq in a network namespace: 17/17, 3/3 and 5/5
assertions green, a deliberately wrong expectation failing with expected and
received on one screen, and host networking untouched.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-10 08:51:10 -03:00
Daniel Hilst 008f83d2e4 test(xCAT-test): add dhcptest, a wire-level DHCP client test tool
xCAT's DHCP behaviour is tested from two directions today, neither of which
touches the wire: xCAT-test/unit/dhcp_*.t assert on generated config text, and
xCAT-test/integration/dhcp_{isc,kea}_config_validation.t feed that config to a
real daemon and check it parses. Both stop at "the server accepted our
config". Nothing verifies that a client sending option 93 = 0x000b actually
gets an aarch64 loader back.

The one existing wire tool, xCAT-probe/subcmds/detect_dhcpd, hand-packs a
fixed DISCOVER carrying no option 60, 77 or 93 (so it never reaches any arch
branch), parses replies by regexing tcpdump output, and binds an
IO::Socket::INET to a local address, so it cannot run on a provisioning NIC
that has no address yet.

dhcptest fills that gap. It is a Python 3 tool using only the standard library
plus Scapy, speaking raw Layer 2, driving DHCP transactions declared in INI
files and asserting on what came back:

  dhcptest run      -i eth1 --set net=10.0.0.0/24 conf/full-lease.conf
  dhcptest validate conf/*.conf      # no root, no network, no scapy
  dhcptest list     conf/*.conf
  dhcptest discover -i eth1          # ad-hoc, no conf file

Scenarios cover DISCOVER/OFFER, the full lease, renew and rebind, reserved
versus pooled addresses, deliberate silence, the PXE architecture matrix and
the iPXE user-class split. Assertions are a small DSL -- target, operator,
value -- over message type, BOOTP header fields, options by number or name,
subnet membership, and the boot file a client would actually use (option 67
or the header, since servers differ on which they fill).

Two boundaries are enforced by unit test rather than by convention:

  - The tool is implementation-agnostic. It exchanges packets and checks
    fields; it never reads the xCAT database, never runs an xCAT command, and
    names no DHCP implementation. Anything that differs between servers is
    expressed by whoever writes the .conf.
  - Only runner.py imports scapy, so validate and list stay usable in CI on a
    host with neither scapy nor root.

Host networking is never modified. The client MAC defaults to a synthetic
locally-administered address, everything goes over a raw L2 socket, no leased
address is ever configured, and the ARP responder answers only for addresses
the session was actually granted.

Variables in .conf files are configparser's own BasicInterpolation, %(name)s,
fed from [vars] and --set. $offer.address and friends are resolved separately
at step-execution time, because they name a reply that has not arrived when
the file is read; BasicInterpolation gives $ no meaning, so the two coexist
without escaping.

Packaging and CI:

  - xCAT-test.spec installs dhcptest alongside unit/ and integration/, and
    only recommends python3-scapy: it lives in EPEL on EL, and a hard Requires
    would make xCAT-test uninstallable on a management node without EPEL.
    The deb depends on it outright.
  - github_action_xcat_test.pl runs the unit tests and `dhcptest validate`
    last, after the fast regression.
  - autotest/testcase/dhcptest/cases0 carries the two offline cases plus a
    wire case that stands down with an explanation when no provisioning NIC
    is available.

Verified against a real DHCP server in a network namespace: 34 lease, renew,
rebind, reservation and silence assertions and 22 PXE/iPXE assertions all
pass, and host networking is untouched before and after.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-10 08:19:41 -03:00
Vinícius Ferrão 17126c74a9 Merge pull request #7823 from VersatusHPC/feat/bats-shell-tests
test(xcat-core): Introduce BATS & convert shell scripting tests to it
2026-09-09 19:48:49 -03:00
Daniel Hilst 3cade9651c Merge pull request #7821 from VersatusHPC/feature/ubuntu-riscv64
feat(ubuntu): support riscv64 management and provisioned nodes on 24.04 and 26.04
2026-09-09 10:50:07 -03:00
Vinícius Ferrão 166da7c1b6 test(xCAT-test): cover the install architectures debian.pm accepts
The plugin is loaded and install_darch is called for the architectures
xCAT installs Ubuntu on and two it does not. Against the previous plugin the
test fails on the missing function: the mapping was inline in mkinstall.

Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
2026-09-08 22:01:34 -03:00
Vinícius Ferrão a6e69e88a4 fix(debian): stop reporting riscv64 as an unknown install architecture
mkinstall mapped x86_64 and x86 to their Debian names and accepted ppc64le
and ppc64el. Every other architecture, riscv64 included, was logged as
"Unknown arch" on each diskful install, although the install went on with
the name unchanged, which is right for riscv64.

Move the mapping into install_darch, which takes the Debian name from
xCAT::Utils::debian_arch and knows the architectures xCAT installs Ubuntu
on. riscv64 is one of them.

Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
2026-09-08 22:01:34 -03:00
Vinícius Ferrão 5a541814df test(xCAT-test): pin what a riscv64 management node installs and serves
The grub-common assertion fails against the previous metapackage. The rest pins
what must NOT change: a riscv64 management node still recommends the x86 boot
payload and the Genesis images of the other architectures, because it serves
them to the nodes it provisions.

Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
2026-09-08 22:01:33 -03:00
Vinícius Ferrão f0c7803595 test(xCAT-test): cover the declaration of the loader build tool
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
2026-09-08 22:01:33 -03:00
Vinícius Ferrão 53692323b4 fix(xCAT-server): declare the tool copycd builds the riscv64 loader with
copycd builds the riscv64 boot loader by running grub-mkimage, which
grub-common ships on every supported Ubuntu release. Nothing declared it, so a
management or service node installed without that package copies riscv64 media
and produces no loader, while DHCP keeps pointing every riscv64 node at the path
where the loader should be.

The declaration belongs to xcat-server, which carries the plugin that runs the
command, so both metapackages inherit it.

Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
2026-09-08 22:01:32 -03:00