2
0
mirror of https://github.com/xcat2/xcat-core.git synced 2026-10-06 17:46:55 +00:00
Commit Graph

28153 Commits

Author SHA1 Message Date
Daniel Hilst 4ccac9abae test(xcat-core): a failed extraction stops the whole suite, and the DHCP comments repeat
dhcp_ddns_policy.t and dhcp_isc_expression_grouping.t called BAIL_OUT
when a routine was missing or the plugin could not be read. prove stops
every remaining file on a bail-out, so one of them hides the results of
every test that would have run after it. die is just as loud and costs
only its own file.

Three facts in this branch were each written out in four places. That
SIGHUP reports success before Kea reads the file appears in Kea.pm twice
and in dhcp.pm again; that an installed node netboots when no boot file
reaches it appears in BootPolicy.pm and three times in dhcp.pm; the
dhcpd grouping rule appears in BootPolicy.pm and in the test header. Each
now stands once, without the symptom narration and the defect history
around it.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-14 08:44:38 -03:00
Daniel Hilst a4f970d9be test(dhcp): the DDNS scenarios in the spec have no assertion
specs/dhcp-wire.md gained a feature for the name the cluster's DNS learns --
S-74 to S-78. Four of the five had nothing asserting them, and dhcptest could
not send option 81 at all, so the scenario a client uses to claim a name could
not be written.

dhcptest gains a `fqdn` step field carrying option 81, with the RFC 4702 flags
written as a prefix (`S:node01`, `N:node01`, `E` for wire format). It gains two
targets: `fqdn_flags` for the flag letters, and `dns_name` for the name the
reply says the node will be known by -- option 81 when the server sent one,
option 12 otherwise, because Kea answers an option 81 request in option 81 and
ISC under `ignore client-updates` answers in option 12 only. That is the same
shape as `bootfile`.

netboot-methods.conf asserts S-75 twice, once for a name claimed in option 12
and once in option 81. dhcp_ddns_policy.t asserts S-76, S-77 and S-78 on both
backends.

Checked against kea-dhcp4 2.4.1 and isc-dhcpd 4.4.3-P1 on xcat26-mn: both
scenarios fail on a server with no per-node name and pass on one that has it.
Four mutations of the code under dhcp_ddns_policy.t are each caught.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 10:15:24 -03:00
Daniel Hilst 18d5c78121 refactor(dhcp): the DDNS zone block is written twice and the copies disagree
addnet and addnet6 each build the zone and key statements that let dhcpd
update the cluster's DNS. The two copies drifted: addnet writes "primary" only
when a server is known, addnet6 writes it unconditionally and emits
"primary ; key xcat_key;" for a network whose nameservers are unset. Neither
copy can be driven from a test, because both sit in the middle of a routine
that needs the whole plugin's globals.

isc_ddns_zone_statements takes the decision and returns the lines. The callers
keep the side effect and the _omapi_settings lookup. addnet6 gains addnet's
guard, and a network with neither a ddnsdomain nor a domain now writes no zone
at all rather than "zone . {".

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 10:13:49 -03:00
Daniel Hilst 5ca83b991f test(dhcp): name the spec scenario each hostname check asserts
S-35 in specs/dhcp-wire.md now requires the node's own name whatever the
client advertised, and new S-74 requires the same name in the lease and in the
DNS update. The two checks that assert them cite neither.

The Kea reservation check in dhcp_kea_plugin_intent.t cites S-74, which is
marked [config] because no client can observe the recorded name. The
netboot-methods.conf step cites S-35 and says which half it cannot see.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 09:56:43 -03:00
Daniel Hilst a2b4b5a690 fix(dhcp): a client naming itself takes over a Kea node's name
kea_node_reservations in xCAT-server/lib/xcat/plugins/dhcp.pm writes the node
name only as the reservation's host-name option. Kea reads the lease name and
the DDNS name from the reservation's "hostname" field, so with that field
absent it uses the name the client put in option 12. A node that calls itself
"ubuntu" -- which the Ubuntu installer does -- gets the DNS record for the
node's address and is told "ubuntu" back.

The reservation carries the hostname field again. The host-name option stays,
so a server that sends it verbatim still sends the short name.

The field is qualified by Kea 2.4 before it reaches option 12, so a node on a
cluster with DDNS is told "node01.cluster" where ISC says "node01". S-35 in
netboot-methods.conf asserts the label rather than the whole option for that
reason, and dhcp_kea_plugin_intent.t asserts the field is present.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 09:29:57 -03:00
Daniel Hilst 838d2f6870 test(dhcp): a client naming itself takes over a Kea node's name
A Kea host reservation written by kea_node_reservations carries the node's
name only as a host-name option. Kea takes the name for the lease and for the
DDNS update from the reservation's "hostname" field, and with that field
absent it takes the name the client advertised instead.

Measured with kea-dhcp4 2.4.1 on a veth pair, ddns-qualifying-suffix
"cluster.local.", a reservation for 192.0.2.50 and a client sending option 12
"otherdesk":

  hostname field absent   option 12 -> otherdesk.cluster.local
                          lease     -> otherdesk.cluster.local (fqdn fwd+rev)
  hostname field present  option 12 -> node01.cluster.local
                          lease     -> node01.cluster.local

dhcp_kea_plugin_intent.t now asserts the reservation carries the hostname
field, and netboot-methods.conf adds S-35 "a node advertising another name is
still told its own". The existing S-35 step matches the node's label rather
than the whole option, because a server that qualifies the name with the
cluster domain has still answered with the node's own name.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 09:25:55 -03:00
Daniel Hilst 9adbb6ae67 fix(provtest): a fetch is ok only when curl finished it
_decode in provtest_lib/httpc.py builds the reply's "ok" field from the HTTP
status alone. curl writes the status as soon as the response header arrives,
so a connection closed in the middle of the body, or a transfer killed by
--max-time, still produces status 200. The http.conf scenario that fetches the
kernel and initrd "byte for byte" then reads a partial file as served.

The reply already carries the reason curl gave. "ok" now also requires that
reason to be absent.

xCAT-test/provtest/tests/test_parsers.py asserts it for a body cut short
(curl exit 18) and for a timeout while reading the body.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 09:20:24 -03:00
Daniel Hilst 5212a75179 test(provtest): a download cut short passes the HTTP check
The provtest HTTP client reports success from the status line alone. curl
prints the status as soon as the response header arrives, so a server that
closed the connection in the middle of a kernel, or a transfer that ran out
of time, still leaves a reply that reads 200 and ok.

Two cases assert that a reply is not ok when curl did not finish: a body cut
short (curl exit 18) and a timeout while the body was being read. Both fail
before the fix.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-12 09:19:57 -03:00
Daniel Hilst cddf4f3bef docs(provtest): drop the pointer to the internal specification
The README named a file in a repository most readers of this one cannot
open. The clause tags on the scenarios already carry what the pointer was
for: a failure names the clause, and anyone holding the specification can
find it from the tag alone.
2026-09-11 13:08:19 -03:00
Daniel Hilst 69f53644d0 docs(provtest): correct the boot file's provenance and list what is not covered
The dhcp-named-file scenario said its file name came from an observed
DHCP acknowledgement. Nothing here speaks DHCP: the caller spells the
name, dhcptest asserts that xCAT sends it, and this asserts it can be
fetched.

The scenario table had fallen behind four files, and two real gaps were
unwritten anywhere: the x509cert form of getcredentials, whose CSR a
wire test cannot fabricate meaningfully, and the absence of copycds --
the install tree the nodes point at is fabricated.
2026-09-11 12:02:38 -03:00
Daniel Hilst 9c515a596d test(provtest): refuse a callback port any user on the node could bind
credentials.pm:138-143 signs nothing unless the callback port named in
the request is below 1024, because only root can bind one -- a high
port would let any user on a node collect the cluster's keys.

The suite asserted the closed-port half of that and not this one. The
new scenario puts a listener up on a high port and asks to be called
back there, so a callback arriving is one xcatd chose to make rather
than one the client failed to answer.
2026-09-11 12:02:33 -03:00
Daniel Hilst 5ac224bc25 test(provtest): name the values a node is given, not merely that it got some
P-55 chooses between four sources for imgserver and every one of them
held the same address, so the assertion could not say which won. The
node now has a tftpserver of its own on a second address, and imgserver
is asserted to be it and not site.master.

chain-advances asserted only that two answers differ; it now names the
state, boot, which nodeset leaves in chain.currchain (destiny.pm:550).

And the postscript is read for NODE, MASTER, SITEMASTER and DOMAIN --
four renderings from three sources -- and for any directive that
survived into the body.
2026-09-11 12:02:27 -03:00
Daniel Hilst 378ba7eba1 test(provtest): assert what the renderers produced, not that they ran
The kickstart template carried no directive, so mkinstall rendered it
correctly by doing nothing; it now carries a #TABLE: per node and an
#INCLUDE:, and the fetched autoinst is asserted to hold the expanded
values and no directive or #INCLUDEBAD:.

The postscripts listing is walked as a node walks it: a name taken out
of the index, then that leaf fetched -- Indexes and the alias are
separate directives, and a tree listed but unreadable fails every
postscript at once.

And the kernel and initrd the config names are fetched over both
transports and compared by sha256.
2026-09-11 12:02:14 -03:00
Daniel Hilst 1a7386e978 fix(dhcptest): remove every node teardown defined, not just the first
The two loops that undefine the netboot and extra nodes read their list on
standard input, and the xCAT clients called inside them read standard input
too: makedhcp swallowed the rest of the file on the first iteration, so the
loop ended after one name. A full run left fourteen node definitions on the
machine, and the state directory was deleted regardless, so nothing
recorded what had been left.

Both loops read on file descriptor 3, as the provtest teardown already
does. Each removal is checked and teardown keeps its state directory when
one fails. Verified on both backends: no dhcptest node survives a run.
2026-09-11 11:42:25 -03:00
Daniel Hilst 02f354d4bc test(ci): fail a phase whose cases produced no verdict
run_cases counts the Passed and Failed lines a run prints, so a case that
crashed before either one was counted nowhere and the phase reported
success over cases that never ran. Each phase now knows how many cases it
asked for and fails when the verdicts do not add up.

A failed query for the case list returned an empty list, which read as a
phase with nothing to do; it is a failure now. A backend whose setup failed
runs its teardown before the next one starts, rather than leaving the
machine configured for a backend that is no longer under test.
2026-09-11 11:26:01 -03:00
Daniel Hilst bcc6acab5f fix(provtest): build the grub2 loader on Debian too, and run without one
The CI runner is Ubuntu, where the program is grub-mkimage and not
grub2-mkimage, so the fixture built no loader; grub2.pm writes no
configuration at all for a node whose loader is missing, setup died on the
node every later stage needs, and all eight wire cases failed for that one
reason.

Both names are tried now, and the EFI and BIOS module trees under either
/usr/lib/grub or /usr/lib/grub2. Where none of that exists the node is set
with xnba instead, and the stages that read a grub2 config say which
assertions they left out. The workflow installs grub-common and
grub-efi-amd64-bin so the runner tests the grub2 path itself.
2026-09-11 11:25:55 -03:00
Daniel Hilst 0cb9b4735b test(provtest,dhcptest): fail a dirty machine instead of skipping it green
check answered only yes or no, so a machine holding an earlier run's nodes,
interface or namespace passed every wire case without running one -- and
went on doing that until someone looked. It now answers three ways: run,
cannot run and the case passes, or refuse, which fails the case. Leftovers
and a cluster with nodes of its own are refusals; PROVTEST_ALLOW_LIVE and
DHCPTEST_ALLOW_LIVE opt in.

The empty-cluster message from lsdef is no longer counted as a node. ss is
required rather than assumed, since without it every port probe answered
yes. Each skipped stage is counted and the total printed. The DNS removal
stage puts the name back and proves it resolves before asserting that it
does not.
2026-09-11 11:25:48 -03:00
Daniel Hilst ecbdd53e83 fix(provtest,dhcptest): make a failed teardown safe to fail
Teardown removed nodes, restored tables and unpacked the saved zone tree
without checking any of it, then deleted the state directory regardless.
A restore that half-failed left the machine wrong and the record of what
to put back gone.

Every removal and restore is checked now. The zone archive is verified
when written and listed again before the live tree is removed, so an
unreadable archive leaves the tree alone. Anything that fails keeps the
state directory and returns non-zero. The kickstarts mkinstall renders
under autoinst are removed by name, and the loader moved aside during the
absent-loader stage stays in tftpdir rather than under /tmp.
2026-09-11 11:25:25 -03:00
Daniel Hilst 68075b6b44 fix(provtest): give makedns a name it can resolve for this machine
makedns looks up the server of each zone it writes and qualifies it with
site.domain. On a cluster whose networks name their server as
<xcatmaster> that name is this machine's own, inside the fixture's
domain, which nothing resolves -- so makedns failed and every wire case
died in setup. It passed where the other networks name no server and the
lookup fell back to site.nameservers, an address.

Add the qualified name to /etc/hosts, which teardown already restores.
2026-09-11 10:40:38 -03:00
Daniel Hilst f207c33777 fix(provtest,dhcptest): refuse a second setup over a live fixture
Both fixtures record what they are about to change so teardown can put
it back, and both truncated that record on entry. A setup run while one
was already up therefore saved the fixture's own site table, dhcpd.conf
and /etc/hosts as the originals, and the veth and namespace it no longer
knew about survived the teardown that followed.

Refuse when the state directory is there, naming the teardown to run.
2026-09-11 10:40:17 -03:00
Daniel Hilst 1832572a20 fix(provtest): fetch the boot file DHCP actually names
P-70 asserts that the name DHCP sends can be fetched, but the fixture
passed the architecture's loader as that name. DHCP hands a grub2 node
the per-node name instead, so deleting it left the whole suite green
while a real node would TFTP a file that is not there.

Pass boot/grub2/grub2-<node>, which is what dhcptest asserts is sent.
The scenario stays out of the run where no loader is on disk, since the
per-node name is a link to it.
2026-09-11 10:40:02 -03:00
Daniel Hilst a47e63141d docs(dhcp,provtest): cut the comments back to what a reader needs
The branch's own comments had grown to explanations of the reasoning
behind each decision. Trim them to the fact and its consequence: why a
line exists, and what breaks without it. Long blocks come down to a
hundred words, most to twenty or fewer, and the file headers keep only
their usage tables and the invariant each suite rests on.

No code changes. Every unit test, scenario validation and syntax check
still passes.
2026-09-11 10:13:23 -03:00
Daniel Hilst 9cae1ea9f8 refactor(provtest): say each thing in the fixture once
Four repetitions, each of them a place for a later edit to be made in
one copy and not the others:

- the settings every scenario is given, retyped at nineteen calls
  beside a helper that built them and was never called;
- three unit-name lookups with identical bodies;
- the node reset two stages open with;
- the eight settings the three ordering scenarios share.

tftpdir and installdir are asked of xcatd once rather than nine times
in a stage, and the name encodings now call provtest's own netutil
rather than a second implementation in shell that no test covers.
2026-09-11 09:37:50 -03:00
Daniel Hilst 9d63035a4d docs(ci): say why neither wire set runs in the fast regression phase
The comment named only the DHCP cases; the provisioning ones are kept
out of that phase for the same kind of reason -- they add an
interface, nodes, a network and a rewritten zone.
2026-09-11 09:25:08 -03:00
Daniel Hilst e70166d694 docs(provtest): cite P-76 on the scenario that pins it
The nameless-request scenario was still labelled P-61, the scenario
that found it, rather than the one written for it.
2026-09-11 09:25:08 -03:00
Daniel Hilst f46261df90 test(provtest): assert the xnba kernel fetch only where xCAT serves it
A management node can end up with a web server that is not the one
xcatconfig configured -- a stock nginx holding port 80 in front of the
httpd whose xcat.conf carries the aliases.  It answers 404 to every
path a node is ever given, so P-22's kernel fetch failed in a way that
read like a broken boot script, and the whole HTTP stage would have
failed nine times for one reason.

The fixture now probes the aliased directories before asserting
against them.  P-21 and P-22 are separate scenarios, so the script is
still asserted where the kernel cannot be, and the HTTP stage skips
with the reason rather than failing.
2026-09-11 09:25:01 -03:00
Daniel Hilst 4d8d686747 fix(provtest): make the fixture's teardown finish what it started
Every loop over a list of names ran the xCAT clients on standard
input, so the first client read the rest of the file and the loop
ended after one name.  Teardown said it had removed the fixture while
five nodes, an osimage, a network and an unrestored site table were
still there, and the next setup recorded those leftovers as the state
to restore to.  The loops now read on descriptor 3.

`nodeset offline` is not enough either: it gives up when it cannot
reach the DHCP backend and never removes the staged kernel and initrd.
The artefacts the fixture's own nodes and images can produce are now
removed by name.

check refuses on any leftover node, not just the first.
2026-09-11 09:24:55 -03:00
Daniel Hilst a1bc164b5a ci: run the provtest checks after ci_test and before the DHCP wire test
The provisioning cases go where a node meets them: after the whole
ci_test set, so they are not reconfiguring the machine underneath it, and
before the DHCP phases.

run_provtest_unit_tests runs the offline half first -- the Python suite
and a parse of every scenario -- so a file that does not even load is
reported as that rather than as a case failure, and skips rather than
fails when python3 is absent. run_prov_wire_test then runs the prov_wire
cases once, not once per backend: nothing in them depends on which DHCP
server is installed, and what a node is told over DHCP is dhcptest's
subject.
2026-09-11 09:01:00 -03:00
Daniel Hilst 711c4870db test(provtest): add the autotest cases and ship the tree in the packages
Ten cases: two offline ones labelled ci_test, the unit suite and a parse
of every shipped scenario, and eight wire cases labelled prov_wire, one
per stage of the chain a node walks -- DNS, loader and config, install
tree, genesis and discovery artefacts, discovery requests, the xcatd
protocols, the install monitor, and the three failures that look alike
from a console.

The wire cases are deliberately not ci_test: while one runs the
management node has an extra veth pair and namespace, four nodes, a
rewritten zone and possibly a moved http port. Each restores all of it
from a trap, and each passes with a message when the host cannot host it.

provfixture.sh holds every xCAT command; the scenarios read none of it.
2026-09-11 09:00:54 -03:00
Daniel Hilst 2efcbcde08 test(provtest): assert what xCAT does rather than what the spec assumed
Running the scenarios against a live management node found nine places
where they asserted something xCAT does not do. Corrected here, each with
the reason in the file:

- install configs carry no xcatd= or destiny=; those are written for the
  genesis states only.
- #END OF SCRIPT is plain-transport framing; over TLS a reply ends with
  <serverdone>.
- getcredentials is named in <arg> and answered as <data><content/>
  <desc/></data>, so both fields are addressable now.
- currstate for an install is "install <osimage>", so the operator is
  starts-with.
- bootparams is empty for an ordinary install, so kernel and initrd
  cannot be asserted from getdestiny.
- a server never reached sends no <error>; assert the handshake instead.
- an NXDOMAIN answer is the expected result of a removal scenario.
- P-75 has to precede P-74: a withdrawn name cannot be nodeset again.

Also adds netutil.hex_net: a per-network config is named by as many hex
digits as the netmask covers, not by all eight.
2026-09-11 09:00:47 -03:00
Daniel Hilst 721bfe710e fix(dhcptest): find the dhcptest tree from a checkout three levels up
The fixture falls back to a checkout when the installed tree is absent,
and the relative path was one level short: from
autotest/testcase/dhcptest it has to climb to xCAT-test, not to
autotest. In a checkout the fallback therefore resolved to nothing and
every case skipped, silently and with a pass.
2026-09-11 09:00:40 -03:00
Daniel Hilst 5a969c8eb7 fix(credentials): refuse a getcredentials request naming no credential
Which credential is wanted is named in <arg>. A request without one
reached a dereference of an undefined array ref, so xcatd died inside the
plugin and shipped the Perl error back to the client instead of ignoring
the request.

Check for the element before the callback is made, and log and drop the
request the same way a bad callback port is dropped. A client that sends
a malformed request now learns nothing about the server, which is the
point: this path hands out cluster certificates to anything whose address
resolves.
2026-09-11 09:00:35 -03:00
Daniel Hilst 9f03656318 docs(test): move the wire specifications out of the test tree
Both spec.md files are design documents for this organisation: they argue
about what xCAT ought to put on the wire, cite the plugin lines that decide
it, and record where the original proposal was wrong. Shipped under
xCAT-test/ they would land in an RPM on every management node and be offered
upstream as part of a test directory, which is not what they are for.

They now live in the internal repository as specs/dhcp-wire.md and
specs/provision-chain.md. The suites keep their clause tags, so a failing
case still names the clause it belongs to.
2026-09-11 08:09:54 -03:00
Daniel Hilst 86533527e0 test(provtest): unit-test the offline half, and the rules it relies on
154 tests, standard library only, no network and no root: the config parser,
the assertion language, $step.field resolution, the offline machine, TAP
output, address arithmetic, and the decoding of captured dig, curl and tftp
output -- including an NXDOMAIN whose only record is an SOA, and a TFTP fetch
that created an empty file before finding the server had nothing to send.

test_boundaries.py asserts rules about this code rather than about xCAT: only
proc.py starts a process, no module opens a database or names an xCAT command,
and no step type can run one. A suite that asked xCAT what to expect would
pass on a cluster that could not boot a node.
2026-09-11 08:04:31 -03:00
Daniel Hilst 70a114e2c8 fix(provtest): assert against a whole list, not its first element
An assertion resolved its target through Reply.field, which collapses a list
to its first item. A name with two A records could therefore never satisfy
`data == <the second one>`, and _equal's list branch -- written for exactly
that case -- was unreachable.

Reply.whole returns the container intact and assertions use it. References
keep the old behaviour: $step.data substitutes one value into a command line,
so the first is the right answer there.
2026-09-11 08:04:31 -03:00
Daniel Hilst 1f06d22883 test(provtest): assert the 75 provision clauses as wire scenarios
Sixteen files, fifty scenarios, one per behaviour the specification names.
The cross-stage ones are the point: a config is fetched over TFTP, the
kernel path is read out of it, and that path is fetched. No file says what
the kernel will be called, which is what separates testing xCAT from
comparing xCAT's output against a copy of it.

The four tftp-*.conf files are alternatives, not a set -- a node has one
netboot method -- and the fixture selects the one it was defined with.

conf/tftp-grub2.conf @P-70 meets dhcptest at the only shared assertion:
the boot file DHCP named is fetchable under exactly that name.
2026-09-11 07:55:59 -03:00
Daniel Hilst 7ea4bbdfb4 test(provtest): add provtest, a wire-level provision chain client
provtest asks the management node the questions a booting node asks and
asserts on the answers. DNS, TFTP and HTTP are driven by dig, curl and
tftp, so the result is what a real client sees rather than what this
tool's own protocol parser believed. The xcatd stages use sockets: they
must choose their source address, because xcatd names a client by the
reverse lookup of the address a connection arrived from.

Scenarios are declarative .conf files sharing dhcptest's assertion
grammar, exit codes and TAP output. Nothing here reads the xCAT database
or runs an xCAT command; every expected value arrives via --set.

validate and list run offline, in a checkout, with no root.
2026-09-11 07:55:59 -03:00
Daniel Hilst 6f6acc1b08 docs(provtest): specify the provision chain a booting node observes
The provision chain has no specification anywhere: what a node fetches, in
what order, and what counts as a correct answer live only in the plugins
that emit the artefacts. A cluster can pass every ci_test case and still
boot nothing, because no test in the tree ever acts as a client.

spec.md states 75 clauses, @P-01 to @P-75, each citing the source that
decides it, covering DNS, TFTP, HTTP, discovery, findme, the xcatd request
verbs and the install monitor. Three clauses correct the proposal against
what xcatd actually does; Appendix A records which and why.

README.md documents the tool that runs them.
2026-09-11 07:55:46 -03:00
Daniel Hilst a64dfe149a fix(dhcp): stop handing an installed node a boot file on either backend
A node that has finished installing has chain.currstate "boot", and from
that point the server must stop naming it a boot file: a node given a
netboot script every time it powers on reinstalls itself forever, and does
so silently, because each individual boot looks like a successful one.

Neither backend did that. The wire suite found both.

ISC gated the whole boot-from-disk branch on $doiscsi, so the rule only
fired for a node with an iscsi table row. An ordinary installed node fell
through to the netboot branches below and its xNBA second stage was handed
its own install script. Both the xnba and the pxe branch had it; the pxe
one handed out pxelinux.0.

Kea reads "boot-file-name": "" as *unspecified*, not as "no boot file", so
the empty boot file on the node's reservation did not outrank the classes
every client shares. The always-evaluated xcat-bios class supplied
xcat/xnba.kpxe, the loader came back announcing user class xNBA, and the
per-network class then supplied the network's boot script -- so withholding
only the per-node script withheld nothing. The fix is the same mechanism the
DROP and NOIP classes already use: a xcat-localboot class holding the MACs,
and a "not member('xcat-localboot')" guard on every class that names a boot
file. Kea requires a class to be defined before it is referenced, so
xcat-localboot is written first.

iSCSI is the exception on both backends. Its root disk is on the network and
gPXE is what attaches it, so it still gets a loader: ISC keeps its $doiscsi
branch, and Kea leaves those MACs out of the class.

spec.md S-31 is rewritten to say what correct is -- no boot file at all, on
either request of an xNBA boot, whatever the netboot method -- rather than
the weaker "not its own script", which the second stage above shows is not
the same statement.

conf/localboot.conf asserts the stronger form and gains a netboot=pxe node,
since that branch is written separately and the xnba node passing says
nothing about it. dhcpfixture.sh defines it.

An assertion with != or not-in now holds against a target that is absent
from the reply: a reply naming no boot file has certainly not named the
node's install script. Every other operator still fails on absence, and
"absent" remains the way to assert absence itself.
2026-09-10 17:36:44 -03:00
Daniel Hilst cb345bba06 fix(dhcp): let a Kea reservation name a node without qualifying it
S-35 asks both backends to tell a node its own name in option 12. The
previous attempt wrote the reservation's hostname with a trailing dot, on
the reading that a name already fully qualified is not qualified again.
That is true of Kea 3.0, which is what it was tested against, and false of
Kea 2.4, which is what Ubuntu 24.04 ships and what CI runs: it appends
ddns-qualifying-suffix regardless and the node was told
"node01.pok.stglabs.ibm.com" where ISC said "node01".

Confirmed against both versions with a reservation on a veth pair, DDNS
configured, asking for option 12:

  reservation hostname "probenode."          2.4.1 -> probenode.example.test
                                             3.0.3 -> probenode
  hostname "probenode." and option-data      2.4.1 -> probenode.example.test
  option-data host-name only                 2.4.1 -> probenode
                                             3.0.3 -> probenode

So a reservation's option-data cannot override the field either --
processHostnameOption adds option 12 before appendRequestedOptions runs,
and appendRequestedOptions only fills in what is not already there.
Leaving the hostname field out is the one form that answers with the
node's own name on both versions, and it is also the shape ISC has: an
option statement, separate from the ddns-hostname statement beside it.

kea_boot_for_node already emits the host-name option for every node, so
nothing needed to be added -- only the field removed.

Two consequences, both intended. Kea no longer has a per-reservation DDNS
name and will use whatever the client sends, which is what a cluster with
xCAT's own DNS already relies on; ddns-qualifying-suffix still qualifies
the dynamic clients that have no reservation. And `makedhcp -q` now reads
the name back out of the host-name option, because that is where it is.
2026-09-10 16:44:42 -03:00
Daniel Hilst 8f90c68c82 test(dhcptest): assert an installed node is sent no script, and cite the spec
Two gaps in what the wire suite proves, both about traceability rather than
about a new behaviour.

S-31 had no wire case. A node whose chain.currstate is "boot" has an
operating system and must be left to start it; a node handed a netboot
script every time it powers on reinstalls itself forever, and does so
silently, because each individual boot looks like a successful one. The
Kea side of this is covered by dhcp_kea_plugin_intent.t, but the ISC side
is built inline in addnode against a live database and cannot be reached
by a unit test at all.

The new case asks twice with the same MAC: once as firmware with no user
class, once announcing the xNBA user class the first stage sets. Only the
second request can be answered with the node's script, so only the second
request can see the bug. What is asserted is S-31's own wording -- not
"no boot file", which would forbid the stage-1 binary a server is entitled
to send, but "not the node's xNBA script". The node is defined with
netboot=xnba because only that method generates a second stage.

The remaining ten scenarios already proven on the wire now say which
specification scenario they prove, in the same "S-nn:" form the rest of
the suite uses. That raises the scenarios cited by a wire case from 36 to
47 of the spec's 73, without running anything new -- the coverage was
there and was not traceable.
2026-09-10 16:26:18 -03:00
Daniel Hilst 05185b13db fix(dhcp): make Kea and ISC agree on the wire where the spec says they must
The wire suite compared the two backends against xCAT-test/dhcptest/spec.md and
found four places where a node was told different things depending on which
daemon answered. Each is fixed at the side the spec calls correct.

Option 12 (S-35): Kea builds the host-name option out of a reservation's single
`hostname` field and runs it through ddns-qualifying-suffix on the way out, so a
node asking who it was got "node01.cluster.example.com" while ISC, which writes
option 12 and the DDNS name from separate statements, said "node01". A name that
already ends in a dot is fully qualified and is not qualified again, so the
reservation is now written "node01." -- the wire agrees with ISC, and the suffix
still qualifies the dynamic clients that have no reservation. A reservation's
option-data cannot be used for this: processHostnameOption adds option 12 before
appendRequestedOptions runs, and appendRequestedOptions only fills in options
that are not already present. Reservation lookups are made dot-insensitive so
`makedhcp -d` and `-q` keep matching a node by its bare name.

Option 43 / ISAN (S-20): Kea appends an encapsulated space to a reply only when
the option that carries it is configured too, and option 43 is `type: empty`
with `encapsulate: isan`. Naming only the isan-space sub-options left them with
nothing to travel in and the initiator was offered an address with no target, so
the empty container is now named alongside them. ISC needs no equivalent --
declaring `option isan.iqn` builds option 43 for it.

Boot file fallback (S-12): dropping an architecture branch from the ISC if/else
chain let that client fall through to the final `substring(filename,0,1) = null`
catch-all and be handed /yaboot, which is the loader substitution the spec
forbids. Suppressed architectures now get an explicit empty branch. Kea cannot
have this bug: its classes are independent and xcat-fallback excludes every
recognised arch.

Reload (S-40): `systemctl reload kea-dhcp4` returns 0 once the signal is
delivered, and Kea can then reject the config and keep serving the old one while
systemctl still reads active -- a node added by discovery was never adopted.
dhcp4 and dhcp6 intents therefore always carry a control-socket, and a reload
goes over it so the daemon reports whether it took, falling back to
reset-failed + restart when it cannot confirm.

The fixture now proves a daemon actually holds port 67 before it believes the
backend is serving, which is what turned the reload failure from a flake into a
reproducible case.

Verified on the wire against both backends on EL10 with dhcpd 4.4.3 and
kea 3.0.3: all ten cases rc=0 on each. Unit suite 848/848.
2026-09-10 15:57:33 -03:00
Daniel Hilst 9da497cac4 fix(dhcp): let Kea start, let ISC refuse, and ask for what is asserted
Three things the wire cases turned up, only the first of which is xCAT
answering a client wrongly -- the other two never got as far as an answer.

kea-dhcp4 would not start at all. The ONIE class named option "www-server",
which xCAT's dhcpd.conf declares as code 114 = string but which means the
standard option 72 to Kea -- a list of IPv4 addresses, so an installer URL
in one is a configuration error and the whole Kea pass of the wire cases
never ran. Naming the code instead says the same thing to both.

ISC answered a request for an address on a network it has never heard of
with silence, where the spec says DHCPNAK, because "authoritative" was
written into each subnet declaration and read from the subnet the
requested address belongs to -- which, for this case, is none of them. It
is now global as well, which is what Kea's authoritative:true already
covered.

The remaining two were the cases asking the wrong question. ISC sends an
option when the client asks for it, so an ONIE case that never put 114 in
its parameter request list could not see it however the class was
written; and the common-option case announced a PXEClient vendor class
while asserting the cluster's lease time, which is the ten-minute
firmware lease of S-51 rather than S-50. S-50 now says which clients it
is about.
2026-09-10 14:18:41 -03:00
Daniel Hilst 26e8e277ef fix(dhcp): give both backends one answer for where a node is sent
siaddr comes from noderes, in an order neither backend had right. ISC
read xcatmaster only for petitboot and onie, so every other node in a
hierarchical cluster was sent to the management node rather than to its
service node -- and even for those two the address went into the URL
without going into siaddr, so one reply named two different machines.
Kea read xcatmaster for every node, but when nothing named a server it
used my_ip_facing rather than the subnet's value, which is a different
answer whenever networks.tftpserver names a third machine.

Both now read next_server_for_node: the node's tftpserver, then its
xcatmaster, then -- only for the methods that build a URL and so need an
address in hand -- the interface facing the node. A node that named none
of them inherits the subnet's value, which ISC states in the subnet and
Kea states by leaving next-server out of the reservation.

An xcatmaster that does not resolve is now an error on both rather than
silence on one: it is a misconfiguration, and sending the node somewhere
else instead hides it.

Appendix A rows 26 and 27. Unlike the rest, these two are not in
xcat-internal#175 -- the wire cases found them.
2026-09-10 14:01:22 -03:00
Daniel Hilst 980b7de169 fix(dhcp): close the last four ISC/Kea parity gaps
Appendix A rows 15 to 18, the remaining spec decisions where the two
backends answered the same client differently.

Row 15, a loader that is not on disk is not named. The ISC architecture
chain named every loader unconditionally, so a client whose loader was
never built spent a full TFTP timeout it could not diagnose; the Kea side
had always left the class out. The chain is now built from a list of
gated branches rather than a literal block, because dropping a branch
from a literal if/else if chain can leave a leading "} else if", which
dhcpd rejects outright.

Row 16, a *NOIP* interface draws no reply. Kea discards a packet assigned
to the class named DROP, and only one such class may exist, so every
dropped MAC in the cluster shares it and the user-context records which
node each belongs to. Syncing one node merges into that list and removing
one node prunes only its own entries, so makedhcp for a single node
cannot bring another node's interface back.

Row 17, a node that boots from disk, and row 18, a node deferring to
proxydhcp, are decided per node. Kea host reservations outrank every
client class, so the reservation has to fall silent -- an empty
boot-file-name -- and let the class carry the answer.

Rows 19 and 20 landed earlier and dictated the same shape: the
reservation names nothing and two mutually exclusive classes, one for the
vendor and one for "not the vendor", decide between them, because Kea
evaluates every class independently and has no else.
2026-09-10 13:56:30 -03:00
Daniel Hilst ae7a572132 fix(dhcp): give Kea the ScaleMP and ISAN vendor forms
Appendix A decisions 19 and 20. ISC writes both as an if/else inside the
node's own host block:

  if option vendor-class-identifier = "ScaleMP" { filename = "vsmp/pxelinux.0"; }
  else { filename = "pxelinux.0"; }

Kea has no else, and a reservation outranks every client class, so
anything the reservation names cannot be overridden by the vendor class
that follows it. A ScaleMP hypervisor was therefore handed the
reservation's pxelinux.0 and booted the wrong binary, and an ISAN
initiator -- which reads its initiator name and root path out of option
43 and not out of option 17 -- was sent the standard form it ignores and
nothing it could use.

Both are now written as pairs of mutually exclusive per-node classes,
with the reservation deliberately naming neither the pxe boot file nor an
ISAN node's root path so that a class can decide. This is the same
mechanism the xNBA second stage already used, so the generator, the sync
and the removal are generalised from xnba to per-node classes and each
class carries the purpose that identifies it.

The isan option space ISC declares as "option space isan" is declared to
Kea as an option-def encapsulating option 43, with the initiator name in
sub-option 203 and the root path in 201.
2026-09-10 13:43:42 -03:00
Daniel Hilst 1a531128de fix(dhcp): close four more ISC/Kea drifts from the parity spec
Appendix A decisions 9, 21, 22 and 24. Each one is a difference an
operator never chose: the backend is picked by distribution version, so
whichever side is wrong is wrong on half the clusters.

9  -- Kea never loaded libdhcp_bootp.so, so a BOOTP-only client that ISC
      answered got nothing. The hook now sits alongside host_cmds in one
      hooks-libraries array, and when the hooks package is absent the
      operator is told which clients that leaves unanswered rather than
      being left to find out from a node that never boots.

21 -- The ISC host statements sent "send host-name", which is a dhcpd
      *client* keyword: it never reached option 12 and the node was
      handed no hostname at all. Kea has always sent one.

22 -- The PXE short lease existed on ISC in name only. dhcpd applies
      min-lease-time after max-lease-time, so the cluster default won and
      a pool address taken by a PXE ROM was held for half a day. The
      class now names all three bounds, and Kea is given the same 600
      seconds through xcat-pxe-lease.

24 -- ISC writes "authoritative;" into every generated subnet; Kea
      defaults to the opposite. A node that moved rack asked to keep an
      address from the network it had left and was answered with silence,
      so it waited out a lease that would never be renewed instead of
      being told to start over.

Also makes $xCAT_plugin::dhcp::callback a package variable. It was a file
lexical, so the tests' local() set something nothing read, and warnings
were only ever captured because an earlier process_request had left its
collector behind.
2026-09-10 13:38:32 -03:00
Daniel Hilst e7f7717627 fix(dhcp): give both backends the same answer for eight kinds of client
Eight of the parity decisions in the spec's Appendix A, where ISC dhcpd and
Kea answered the same frame differently and the operator never chose which
backend they got.

ISC gains two branches its if/else chain never had, so the client falls
through to /yaboot no longer:

  - 0x000c ppc64 is given /boot/grub2/grub2.ppc.  yaboot is not a UEFI
    loader and cannot boot one of these machines (decision 2).
  - 0x0010 is the same x86-64 UEFI firmware and the same loader as 0x0007,
    announced by a machine set to fetch it over HTTP (decision 3).

Kea gains what ISC has always had:

  - Etherboot, which predates option 93 and says what it is in option 60
    alone, is recognised and given the BIOS loader (decision 10).
  - onie_vendor is answered per subnet, not only per node: a switch
    announces it on its first boot, before anyone has defined it as a node
    (decision 11).
  - A client nothing else recognises is given /yaboot.  Kea has no else, so
    the condition is the negation of every architecture and vendor class
    another rule answers, rather than a dependence on class ordering
    (decision 7).
  - netboot=nimol is given /vios/nodes/<node> (decision 6).

and drops two answers ISC never gave:

  - No BIOS loader on disk no longer means pxelinux.0 in its place.  Naming
    a file that is not there costs the client a timeout it cannot diagnose,
    and a different loader boots something nobody asked for; the class is
    simply not written (decision 8).
  - netboot=petitboot sends the conf-file and nothing else.  petitboot acts
    on a boot file name when it sees one, so naming one as well sent the
    machine after a TFTP fetch that never happens on ISC (decision 12).
2026-09-10 13:27:58 -03:00
Daniel Hilst 72bae42096 test(dhcptest): write "no boot file" as absent, which is what a reply carries
Four scenarios asserted `bootfile matches ^$` for the cases where the point
is that the client is handed nothing to fetch: OPAL-v3 and s390x, which get
a conf-file instead, a petitboot node, and the architecture whose loader was
taken off disk.

That assertion cannot pass. A server naming no boot file sends an empty
BOOTP file header and no option 67, and an empty value is absent -- so the
run reported "bootfile is not present in the reply" against both backends,
in every case, whatever the server had done. Four failures that said nothing
about either implementation.

`bootfile absent` is the assertion those scenarios wanted. A unit test pins
it, so the unsatisfiable form cannot come back unnoticed.
2026-09-10 13:16:13 -03:00
Daniel Hilst 2abcec617b test(dhcptest): assert the spec scenarios that had no wire case at all
The previous commit deleted the scenarios nothing ran. This one covers the
other half of the same review: places where spec.md states a behaviour and
nothing on the wire ever checked it.

Eight scenarios, each with the fixture work it needs:

  S-36/37/38  next-server follows a node's own tftpserver, then its
              xcatmaster, then the subnet -- three nodes, three sources,
              three different answers, which is the only arrangement in
              which a server that ignores the node attributes can be told
              apart from one that honours them
  S-08        a node whose mac attribute names two ports is reserved on
              both, at the two addresses its two hostnames resolve to
  S-33/34     a diskless node is told its iSCSI target: option 17 for a
              normal client, the ISAN vendor space for one announcing that
              vendor class
  S-12        an architecture whose loader is not on disk is served an
              address and no boot file, with a second architecture whose
              loader is still there as the control
  S-14        a URL names the port the web server actually listens on
  S-05        a dynamic range written as a CIDR block serves out of it
  S-59        a withdrawn reservation stops being handed out
  S-54        a client that speaks BOOTP and not DHCP is still served

Three of these needed the tool to grow: BOOTP has no option 53, so a
BOOTREQUEST step type and a BOOTREPLY message type; option 119 is a name
list that may be compressed against itself, so RFC 3397 decoding, with a
fall back to hex on anything malformed rather than a misreading to assert
against.

RELEASE exposed a state-machine bug on the way. A step that expects no
reply left the client where it was, which is right for a DISCOVER nobody
answered and wrong for a RELEASE, which is never answered and whose whole
point is giving the lease up. A scenario could not boot a node, release
and boot it again -- which is exactly how S-02 has to be asserted, since
a session lasts one scenario and a second boot cannot reference the first
across that boundary.

Every new conf is reachable: three cases0 entries drive the eight fixture
subcommands, so none of them is dead on arrival.
2026-09-10 13:11:51 -03:00