Appendix A rows 15 to 18, the remaining spec decisions where the two
backends answered the same client differently.
Row 15, a loader that is not on disk is not named. The ISC architecture
chain named every loader unconditionally, so a client whose loader was
never built spent a full TFTP timeout it could not diagnose; the Kea side
had always left the class out. The chain is now built from a list of
gated branches rather than a literal block, because dropping a branch
from a literal if/else if chain can leave a leading "} else if", which
dhcpd rejects outright.
Row 16, a *NOIP* interface draws no reply. Kea discards a packet assigned
to the class named DROP, and only one such class may exist, so every
dropped MAC in the cluster shares it and the user-context records which
node each belongs to. Syncing one node merges into that list and removing
one node prunes only its own entries, so makedhcp for a single node
cannot bring another node's interface back.
Row 17, a node that boots from disk, and row 18, a node deferring to
proxydhcp, are decided per node. Kea host reservations outrank every
client class, so the reservation has to fall silent -- an empty
boot-file-name -- and let the class carry the answer.
Rows 19 and 20 landed earlier and dictated the same shape: the
reservation names nothing and two mutually exclusive classes, one for the
vendor and one for "not the vendor", decide between them, because Kea
evaluates every class independently and has no else.
Appendix A decisions 19 and 20. ISC writes both as an if/else inside the
node's own host block:
if option vendor-class-identifier = "ScaleMP" { filename = "vsmp/pxelinux.0"; }
else { filename = "pxelinux.0"; }
Kea has no else, and a reservation outranks every client class, so
anything the reservation names cannot be overridden by the vendor class
that follows it. A ScaleMP hypervisor was therefore handed the
reservation's pxelinux.0 and booted the wrong binary, and an ISAN
initiator -- which reads its initiator name and root path out of option
43 and not out of option 17 -- was sent the standard form it ignores and
nothing it could use.
Both are now written as pairs of mutually exclusive per-node classes,
with the reservation deliberately naming neither the pxe boot file nor an
ISAN node's root path so that a class can decide. This is the same
mechanism the xNBA second stage already used, so the generator, the sync
and the removal are generalised from xnba to per-node classes and each
class carries the purpose that identifies it.
The isan option space ISC declares as "option space isan" is declared to
Kea as an option-def encapsulating option 43, with the initiator name in
sub-option 203 and the root path in 201.
Appendix A decisions 9, 21, 22 and 24. Each one is a difference an
operator never chose: the backend is picked by distribution version, so
whichever side is wrong is wrong on half the clusters.
9 -- Kea never loaded libdhcp_bootp.so, so a BOOTP-only client that ISC
answered got nothing. The hook now sits alongside host_cmds in one
hooks-libraries array, and when the hooks package is absent the
operator is told which clients that leaves unanswered rather than
being left to find out from a node that never boots.
21 -- The ISC host statements sent "send host-name", which is a dhcpd
*client* keyword: it never reached option 12 and the node was
handed no hostname at all. Kea has always sent one.
22 -- The PXE short lease existed on ISC in name only. dhcpd applies
min-lease-time after max-lease-time, so the cluster default won and
a pool address taken by a PXE ROM was held for half a day. The
class now names all three bounds, and Kea is given the same 600
seconds through xcat-pxe-lease.
24 -- ISC writes "authoritative;" into every generated subnet; Kea
defaults to the opposite. A node that moved rack asked to keep an
address from the network it had left and was answered with silence,
so it waited out a lease that would never be renewed instead of
being told to start over.
Also makes $xCAT_plugin::dhcp::callback a package variable. It was a file
lexical, so the tests' local() set something nothing read, and warnings
were only ever captured because an earlier process_request had left its
collector behind.
Eight of the parity decisions in the spec's Appendix A, where ISC dhcpd and
Kea answered the same frame differently and the operator never chose which
backend they got.
ISC gains two branches its if/else chain never had, so the client falls
through to /yaboot no longer:
- 0x000c ppc64 is given /boot/grub2/grub2.ppc. yaboot is not a UEFI
loader and cannot boot one of these machines (decision 2).
- 0x0010 is the same x86-64 UEFI firmware and the same loader as 0x0007,
announced by a machine set to fetch it over HTTP (decision 3).
Kea gains what ISC has always had:
- Etherboot, which predates option 93 and says what it is in option 60
alone, is recognised and given the BIOS loader (decision 10).
- onie_vendor is answered per subnet, not only per node: a switch
announces it on its first boot, before anyone has defined it as a node
(decision 11).
- A client nothing else recognises is given /yaboot. Kea has no else, so
the condition is the negation of every architecture and vendor class
another rule answers, rather than a dependence on class ordering
(decision 7).
- netboot=nimol is given /vios/nodes/<node> (decision 6).
and drops two answers ISC never gave:
- No BIOS loader on disk no longer means pxelinux.0 in its place. Naming
a file that is not there costs the client a timeout it cannot diagnose,
and a different loader boots something nobody asked for; the class is
simply not written (decision 8).
- netboot=petitboot sends the conf-file and nothing else. petitboot acts
on a boot file name when it sees one, so naming one as well sent the
machine after a TFTP fetch that never happens on ISC (decision 12).
Four scenarios asserted `bootfile matches ^$` for the cases where the point
is that the client is handed nothing to fetch: OPAL-v3 and s390x, which get
a conf-file instead, a petitboot node, and the architecture whose loader was
taken off disk.
That assertion cannot pass. A server naming no boot file sends an empty
BOOTP file header and no option 67, and an empty value is absent -- so the
run reported "bootfile is not present in the reply" against both backends,
in every case, whatever the server had done. Four failures that said nothing
about either implementation.
`bootfile absent` is the assertion those scenarios wanted. A unit test pins
it, so the unsatisfiable form cannot come back unnoticed.
The previous commit deleted the scenarios nothing ran. This one covers the
other half of the same review: places where spec.md states a behaviour and
nothing on the wire ever checked it.
Eight scenarios, each with the fixture work it needs:
S-36/37/38 next-server follows a node's own tftpserver, then its
xcatmaster, then the subnet -- three nodes, three sources,
three different answers, which is the only arrangement in
which a server that ignores the node attributes can be told
apart from one that honours them
S-08 a node whose mac attribute names two ports is reserved on
both, at the two addresses its two hostnames resolve to
S-33/34 a diskless node is told its iSCSI target: option 17 for a
normal client, the ISAN vendor space for one announcing that
vendor class
S-12 an architecture whose loader is not on disk is served an
address and no boot file, with a second architecture whose
loader is still there as the control
S-14 a URL names the port the web server actually listens on
S-05 a dynamic range written as a CIDR block serves out of it
S-59 a withdrawn reservation stops being handed out
S-54 a client that speaks BOOTP and not DHCP is still served
Three of these needed the tool to grow: BOOTP has no option 53, so a
BOOTREQUEST step type and a BOOTREPLY message type; option 119 is a name
list that may be compressed against itself, so RFC 3397 decoding, with a
fall back to hex on anything malformed rather than a misreading to assert
against.
RELEASE exposed a state-machine bug on the way. A step that expects no
reply left the client where it was, which is right for a DISCOVER nobody
answered and wrong for a RELEASE, which is never answered and whose whole
point is giving the lease up. A scenario could not boot a node, release
and boot it again -- which is exactly how S-02 has to be asserted, since
a session lasts one scenario and a second boot cannot reference the first
across that boundary.
Every new conf is reachable: three cases0 entries drive the eight fixture
subcommands, so none of them is dead on arrival.
node-specific-second-stage.conf and no-reply.conf were never invoked by any
dhcpfixture.sh subcommand or any case in cases0. A scenario file that nothing
runs is not coverage; it is a claim about the server that no CI failure will
ever contradict. Both are gone, along with the README rows and the header
comments in static-vs-dynamic.conf that pointed readers at them.
What node-specific-second-stage.conf described is real behaviour, so it is
asserted where it actually runs: netboot-methods.conf already defines a node
whose netboot method is xnba, and now asserts that the same node announcing
user class xNBA is handed .../xcat/xnba/nodes/<node> -- in both encodings of
option 77. That is S-25, and it makes the neighbouring first-stage scenario
S-26 as well: the second-stage rule must not change what the firmware sees.
The chainload scenarios asserted only that the second stage differed from the
first. An empty boot file satisfies that, and so does a wrong URL, so both
backends could pass while handing the loader something it cannot run. They now
assert the exact per-network script URL, and a UEFI pair is added for S-23,
where the URL takes a .uefi suffix that a server keying on the user class
alone will not produce. hierarchy-dhcpserver.conf had `bootfile !=` with
nothing on the right-hand side; it asserts the node's own loader instead.
Three header comments still said the backends were free to differ over what an
unknown machine is told to boot. The specification settled that in S-56 -- the
answer has to come from the subnet, on both backends, or a machine cannot
reach the state where anyone could define it -- so the comments now say so.
Also drops the riscv64http branch of arch_loader, which had no caller left
once the HTTP-boot scenario began asserting on the TFTP loader's name as a
substring of the URL.
Appendix A decision 22. xCAT writes `class "pxe" { max-lease-time 600; }`
into the ISC config and `min-lease-time` equal to the cluster default into the
subnet around it, so what firmware is actually granted is not stated anywhere
and has never been checked. Kea has no counterpart at all.
Asserting the number rather than "shorter than the default" is the point: a
lease time nobody wrote down is a lease time nobody can tell has regressed.
The specification had a scenario headed "Known asymmetries between the
backends" and several clauses of the form "on Kea ... on ISC ...". A node
does not choose its management node's DHCP backend -- Backend.pm picks kea
on Ubuntu >= 22.04 and EL >= 10 and isc below -- so every one of those
clauses describes a machine that boots on one distro and hangs on the next,
written down as though it were a requirement.
Each is now decided one way, per the review in VersatusHPC/xcat-internal#175,
and the losing side is recorded as work rather than as behaviour. Appendix A
lists all 25 decisions with the reason and the side that has to change;
appendix B summarises what moved in the document itself. Every scenario also
carries a review number, @S-nn, so an issue can name one behaviour instead of
a paragraph.
The suite follows the specification rather than either implementation:
pxe-arch-matrix.conf grows ia64, ppc64 UEFI, x86-64 UEFI HTTP boot, OPAL-v3,
QEMU s390x, Etherboot by vendor class alone, an ONIE switch that is not yet
a node, and a client that announces nothing at all. It also asks for options
12, 114 and 209, which ISC sends only to a client that asked.
netboot-methods.conf is new: one defined node per netboot method, each with
its own MAC, plus the ScaleMP vendor class, the node's own name in option 12,
and an interface the operator marked *NOIP* that must not be answered.
dhcpfixture.sh defines that node matrix, supplies the loaders and URLs, and
no longer branches on the backend when deciding what an unknown machine is
told to boot.
The cases are expected to fail until the implementations are brought into
line: the failures are the divergence list, on the wire.
The CI run is now three phases in order: the perl and bash unit tests,
the ci_test set, then the DHCP on-wire cases.
The wire cases carry dhcp_wire and deliberately not ci_test. They are the
only cases that do not leave the machine alone while they run -- a veth
pair appears, site.dhcpinterfaces and site.dhcpbackend change, the DHCP
daemon is stopped and started -- and although each one restores all of it
from a trap, a ci_test case running in between would be sharing a
management node that is mid-reconfiguration. Holding them out of the set
is also what stops them being run twice.
run_dhcp_wire_test runs the whole set once per backend installed: against
ISC, then against Kea, each pass bracketed by `dhcpfixture.sh
backend-setup` and `backend-teardown`. It reports separately from the
fast regression test and counts case runs rather than cases, since each
case is run once per backend and each failure carries the backend name.
The phase writes its own xcattest configuration without a dhcpbackend
key: xcattest applies [Table_site] before every case, so the regression
conf's `dhcpbackend=isc` would have reset the backend case by case and
left the run silently testing ISC twice.
arch_loader branched on the running backend, so each backend was measured
against its own behaviour and the two could disagree for ever with the
case still green. That is the drift the case exists to catch.
It now returns one loader per architecture. The backends do read one
thing off the machine before deciding -- dhcp.pm:3469 keys the BIOS and
UEFI classes on xnba.kpxe and xnba.efi being unpacked under tftpdir, and
BootPolicy.pm:84 keys the riscv64 HTTP class on grub2.riscv64 -- while
ISC names all three unconditionally. So setup places the three loaders it
finds missing and teardown removes exactly those, and both backends are
asked the same question. An empty file is enough: dhcptest asserts the
name in the reply and never opens a TFTP session.
Without that, a parity failure would only mean the runner had not
unpacked a loader, which is a fact about the runner and not about xCAT.
The chainload case loses a guard that can no longer fire, and run-arch
loses the branch that skipped an architecture the backend "does not
serve".
Which DHCP backend a management node runs is an implementation default of
its distro -- kea on the newer releases, isc on the older ones -- so a
node booting on the same network has to be told the same things either
way. Until now CI exercised whichever backend the runner happened to
configure, which proves half of that and hides every drift between the
two.
The wire cases now carry a dhcp_wire label, and the CI driver holds them
back from the main pass: it asks dhcpfixture.sh which backends are
installed, then runs the whole set against one and again against the
other. Each pass is bracketed by backend-setup, which points
site.dhcpbackend at the backend and stops the other daemon -- two servers
on one wire both answer the same DISCOVER -- and backend-teardown, which
restores the site table and restarts what was running before.
Running every case under one backend and then every case under the other,
rather than switching inside each case, reconfigures the daemon once per
pass instead of once per case, and gives each failure in the summary the
backend name it belongs to. The summary counts case runs rather than
cases, since a wire case is run once per backend.
The check gate is relaxed from [ -x ] to [ -f ]: the fixture invokes
dhcptest as `python3 src/dhcptest`, so the execute bit only matters to
someone running it directly, and testing it made every wire case skip on
Debian.
dhcptest_backend_switch is dropped, along with the fixture's backend and
alt-backend actions: with every case running under both backends there is
nothing left for a case that switches one.
src/dhcptest was mode 100644 in git. xCAT-test.spec chmods it to 755 on
install, but the Debian package ships it through dh_install, which keeps
the source mode -- so on Ubuntu it arrived non-executable.
dhcpfixture.sh's check gate tested [ -x ], so every wire case skipped and
exited 0. The CI runner is Ubuntu, which means no dhcptest case has ever
exchanged a packet there: each one took about a second, which is less
than the fixture's setup alone, and passed.
Setting the bit in git makes both packages ship it the same way. The
gate itself is relaxed to [ -f ] in the following commit.
The per-node statement for a UEFI node running netboot=xnba matched its two
architecture ids with a parenthesised alternation:
else if <user class> and (option client-architecture = 00:09
or option client-architecture = 00:07) { ... }
ISC dhcpd has no parenthesised grouping in its expression grammar, so that
condition cannot parse. The node-level second stage it guards -- the
per-node .uefi script that keeps two machines chainloading at the same
moment from running the same script -- has never been reachable.
Written as two branches instead, one per architecture id, which is how the
per-network chain in BootPolicy.pm already spells the same test.
dhcp_isc_expression_grouping.t is a guard rather than a test of this branch:
the per-node statements are built inline in addnode against a live database
and cannot be called from a unit test, so it scans the plugin for the one
construct that produces the failure and checks the rendered per-network
chain and the user class test for it as well. dhcpd's tokens only -- Perl's
own parenthesised `exists` and `not` are excluded, and a bareword `option`
after an opening paren cannot be Perl.
isc_xnba_user_class_test wrapped its two alternatives in parentheses so the
trailing `and option client-architecture = ...` would bind to the whole
alternation rather than to the second branch alone. ISC dhcpd has no
parenthesised grouping in its expression grammar, so every generated
dhcpd.conf became unparseable:
/etc/dhcp/.../dhcpd.conf line 11: left brace expected.
if (option user-class-identifier = "xNBA" or substring(option user-cl
^
followed by a cascade of "expecting a parameter or declaration" at each
`} else`, and a daemon that will not start at all.
suffix() gives the same coverage in a single expression that needs no
grouping: it returns the last four bytes, which are "xNBA" whether the
client sent the bare string or the RFC 3004 length-prefixed "\x04xNBA".
The unit test now pins the whole expression rather than each alternative,
and asserts the two properties that broke the daemon -- no leading paren,
no `or` -- so a future rewrite cannot reintroduce grouping unnoticed.
The examples table said "see backend parity below" and nothing below is called
that. The asymmetries are enumerated under "Feature: The two backends behave
the same on the wire", so name that heading and its scenario instead.
Measured against spec.md, the wire cases asserted one boot file, for one
architecture, on one backend. Seven of the eleven shipped .conf files were
executed by nothing at all -- only parsed by `dhcptest validate` -- and the CI
run pins the backend to isc, which skipped the one discovery assertion as well.
A green run therefore said very little about whether a node can boot.
What the fixture now drives:
- One DISCOVER per client architecture (option 93), asserting the loader each
is handed. This is where the plugin branches most and where a mistake is
most expensive. The backends genuinely differ in which architectures they
will answer for at all -- Kea emits no UEFI or HTTP-boot class unless the
loader is present -- so arch_loader holds that policy in one place and the
run skips by name rather than asserting a filename nothing configured.
- Both encodings of the user class, against the fix in the previous commit.
- The lease itself: the four-way handshake, renewal, rebinding, and a DHCPNAK
for an address off this network. A node moved between racks comes back
asking for the address it still holds; a server that stays silent leaves it
retrying forever, which looks exactly like a node that will not boot.
- Discovery end to end: a machine nobody has heard of takes a pool address,
is defined, and from the next DISCOVER is served its own address -- with the
daemon's pid asserted unchanged across the adoption. xCAT injects ISC
reservations over OMAPI precisely so adopting one node does not drop every
other node part-way through discovery, and a regression to rewriting
dhcpd.conf and restarting would pass every config-level test in this tree.
- The options a deployment actually needs: gateway, resolver, domain, MTU,
log server and lease time. An address on its own does not install an
operating system; it fails later and less obviously. The fixture's network
now sets an MTU so there is something to assert.
- The hierarchy case, which was mn_only, now runs in CI too.
Option 12 is deliberately not asserted: ISC emits `send host-name` through
OMAPI and only rewrites it to `option host-name` on the static host fallback
path, so it is not reliably on the wire. spec.md records that, along with what
this specification does not cover -- DHCPv6, relay agents -- so a green run is
not read as more than it is.
A chainloaded second stage announces itself in the user class, option 77, and
must be handed a script URL rather than the loader it just ran -- otherwise it
chainloads itself forever and the machine never finishes booting.
RFC 3004 length-prefixes each string in option 77. Plenty of firmware sends the
bare string instead, and the same loader sends either depending on how it was
built, so both encodings have to be recognised. The Kea policy has always
accepted both. The ISC side compared only `option user-class-identifier =
"xNBA"`, which is the bare form, so a conforming loader booting against an ISC
management node loops -- and does so silently, since from the server's side
every exchange looks like a normal first-stage boot.
Put the test in one place, isc_xnba_user_class_test, and use it from both the
per-network architecture chain and the three per-node statements. The per-node
ones reach dhcpd through omshell, so the helper takes the quoting its caller
needs. `substring(option user-class-identifier, 1, 4)` skips the length byte;
option 77 is declared as a plain string, and the same construct is already used
by the onie_vendor branch a few lines below.
debian_sysconfig_interface_keys fell back to querying the local dpkg when it
was handed no version. The caller always passes one, so the fallback bought
nothing -- and it made the "version could not be established" case answer
differently depending on whether the machine running the code happened to have
isc-dhcp-server installed. It passed on a developer's EL workstation and failed
on the Ubuntu CI runner, which is the failure it was always going to produce.
An undefined version now means what the caller means by it: could not be
established, so write every spelling. Covered by a test that stubs the lookup
to prove the answer no longer depends on the machine.
Write down, as scenarios, what xCAT's DHCP server is supposed to do from the
point of view of a machine booting on a provisioning network: which address it
gets, what it is told to boot, where it fetches the loader from, and what else
the reply has to carry for a deployment to complete.
Read out of the source rather than from the documentation, and cited back to it
line by line, so a scenario that stops matching xCAT is a bug in one of the two
and it is clear which.
Scenarios that cannot be seen from the wire -- a generated file, a daemon's
command line -- are marked [config] and belong to the Perl unit tests instead.
The rest defines the domain the dhcptest wire cases are measured against,
including the ISC/Kea asymmetries the source shows but nothing yet asserts.
makedhcp wrote the interface list into /etc/default/isc-dhcp-server as
INTERFACES="..."
which is no longer always the variable the daemon is started with. The
systemd unit sources that file and expands exactly one variable onto dhcpd's
command line, and which one changed with the package:
14.04 4.2.4-7ubuntu12 sysvinit only $INTERFACES
16.04 4.3.3-5ubuntu12 unit $INTERFACES
18.04 4.3.5-3ubuntu7 unit $INTERFACES
20.04 4.4.1-2.1ubuntu5 unit $INTERFACES
22.04 4.4.1-2.3ubuntu2 unit $INTERFACESv4
24.04 4.4.3-P1-4ubuntu2 unit $INTERFACESv4
26.04 4.4.3-P1-4ubuntu2 unit $INTERFACESv4
with the matching v6 unit reading $INTERFACES before the change and
$INTERFACESv6 after it. Note the boundary is not an upstream ISC release:
20.04 and 22.04 both ship upstream 4.4.1 and are told apart only by the
Debian revision, so it has to be decided on the whole package version.
The sysvinit script does copy INTERFACES into INTERFACESv4 when the latter is
empty, but the unit is what starts the daemon on any of these releases and it
has no such bridge. So from 22.04 on, the unit ran
exec dhcpd -user dhcpd -group dhcpd -f -4 ... -cf $CONFIG_FILE $INTERFACESv4
against a variable xCAT never set. An unset variable expands to nothing, so
dhcpd was launched with no interface argument at all. It does not fail for
that: it binds every interface it can find and only warns about the ones with
no subnet declaration. site.dhcpinterfaces and servicenode.dhcpinterfaces
were therefore silently inert -- the provisioning NIC was served because
makedhcp had also emitted a subnet stanza for it, not because anything
honoured the setting, and any other interface on that subnet was served
alongside it.
The line match compounded it. m/^$dhcpd_key/ is not anchored on the
assignment, and INTERFACES is a prefix of both variables the package ships.
Since 18.04 the postinst seeds only INTERFACESv4="" and INTERFACESv6="" --
no INTERFACES line exists to match -- so the rewrite claimed those two lines
instead and overwrote both, discarding whatever debconf or the administrator
had put there and leaving a file with two INTERFACES assignments and nothing
either unit reads.
Pick the variables through a dispatcher keyed on the installed
isc-dhcp-server version, so both behaviours are served rather than one being
traded for the other, and anchor the match on the '=' so a key cannot claim a
line it is merely a prefix of. When the version cannot be established the
dispatcher writes every spelling, since an unset variable is the outcome that
leaves dhcpd bound to everything. Duplicate assignments left behind by the
old writer are collapsed to one, which repairs a file already damaged on an
upgraded management node.
The EL and SLES paths keep their single DHCPDARGS / DHCPD_INTERFACE /
DHCPD6_INTERFACE key and are unchanged.
On Debian and Ubuntu, makedhcp writes the list of interfaces it is serving
into /etc/default/isc-dhcp-server using the key set at dhcp.pm:2296:
$dhcpd_key = "INTERFACES";
That key stopped being the one the daemon is started with. The variable the
isc-dhcp-server systemd unit expands onto dhcpd's command line changed with
the package; verified by unpacking the archive's own debs:
trusty 4.2.4-7ubuntu12 sysvinit only INTERFACES
xenial 4.3.3-5ubuntu12 unit $INTERFACES
bionic 4.3.5-3ubuntu7 unit $INTERFACES
focal 4.4.1-2.1ubuntu5 unit $INTERFACES
jammy 4.4.1-2.3ubuntu2 unit $INTERFACESv4
noble 4.4.3-P1-4ubuntu2 unit $INTERFACESv4 (+ a v6 unit
reading $INTERFACESv6)
and every package from bionic onward ships a default file whose only
interface variables are INTERFACESv4 and INTERFACESv6, both seeded empty by
postinst. No INTERFACES line ships at all.
Two things go wrong on jammy and later.
The unit runs `exec dhcpd ... -cf $CONFIG_FILE $INTERFACESv4`. That variable
is never set, so it expands to nothing and dhcpd is launched with no
interface argument. dhcpd does not fail for this: it binds every interface
it can find and merely warns about the ones with no subnet declaration. So
site.dhcpinterfaces and servicenode.dhcpinterfaces are silently inert. The
provisioning NIC is served only because makedhcp also emitted a subnet
stanza for it, not because anything honoured the setting, and any other
interface on that same subnet -- a bridge port, a bond member, a second NIC
-- is served too.
The line match is m/^$dhcpd_key/, which is not anchored on the assignment.
"INTERFACES" is a prefix of both variables the package ships, so the rewrite
claims the INTERFACESv4 and INTERFACESv6 lines and overwrites both. The file
is left holding two INTERFACES lines and nothing the units read, discarding
whatever debconf or the administrator had set there.
Extract the rewrite into _sysconfig_interfaces_content() with no change in
behaviour, and add a unit test that asserts the effect rather than the
spelling: it writes the produced file to a temp path and sources it with sh
exactly as the unit does, then checks what would land on dhcpd's command
line. The test is red against the current key and stays red until the writer
sets the variables the daemon is actually started with.
Two things would have kept the wire cases from doing anything on a management
node built from the deb.
src/dhcptest was committed without its executable bit. The spec chmods it at
install time, but dh_install copies the source mode straight through, so on a
deb-installed node the fixture's `[ -x ]` check failed and every wire case
stood down with "dhcptest is not installed" -- passing while testing nothing,
which is the exact failure mode the fixture exists to avoid.
Teardown restarted the backend that was running before the fixture started but
left the one the backend-switch case switched to still running. Two servers on
one network answer the same DISCOVER, so the next case to look at that network
would see an offer from a daemon nobody asked for.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
dhcptest could already assert on what a server puts on the wire, but on a
single-node management node -- the CI runner above all -- there was nothing
for it to talk to, so a wire case would have passed while proving nothing.
dhcpfixture.sh builds the missing network out of a veth pair: the server end
carries the management address and is the only interface named in
site.dhcpinterfaces, the client end has no address, which is the state a
provisioning NIC is in when a machine boots on it. It then defines a network
with a dynamic range and one node with a MAC, runs makedhcp so the
configuration comes from xCAT rather than from hand-editing, and restores
everything from a trap. All the xCAT knowledge lives there, next to the
xcattest cases, so dhcptest itself stays agnostic to it.
The new scenarios assert what a provisioning server has to get right:
provision-vs-discovery.conf a known machine gets its reservation and its
own loader, an unknown one gets a pool address
discovery-bootfile.conf an unknown machine is told what to boot too
node-specific-second-stage.conf the chainloaded stage names its host
hierarchy-dhcpserver.conf a subnet whose pool was handed to another
server ignores unknown MACs but still points
known ones at it
Those are four files rather than four scenarios in one because they describe
server configurations that contradict each other, and because --set values are
resolved when a file is loaded, before -s selects anything: an unused
scenario's missing variable would otherwise abort a run that never intended to
use it. ISC dhcpd leaving the unknown-client boot file to the per-host blocks
while Kea puts it on the subnet is exactly such a contradiction, which is why
discovery-bootfile.conf is opt-in and the fixture only asks for it on Kea.
dhcptest_backend_switch reruns the same files after switching site.dhcpbackend
and regenerating, so the backend-agnosticism the tool exists to demonstrate is
itself under test rather than asserted in a comment.
The `in` operator now accepts an inclusive first-last range as well as a CIDR
and a list: a dynamic pool is written as two addresses because it rarely lines
up on a prefix boundary.
Verified against a real dnsmasq in a network namespace: 17/17, 3/3 and 5/5
assertions green, a deliberately wrong expectation failing with expected and
received on one screen, and host networking untouched.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xCAT's DHCP behaviour is tested from two directions today, neither of which
touches the wire: xCAT-test/unit/dhcp_*.t assert on generated config text, and
xCAT-test/integration/dhcp_{isc,kea}_config_validation.t feed that config to a
real daemon and check it parses. Both stop at "the server accepted our
config". Nothing verifies that a client sending option 93 = 0x000b actually
gets an aarch64 loader back.
The one existing wire tool, xCAT-probe/subcmds/detect_dhcpd, hand-packs a
fixed DISCOVER carrying no option 60, 77 or 93 (so it never reaches any arch
branch), parses replies by regexing tcpdump output, and binds an
IO::Socket::INET to a local address, so it cannot run on a provisioning NIC
that has no address yet.
dhcptest fills that gap. It is a Python 3 tool using only the standard library
plus Scapy, speaking raw Layer 2, driving DHCP transactions declared in INI
files and asserting on what came back:
dhcptest run -i eth1 --set net=10.0.0.0/24 conf/full-lease.conf
dhcptest validate conf/*.conf # no root, no network, no scapy
dhcptest list conf/*.conf
dhcptest discover -i eth1 # ad-hoc, no conf file
Scenarios cover DISCOVER/OFFER, the full lease, renew and rebind, reserved
versus pooled addresses, deliberate silence, the PXE architecture matrix and
the iPXE user-class split. Assertions are a small DSL -- target, operator,
value -- over message type, BOOTP header fields, options by number or name,
subnet membership, and the boot file a client would actually use (option 67
or the header, since servers differ on which they fill).
Two boundaries are enforced by unit test rather than by convention:
- The tool is implementation-agnostic. It exchanges packets and checks
fields; it never reads the xCAT database, never runs an xCAT command, and
names no DHCP implementation. Anything that differs between servers is
expressed by whoever writes the .conf.
- Only runner.py imports scapy, so validate and list stay usable in CI on a
host with neither scapy nor root.
Host networking is never modified. The client MAC defaults to a synthetic
locally-administered address, everything goes over a raw L2 socket, no leased
address is ever configured, and the ARP responder answers only for addresses
the session was actually granted.
Variables in .conf files are configparser's own BasicInterpolation, %(name)s,
fed from [vars] and --set. $offer.address and friends are resolved separately
at step-execution time, because they name a reply that has not arrived when
the file is read; BasicInterpolation gives $ no meaning, so the two coexist
without escaping.
Packaging and CI:
- xCAT-test.spec installs dhcptest alongside unit/ and integration/, and
only recommends python3-scapy: it lives in EPEL on EL, and a hard Requires
would make xCAT-test uninstallable on a management node without EPEL.
The deb depends on it outright.
- github_action_xcat_test.pl runs the unit tests and `dhcptest validate`
last, after the fast regression.
- autotest/testcase/dhcptest/cases0 carries the two offline cases plus a
wire case that stands down with an explanation when no provisioning NIC
is available.
Verified against a real DHCP server in a network namespace: 34 lease, renew,
rebind, reservation and silence assertions and 22 PXE/iPXE assertions all
pass, and host networking is untouched before and after.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The plugin is loaded and install_darch is called for the architectures
xCAT installs Ubuntu on and two it does not. Against the previous plugin the
test fails on the missing function: the mapping was inline in mkinstall.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
mkinstall mapped x86_64 and x86 to their Debian names and accepted ppc64le
and ppc64el. Every other architecture, riscv64 included, was logged as
"Unknown arch" on each diskful install, although the install went on with
the name unchanged, which is right for riscv64.
Move the mapping into install_darch, which takes the Debian name from
xCAT::Utils::debian_arch and knows the architectures xCAT installs Ubuntu
on. riscv64 is one of them.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The grub-common assertion fails against the previous metapackage. The rest pins
what must NOT change: a riscv64 management node still recommends the x86 boot
payload and the Genesis images of the other architectures, because it serves
them to the nodes it provisions.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
copycd builds the riscv64 boot loader by running grub-mkimage, which
grub-common ships on every supported Ubuntu release. Nothing declared it, so a
management or service node installed without that package copies riscv64 media
and produces no loader, while DHCP keeps pointing every riscv64 node at the path
where the loader should be.
The declaration belongs to xcat-server, which carries the plugin that runs the
command, so both metapackages inherit it.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The stubbed test proves the decisions copycd makes; this one proves the
artifact. It runs the real grub-mkimage over the grub2 package of copied media
and checks the image against the validation nodeset depends on, so a package
layout change or a grub-mkimage that stops accepting these inputs is caught
where it happens.
It needs media and the grub2 tools, so it skips without
XCAT_TEST_UBUNTU_RISCV64_MEDIA. Run on a management node against the 24.04 and
26.04 riscv64 trees.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The mirror assertions passed a Debian architecture straight in and so never
exercised the conversion genimage performs first. They now start from the xCAT
osarch value, which is what let the 32-bit x86 token reach the wrong archive.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
genimage converts osimage.osarch with debian_arch and uses the result twice: it
picks the apt mirror and it becomes debootstrap --arch. The map knew x86_64 but
not x86, so an image with osarch=x86 selected the ports archive, which carries
no i386, and then asked debootstrap for an architecture it does not know.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Four assertions fail against the previous debian/control: the dependency has no
architecture restriction, and dpkg's own parser still reports it for riscv64.
The last assertion pins that the restriction drops nothing else.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
xcat-genesis-scripts-amd64 is Architecture: all, so the plain Depends installed
the x86 legacy Genesis scripts, and the x86 Genesis base with them, on a riscv64
management node. riscv64 has no legacy Genesis: its image ships as
xcat-genesis-openembedded-riscv64, which the metapackage already recommends and
mknb consumes. xCAT.spec makes the same distinction on the rpm side.
The dependency is now restricted to the architectures that have a legacy
Genesis. amd64 and ppc64el keep it unchanged.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The guides said riscv64 covered EL10 only, and the riscv64 page listed
Ubuntu as unsupported. Both support matrices now carry the architecture for
Ubuntu, and the riscv64 page describes the Ubuntu paths: the loader copycd
builds from the media, the installer needing none of the accommodations
EL10 requires, the ports archive the packages come from, and a management
node running on riscv64. The 26.04 media need the RVA23 profile, which is
recorded as a limitation.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The note shown when boot/grub2/grub2.<arch> is missing told the
administrator it comes from grub2-xcat or the EL installation media. On
Ubuntu neither is true: copycd builds the loader from the grub2 package on
the media, because the image the media carry cannot boot over the network.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The noderes.netboot table gave riscv64 the condition >=el10, so the only
documented way to reach grub2 on the architecture was an EL node, although
lookupNetboot answers grub2 for riscv64 whatever the operating system. The
condition now names Ubuntu as well, the way ppc64 names both of its
families.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
go-xcat stopped on riscv64 before it reached the package manager, so the
installer xCAT documents could not set up the management node the riscv64
packages are built for. The architecture is now accepted alongside the
others.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Calls the selection the template performs. Covers the ports archive for
ppc64el and both spellings of ppc64, the main archive for every x86
spelling, an unknown architecture keeping the previous default, and
site.ubuntu_apt_mirror overriding both without an empty value blanking the
mirror.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
archive.ubuntu.com publishes amd64 and i386 only, so a ppc64el or riscv64
node was given an apt mirror carrying no package for it and the installer
could not fetch what the minimal live media lacks. The default is now the
ports archive for those architectures, chosen from the osimage's
architecture rather than the package directory, which is whatever path the
administrator configured. site.ubuntu_apt_mirror still overrides it.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Calls the escaping the plugin performs. Covers the Ubuntu installer seed
keeping the arguments after it, every separator grub2 recognizes, a command
line without one staying byte for byte the same, escaped and quoted values
surviving unchanged, a separator beside a quoted span in the same word, and
a variable reference being left alone so BOOTIF still expands.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
grub2 reads its configuration as a script, so an unquoted command separator
ends the linux command and everything after it is lost. The Ubuntu
installer seed is written as ds=nocloud-net;s=<url>, so the node booted
without the seed URL and without the arguments that followed it, including
BOOTIF. The installer then found no autoinstall configuration and waited
for someone to answer its questions. A separator that is neither escaped
nor inside a quoted span is now escaped where it stands, which grub2
removes before it hands the line to the kernel. A value the caller escaped
or quoted keeps exactly the form the caller gave it.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Reads the list imgutils returns. Covers the drivers a node needs to reach
its root filesystem, the overlay module, the architectures that already
worked keeping theirs, and an unknown architecture still getting an empty
list.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
The Ubuntu driver table had no riscv64 entry, so genimage was handed an
empty list and built an image carrying no network module. A node whose NIC
is not built into the kernel then has no interface to fetch its root
filesystem with. The architecture now gets the same drivers the enterprise
Linux table lists for it, plus the overlay module every Ubuntu image needs.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Evaluates the selection genimage performs, so the test tracks the script.
Covers the ports archive for ppc64el and riscv64, the main archive for
amd64 and i386, and site.ubuntu_apt_mirror overriding both without an
empty value blanking the mirror.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>