From e57d7ae0b10d5cc9eab166acd341bb1c384a8254 Mon Sep 17 00:00:00 2001 From: Daniel Hilst <392820+dhilst@users.noreply.github.com> Date: Thu, 10 Sep 2026 10:03:46 -0300 Subject: [PATCH] test(dhcptest): close the gaps between the spec and what CI actually runs Measured against spec.md, the wire cases asserted one boot file, for one architecture, on one backend. Seven of the eleven shipped .conf files were executed by nothing at all -- only parsed by `dhcptest validate` -- and the CI run pins the backend to isc, which skipped the one discovery assertion as well. A green run therefore said very little about whether a node can boot. What the fixture now drives: - One DISCOVER per client architecture (option 93), asserting the loader each is handed. This is where the plugin branches most and where a mistake is most expensive. The backends genuinely differ in which architectures they will answer for at all -- Kea emits no UEFI or HTTP-boot class unless the loader is present -- so arch_loader holds that policy in one place and the run skips by name rather than asserting a filename nothing configured. - Both encodings of the user class, against the fix in the previous commit. - The lease itself: the four-way handshake, renewal, rebinding, and a DHCPNAK for an address off this network. A node moved between racks comes back asking for the address it still holds; a server that stays silent leaves it retrying forever, which looks exactly like a node that will not boot. - Discovery end to end: a machine nobody has heard of takes a pool address, is defined, and from the next DISCOVER is served its own address -- with the daemon's pid asserted unchanged across the adoption. xCAT injects ISC reservations over OMAPI precisely so adopting one node does not drop every other node part-way through discovery, and a regression to rewriting dhcpd.conf and restarting would pass every config-level test in this tree. - The options a deployment actually needs: gateway, resolver, domain, MTU, log server and lease time. An address on its own does not install an operating system; it fails later and less obviously. The fixture's network now sets an MTU so there is something to assert. - The hierarchy case, which was mn_only, now runs in CI too. Option 12 is deliberately not asserted: ISC emits `send host-name` through OMAPI and only rewrites it to `option host-name` on the static host fallback path, so it is not reliably on the wire. spec.md records that, along with what this specification does not cover -- DHCPv6, relay agents -- so a green run is not read as more than it is. --- xCAT-test/autotest/testcase/dhcptest/cases0 | 70 +++++- .../autotest/testcase/dhcptest/dhcpfixture.sh | 218 +++++++++++++++++- .../dhcptest/conf/discovery-adoption.conf | 61 +++++ .../dhcptest/conf/nak-foreign-address.conf | 24 ++ .../dhcptest/conf/provision-vs-discovery.conf | 28 ++- xCAT-test/dhcptest/spec.md | 43 +++- 6 files changed, 427 insertions(+), 17 deletions(-) create mode 100644 xCAT-test/dhcptest/conf/discovery-adoption.conf create mode 100644 xCAT-test/dhcptest/conf/nak-foreign-address.conf diff --git a/xCAT-test/autotest/testcase/dhcptest/cases0 b/xCAT-test/autotest/testcase/dhcptest/cases0 index 296bb2835..956a58e5d 100644 --- a/xCAT-test/autotest/testcase/dhcptest/cases0 +++ b/xCAT-test/autotest/testcase/dhcptest/cases0 @@ -73,9 +73,77 @@ cmd:#!/bin/bash check:rc==0 end +start:dhcptest_boot_architectures +description:Each client architecture is handed the loader that belongs to it +label:mn_only,ci_test,dhcp +os:Linux +# Option 93 is where xCAT's DHCP behaviour branches most, and where a mistake +# is most expensive: a machine handed the wrong loader either hangs at the +# firmware or boots the wrong thing entirely. Architectures the running backend +# will not answer for on this machine are skipped by name in the log. +cmd:#!/bin/bash + set -u + FIX=/opt/xcat/share/xcat/tools/autotest/testcase/dhcptest/dhcpfixture.sh + "$FIX" check || exit 0 + trap '"$FIX" teardown' EXIT + "$FIX" setup || exit 1 + "$FIX" run-arch || exit 1 +check:rc==0 +end + +start:dhcptest_chainload_user_class +description:A chainloaded second stage is recognised in both encodings of option 77 +label:mn_only,ci_test,dhcp +os:Linux +# RFC 3004 length-prefixes each string in the user class option; plenty of +# firmware sends the bare string instead, and the same loader sends either +# depending on how it was built. A server that matches only one of them hands +# the loader the first stage again and the machine chainloads itself forever. +cmd:#!/bin/bash + set -u + FIX=/opt/xcat/share/xcat/tools/autotest/testcase/dhcptest/dhcpfixture.sh + "$FIX" check || exit 0 + trap '"$FIX" teardown' EXIT + "$FIX" setup || exit 1 + "$FIX" run-chainload || exit 1 +check:rc==0 +end + +start:dhcptest_lease_lifecycle +description:A lease can be taken, renewed and rebound, and a foreign address is refused +label:mn_only,ci_test,dhcp +os:Linux +cmd:#!/bin/bash + set -u + FIX=/opt/xcat/share/xcat/tools/autotest/testcase/dhcptest/dhcpfixture.sh + "$FIX" check || exit 0 + trap '"$FIX" teardown' EXIT + "$FIX" setup || exit 1 + "$FIX" run-lease || exit 1 +check:rc==0 +end + +start:dhcptest_discovery_adoption +description:A discovered machine is served its own address once defined, without restarting the daemon +label:mn_only,ci_test,dhcp +os:Linux +# The daemon staying up is half the assertion. xCAT injects ISC reservations +# over OMAPI so adopting one node does not drop every other node part-way +# through discovery; a regression to rewriting dhcpd.conf and restarting would +# pass every config-level test in this tree. +cmd:#!/bin/bash + set -u + FIX=/opt/xcat/share/xcat/tools/autotest/testcase/dhcptest/dhcpfixture.sh + "$FIX" check || exit 0 + trap '"$FIX" teardown' EXIT + "$FIX" setup || exit 1 + "$FIX" run-adoption || exit 1 +check:rc==0 +end + start:dhcptest_hierarchy_dhcpserver description:A network whose dynamic pool was handed to another server stops answering unknown MACs -label:mn_only,dhcp +label:mn_only,ci_test,dhcp os:Linux cmd:#!/bin/bash set -u diff --git a/xCAT-test/autotest/testcase/dhcptest/dhcpfixture.sh b/xCAT-test/autotest/testcase/dhcptest/dhcpfixture.sh index 015bde16e..c95889272 100755 --- a/xCAT-test/autotest/testcase/dhcptest/dhcpfixture.sh +++ b/xCAT-test/autotest/testcase/dhcptest/dhcpfixture.sh @@ -18,10 +18,14 @@ # dhcpfixture.sh check is this machine able to run the wire cases # dhcpfixture.sh setup build the network, the node and the config # dhcpfixture.sh run run dhcptest against it +# dhcpfixture.sh run-arch one boot file per client architecture +# dhcpfixture.sh run-lease the lease itself: handshake, renew, rebind, NAK +# dhcpfixture.sh run-chainload first stage versus chainloaded second stage # dhcpfixture.sh backend switch backend and regenerate # dhcpfixture.sh alt-backend name the other backend, if it is installed # dhcpfixture.sh delegate hand the dynamic pool to another server # dhcpfixture.sh run-hierarchy run dhcptest against the delegated network +# dhcpfixture.sh run-adoption discover a machine, define it, serve it its own address # dhcpfixture.sh teardown put everything back # # `setup` records what it changed under $STATE and `teardown` restores it, so @@ -37,6 +41,8 @@ MASK=255.255.255.0 PREFIX=24 SRV_IP=10.99.0.1 POOL=10.99.0.200-10.99.0.250 +MTU=1500 +DOMAIN=dhcptest.cluster NODE=dhcptestcn NODE_IP=10.99.0.11 NODE_MAC=52:54:00:dc:11:01 @@ -44,6 +50,17 @@ NODE_MAC=52:54:00:dc:11:01 # a real machine on a real lab network. UNKNOWN_MAC=02:00:dc:11:00:99 +# The machine that gets discovered and then adopted while the server keeps +# running. Defined by `run-adoption`, not by `setup`, because the whole point +# is what changes between the two DISCOVERs. +ADOPT_NODE=dhcptestcn2 +ADOPT_IP=10.99.0.12 +ADOPT_MAC=02:00:dc:11:00:aa + +# An address on a network this server has never heard of, for the DHCPNAK case. +# 192.0.2.0/24 is TEST-NET-1 and is not routable anywhere. +FOREIGN_IP=192.0.2.77 + STATE=/tmp/dhcptest-fixture DHCPTEST=/opt/xcat/share/xcat/tools/autotest/dhcptest [ -d "$DHCPTEST" ] || DHCPTEST="$(cd "$(dirname "$0")/../../dhcptest" 2>/dev/null && pwd)" @@ -70,19 +87,61 @@ tftpdir() { echo "${dir:-/tftpboot}" } +# The lease time both backends write when site.dhcplease is unset. +lease_time() { + local value + value=$(site_attr dhcplease) + echo "${value:-43200}" +} + # What an unknown machine is told to boot, which is backend policy rather than # protocol: Kea puts an architecture class on the subnet so every client on it # is handed a loader, ISC leaves the boot file to the per-host blocks. An empty # answer means "do not assert this here". discovery_loader() { - local backend + [ "$(current_backend)" = kea ] || return 0 + arch_loader bios +} + +# What a machine of a given client architecture, with no reservation, is told +# to boot. This is the one place in the fixture that has to know each backend's +# policy, because the two genuinely differ in what they will answer at all: +# +# x86 BIOS ISC always names xnba.kpxe; Kea names it only if it is there +# and falls back to pxelinux.0, since Kea has no equivalent of +# dhcpd's "hand it out and let TFTP fail". +# x86-64 UEFI ISC always names xnba.efi; Kea emits no UEFI class at all +# unless the loader exists, so there is nothing to assert. +# aarch64 both, unconditionally. +# riscv64 TFTP both, unconditionally. +# riscv64 HTTP ISC always; Kea only if the loader exists. +# +# An empty answer means "this backend will not answer for this architecture on +# this machine, so do not assert anything". Skipping is the honest outcome: +# asserting a loader that was never configured tests the fixture, not xCAT. +arch_loader() { + local arch=$1 backend tftp backend=$(current_backend) - [ "$backend" = kea ] || return 0 - if [ -f "$(tftpdir)/xcat/xnba.kpxe" ]; then - echo "xcat/xnba.kpxe" - else - echo "pxelinux.0" - fi + tftp=$(tftpdir) + + case "$arch" in + bios) + if [ "$backend" = isc ] || [ -f "$tftp/xcat/xnba.kpxe" ]; then + echo "xcat/xnba.kpxe" + else + echo "pxelinux.0" + fi ;; + uefi) + if [ "$backend" = isc ] || [ -f "$tftp/xcat/xnba.efi" ]; then + echo "xcat/xnba.efi" + fi ;; + aarch64) echo "boot/grub2/grub2.aarch64" ;; + riscv64) echo "boot/grub2/grub2.riscv64" ;; + riscv64http) + if [ "$backend" = isc ] || [ -f "$tftp/boot/grub2/grub2.riscv64" ]; then + echo "http://$SRV_IP/tftpboot/boot/grub2/grub2.riscv64" + fi ;; + esac } current_backend() { @@ -155,9 +214,13 @@ do_setup() { ip link set "$IF_CLI" up || die "cannot bring up $IF_CLI" echo done > "$STATE/veth" + # Every attribute set here is one option the reply has to carry. An + # installer that gets an address and no gateway, resolver or MTU fails much + # later and much less obviously than one that gets no address at all, which + # is why they are configured and asserted rather than left at the default. mkdef -f -t network -o "$NETOBJ" net="$NET" mask="$MASK" mgtifname="$IF_SRV" \ gateway="$SRV_IP" tftpserver="$SRV_IP" nameservers="$SRV_IP" \ - dynamicrange="$POOL" domain=dhcptest.cluster \ + dynamicrange="$POOL" domain="$DOMAIN" mtu="$MTU" \ || die "cannot define the network $NETOBJ" echo done > "$STATE/network" @@ -218,8 +281,18 @@ do_run() { --set node_mac="$NODE_MAC" --set node_ip="$NODE_IP" \ --set node_loader="$(node_loader)" --set pool="$POOL" \ --set next_server="$SRV_IP" --set unknown_mac="$UNKNOWN_MAC" \ + --set gateway="$SRV_IP" --set nameservers="$SRV_IP" \ + --set domain="$DOMAIN" --set mtu="$MTU" --set lease="$(lease_time)" \ conf/provision-vs-discovery.conf || rc=1 + # A reservation is a reservation whichever way the node was reached, so the + # same node has to answer static-vs-dynamic.conf as well: it asks the same + # question from the other end, starting from the pool. + dhcptest_run \ + --set reserved_mac="$NODE_MAC" --set reserved_ip="$NODE_IP" \ + --set unreserved_mac="$UNKNOWN_MAC" --set pool="$POOL" \ + conf/static-vs-dynamic.conf || rc=1 + loader=$(discovery_loader) if [ -n "$loader" ]; then dhcptest_run \ @@ -232,6 +305,128 @@ do_run() { return $rc } +# One DISCOVER per client architecture, asserting the loader each one is handed. +# +# This is the part of xCAT's DHCP behaviour with the most branches and, until +# now, the least wire coverage: a single grub2 assertion for one architecture. +# The architectures a backend will not answer for on this machine are skipped +# by name, so the report says which ones ran rather than quietly passing. +do_run_arch() { + local rc=0 entry arch scenario variable loader + + # arch, the scenario that exercises it, and the variable that scenario + # reads its expected loader from. Every other loader variable is set to a + # value nothing matches, since the .conf declares all of them but only one + # scenario is run at a time. + for entry in \ + "bios:pxe-bios-x86:bios_loader" \ + "uefi:pxe-uefi-x64:uefi_loader" \ + "aarch64:pxe-aarch64:aarch64_loader" \ + "riscv64:pxe-riscv64-tftp:riscv64_loader" \ + "riscv64http:httpboot-riscv64:riscv64_loader" + do + arch=${entry%%:*} + scenario=${entry#*:}; scenario=${scenario%%:*} + variable=${entry##*:} + + loader=$(arch_loader "$arch") + if [ -z "$loader" ]; then + say "skipping the $arch boot file: $(current_backend) does not serve it on this machine" + continue + fi + + # The HTTP scenario asserts a prefix and a substring rather than the + # whole URL, so what it wants is the path inside it. + [ "$arch" = riscv64http ] && loader=boot/grub2/grub2.riscv64 + + dhcptest_run -s "$scenario" --set tftp="$SRV_IP" \ + --set bios_loader=- --set uefi_loader=- \ + --set aarch64_loader=- --set riscv64_loader=- \ + --set "$variable=$loader" \ + conf/pxe-arch-matrix.conf || rc=1 + done + return $rc +} + +# The two halves of a chained network boot. +# +# Firmware PXE sends no user class and must be handed a loader binary. The +# loader that firmware just ran announces itself with user class xNBA and must +# be handed something else -- the per-network script URL -- or it chainloads +# itself forever, and the machine sits at a boot prompt that never advances. +# +# Both encodings of option 77 are sent: the bare string, and the length-prefixed +# form RFC 3004 specifies. The same firmware sends either depending on how it +# was built, so a server that recognises only one of them boots half the fleet +# and loops the other half. Both backends render this branch on the subnet, so +# no node has to be defined with netboot=xnba for it. +do_run_chainload() { + local stage1 + stage1=$(arch_loader bios) + [ -n "$stage1" ] || { say "skipping the chainload cases: no BIOS loader is served here"; return 0; } + + dhcptest_run --set user_class=xNBA --set stage1_loader="$stage1" \ + conf/ipxe-userclass.conf +} + +# The lease itself, rather than what is booted with it: the four-way handshake, +# renewal, rebinding, and the refusal of an address this network cannot give. +# +# The DHCPNAK matters as much as the ACK. A node moved between racks comes back +# asking for the address it still holds; a server that stays silent leaves it +# retrying forever, which on a provisioning network is indistinguishable from a +# node that will not boot. +do_run_lease() { + local rc=0 + dhcptest_run --set net="$NET/$PREFIX" conf/full-lease.conf || rc=1 + dhcptest_run --set net="$NET/$PREFIX" conf/renew-rebind.conf || rc=1 + dhcptest_run --set foreign_ip="$FOREIGN_IP" conf/nak-foreign-address.conf || rc=1 + return $rc +} + +# Discovery, end to end: a machine nobody has heard of takes a pool address, +# gets defined as a node, and from the next DISCOVER on is served its own +# address instead -- with the daemon never restarted in between. +# +# That last part is the assertion with teeth. xCAT injects ISC reservations +# into the leases file over OMAPI precisely so a node being adopted does not +# interrupt every other node still being discovered; a regression to rewriting +# dhcpd.conf and bouncing the daemon would still pass every static test in this +# tree. So the daemon's pid is taken before and after and has to match. +do_run_adoption() { + local rc=0 daemon before after + daemon=$(daemon_of "$(current_backend)") + [ -n "$daemon" ] || die "no DHCP daemon is installed" + + dhcptest_run -s machine-is-unknown \ + --set adopt_mac="$ADOPT_MAC" --set adopt_ip="$ADOPT_IP" --set pool="$POOL" \ + conf/discovery-adoption.conf || rc=1 + + before=$(pgrep -x "$daemon" | head -1) + [ -n "$before" ] || die "$daemon is not running before the node is adopted" + + mkdef -f -t node -o "$ADOPT_NODE" groups=dhcptest ip="$ADOPT_IP" mac="$ADOPT_MAC" \ + arch=x86_64 netboot="$NETBOOT" tftpserver="$SRV_IP" xcatmaster="$SRV_IP" \ + || die "cannot define the node $ADOPT_NODE" + echo done > "$STATE/adopt" + makehosts "$ADOPT_NODE" || die "cannot add $ADOPT_NODE to /etc/hosts" + makedhcp "$ADOPT_NODE" || die "makedhcp $ADOPT_NODE failed" + + after=$(pgrep -x "$daemon" | head -1) + if [ "$before" != "$after" ]; then + say "FAILED: $daemon was restarted to adopt one node (pid $before -> $after)" + rc=1 + else + say "$daemon kept running while $ADOPT_NODE was adopted (pid $before)" + fi + + dhcptest_run -s machine-has-been-adopted \ + --set adopt_mac="$ADOPT_MAC" --set adopt_ip="$ADOPT_IP" --set pool="$POOL" \ + conf/discovery-adoption.conf || rc=1 + + return $rc +} + # Hand the subnet's dynamic pool to another server, the way a management node # does when a service node takes over a rack. Both backends drop a pool whose # networks.dhcpserver is not this host, and the node's own next-server follows @@ -261,6 +456,7 @@ do_teardown() { [ -d "$STATE" ] || return 0 [ -f "$STATE/node" ] && { makedhcp -d "$NODE" >/dev/null 2>&1; makehosts -d "$NODE" >/dev/null 2>&1; rmdef "$NODE" >/dev/null 2>&1; } + [ -f "$STATE/adopt" ] && { makedhcp -d "$ADOPT_NODE" >/dev/null 2>&1; makehosts -d "$ADOPT_NODE" >/dev/null 2>&1; rmdef "$ADOPT_NODE" >/dev/null 2>&1; } [ -f "$STATE/network" ] && rmdef -t network -o "$NETOBJ" >/dev/null 2>&1 [ -f "$STATE/veth" ] && ip link del "$IF_SRV" >/dev/null 2>&1 @@ -299,8 +495,12 @@ case "${1:-}" in alt-backend) if [ "$(current_backend)" = isc ]; then daemon_of kea >/dev/null && echo kea else daemon_of isc >/dev/null && echo isc; fi ;; run) do_run ;; + run-arch) do_run_arch ;; + run-lease) do_run_lease ;; + run-chainload) do_run_chainload ;; delegate) do_delegate ;; run-hierarchy) do_run_hierarchy ;; + run-adoption) do_run_adoption ;; teardown) do_teardown ;; - *) die "usage: $0 {check|setup|generate|backend |alt-backend|run|delegate|run-hierarchy|teardown}" ;; + *) die "usage: $0 {check|setup|generate|backend |alt-backend|run|run-arch|run-lease|run-chainload|delegate|run-hierarchy|run-adoption|teardown}" ;; esac diff --git a/xCAT-test/dhcptest/conf/discovery-adoption.conf b/xCAT-test/dhcptest/conf/discovery-adoption.conf new file mode 100644 index 000000000..43d111590 --- /dev/null +++ b/xCAT-test/dhcptest/conf/discovery-adoption.conf @@ -0,0 +1,61 @@ +# Adopting a machine that is already on the network. +# +# Discovery works like this: a machine nobody has told the cluster about boots, +# takes an address out of the pool, runs a discovery image and reports what it +# is. It is then defined as a node, and from that point on it must come up at +# its own address instead. +# +# The part worth asserting is that the change reaches the wire *while other +# machines are still being discovered*. A server that has to be restarted to +# pick up a new reservation drops every machine part-way through discovery, +# which is why xCAT injects ISC reservations over OMAPI into the leases file +# rather than rewriting dhcpd.conf and bouncing the daemon. +# +# The two scenarios are run separately, against the same MAC, with the node +# defined in between: +# +# dhcptest run -i eth1 -s machine-is-unknown \ +# --set adopt_mac=02:00:dc:11:00:aa --set pool=10.0.0.200-10.0.0.250 \ +# --set adopt_ip=10.0.0.12 conf/discovery-adoption.conf +# ... define the node, run makedhcp for it, restart nothing ... +# dhcptest run -i eth1 -s machine-has-been-adopted \ +# --set adopt_mac=02:00:dc:11:00:aa --set pool=10.0.0.200-10.0.0.250 \ +# --set adopt_ip=10.0.0.12 conf/discovery-adoption.conf + +[defaults] +timeout = 3 +retries = 3 + +[scenario machine-is-unknown] +description = Before it is defined, the machine is served out of the pool + +[step before-discover] +type = discover +mac = %(adopt_mac)s +expect = offer +assert = + msgtype == OFFER + yiaddr in %(pool)s + yiaddr != %(adopt_ip)s + +[scenario machine-has-been-adopted] +description = Once it is defined, the same MAC is served its own address + +[step after-discover] +type = discover +mac = %(adopt_mac)s +expect = offer +assert = + msgtype == OFFER + yiaddr == %(adopt_ip)s + yiaddr not-in %(pool)s + +[step after-request] +type = request +mac = %(adopt_mac)s +requested_address = $offer.address +server_id = $offer.server_id +expect = ack +assert = + msgtype == ACK + yiaddr == %(adopt_ip)s diff --git a/xCAT-test/dhcptest/conf/nak-foreign-address.conf b/xCAT-test/dhcptest/conf/nak-foreign-address.conf new file mode 100644 index 000000000..b97c5689e --- /dev/null +++ b/xCAT-test/dhcptest/conf/nak-foreign-address.conf @@ -0,0 +1,24 @@ +# A machine that comes back asking for an address this network cannot give it. +# +# This is what a node does after it is moved between racks, or after a pool is +# renumbered: it still holds a lease and asks for that address again (INIT-REBOOT, +# a DHCPREQUEST carrying option 50 and no option 54). +# +# An authoritative server must answer DHCPNAK, so the machine gives up on the +# old address and starts over from DISCOVER. A server that stays silent leaves +# the machine retrying an address it will never get, which on a provisioning +# network looks exactly like a node that will not boot. +# +# dhcptest run -i eth1 --set foreign_ip=192.0.2.77 conf/nak-foreign-address.conf + +[scenario foreign-address-is-refused] +description = A request for an address off this network is refused, not ignored + +[step reboot-request] +type = request +requested_address = %(foreign_ip)s +expect = nak +timeout = 3 +retries = 2 +assert = + msgtype == NAK diff --git a/xCAT-test/dhcptest/conf/provision-vs-discovery.conf b/xCAT-test/dhcptest/conf/provision-vs-discovery.conf index b9c587f83..e73e0366c 100644 --- a/xCAT-test/dhcptest/conf/provision-vs-discovery.conf +++ b/xCAT-test/dhcptest/conf/provision-vs-discovery.conf @@ -30,7 +30,7 @@ [defaults] client_arch = 0x0000 vendor_class = PXEClient:Arch:00000:UNDI:002001 -request_options = 1, 3, 6, 43, 51, 54, 60, 66, 67 +request_options = 1, 3, 6, 7, 15, 26, 43, 51, 54, 60, 66, 67, 119 timeout = 3 retries = 3 @@ -61,6 +61,32 @@ assert = bootfile == %(node_loader)s option:1 present +[scenario known-node-gets-a-usable-network] +description = The node is told how to route, resolve and keep time, not just its address + +# An address on its own does not install an operating system. The installer has +# to route off the provisioning network, resolve the server it fetches from, +# and agree with the rest of the cluster about the time; a discovery image has +# to have somewhere to log to. Each of those is one option in this reply, and a +# server that hands out addresses and nothing else fails much later and much +# less obviously. +# +# Asked for as an INIT-REBOOT request -- option 50, no option 54 -- which is +# what a node that already holds this lease sends when it comes back up. +[step known-network-settings] +type = request +mac = %(node_mac)s +requested_address = %(node_ip)s +expect = ack +assert = + msgtype == ACK + option:3 == %(gateway)s + option:6 == %(nameservers)s + option:15 == %(domain)s + option:7 present + option:26 == %(mtu)s + option:51 == %(lease)s + [scenario unknown-node-is-discovered] description = A machine the server does not know gets an address out of the pool diff --git a/xCAT-test/dhcptest/spec.md b/xCAT-test/dhcptest/spec.md index 7858d411c..d88c3cae3 100644 --- a/xCAT-test/dhcptest/spec.md +++ b/xCAT-test/dhcptest/spec.md @@ -225,10 +225,11 @@ Scenario Outline: The user class is recognised however the client encodes it | encoding | | a bare string, as most clients send it | | RFC 3004 length-prefixed, as the RFC says | - # Kea accepts both -- BootPolicy.pm:252 tests option[77].text, the raw hex, - # and the length-prefixed substring. The ISC chain compares - # `option user-class-identifier = "xNBA"` against option 77 declared as a - # plain string (dhcp.pm:4920), which is the bare form only. + # Kea accepts both: kea_xnba_user_class_test tests option[77].text, the raw + # hex, and the length-prefixed substring. ISC accepts both through + # isc_xnba_user_class_test, which pairs the bare comparison with + # `substring(option user-class-identifier, 1, 4)` -- option 77 is declared as + # a plain string there (dhcp.pm), so offset 1 skips the RFC 3004 length byte. Scenario: A known node's second stage is addressed to that node Given a node whose netboot method is xnba @@ -606,8 +607,18 @@ Scenario: Known asymmetries between the backends # Kea offers it only per node. # no match at all ISC falls through to /yaboot; Kea offers # no boot file. - # user class, RFC 3004 form Kea matches the length-prefixed encoding; - # the ISC chain compares the bare string only. + # netboot=xnba, netboot=nimol ISC has a branch for each; + # kea_boot_for_node has neither. + # netboot=petitboot Kea sets boot-file-name as well as + # conf-file; ISC sets conf-file only. + # option 12 (host-name) Kea always sends the node name; ISC emits + # `send host-name` through OMAPI, which is a + # parameter for the reply's sname/file + # handling rather than option 12, and rewrites + # it to `option host-name` only on the static + # host fallback path (dhcp.pm _static_host_ + # statements). Do not assert option 12 in a + # shared .conf. ``` --- @@ -658,3 +669,23 @@ Scenario: The v6 daemon serves the same interfaces as the v4 one pool address versus being ignored -- they belong in separate `.conf` files, chosen by whoever knows how the network under test is configured. Running both against one network will always fail one of them. + +### What this specification does not cover on the wire + +Stated so that a green run is not read as more than it is: + +- **DHCPv6.** `dhcptest` speaks IPv4 only: it builds a BOOTP frame on UDP + 68→67. Everything under *Feature: IPv6* is therefore either asserted from the + generated configuration in `xCAT-test/unit/dhcp_*.t` or not asserted at all. +- **Relay agents.** `giaddr` and option 82 decide which subnet a reply is drawn + from, and no scenario here sends a relayed request: doing it honestly needs a + relay agent on a second network, not a forged `giaddr` from the same wire. + The hierarchy scenarios cover the part that is observable without one -- a + pool handed to another server stops being offered. +- **Service node deployment as a distinct case.** A service node boots exactly + as a compute node does; what differs is what it serves afterwards, which is + `servicenode.dhcpinterfaces` and the delegated pool. Both are covered, as + `[config]` and as the hierarchy scenarios respectively. +- **The architectures no loader exists for on the machine under test.** The + fixture skips those by name rather than asserting a filename that was never + configured; read the log for which ones actually ran.