2
0
mirror of https://github.com/xcat2/xcat-core.git synced 2026-09-25 09:14:05 +00:00
Files
xcat-core/xCAT-test/unit
Daniel Hilst c8a57880d6 fix(xcat-core): cancelling a build left its workers writing the checkout
Killing the command is not killing the build. sh() ran the command through /bin/sh,
and cancellation signalled that shell alone -- but dpkg-buildpackage starts workers
of its own, and those survive their shell. The lock was then released while they
were still writing debian/changelog and debian/control, which is the state the lock
exists to prevent: the next build takes the checkout and the two rewrite it
together.

The command now runs in its own process group, so cancellation can take all of it.
Both sides call setpgid, so neither depends on which runs first, and INT and TERM
are blocked across the fork so cancellation cannot land in the window before the
group exists.

Cancellation escalates from the caught signal to KILL, and then CHECKS: a shell that
has exited is not a build that has stopped, so it waits for the whole group to
disappear rather than for the leader to be reaped. If the group is still there after
that, the locks are RETAINED and the process exits non-zero. Releasing a lock while
a worker may still be writing is worse than leaving a lock behind for a person to
clear -- the first corrupts a build, the second stops one.

cancel_build ignores INT and TERM while it runs, so a second Ctrl-C cannot interrupt
the cleanup half way and release the lock early.

sh() also reports a signalled command as 128+signal instead of 0. $? >> 8 is zero
for a child killed by a signal, so a build stopped mid-way looked to its caller like
one that had succeeded.

Two cases added to builddebs_lock_cancellation.t: a build whose worker is a
grandchild, and a command killed by a signal. Verified by signalling the pid instead
of the group, which leaves the worker running and turns the first red.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-24 06:02:12 -03:00
..

xCAT-test/unit

Unit tests. These run against the source tree only -- no xCAT installation, no running daemons, no management node.

They are executed on every pull request by the xcat_test GitHub Actions workflow, which calls run_unit_tests() in github_action_xcat_test.pl:

prove -r xCAT-test/unit

You can run exactly the same thing from a clean checkout:

cd <xcat-core checkout>
prove -r xCAT-test/unit

What belongs here

A test belongs in unit/ when everything it needs is in the checkout: plugin and library sources, kickstart/preseed/subiquity templates, postscripts, packaging metadata. Such a test asserts on rendered output or module logic and reaches the repository root through FindBin:

use FindBin;
use lib "$FindBin::Bin/../../perl-xCAT";
use lib "$FindBin::Bin/../../xCAT-server/lib/perl";

Because of those FindBin paths the tests only work from a source tree. The copy installed under /opt/xcat/share/xcat/tools/autotest/unit is not a substitute -- ../.. resolves to /opt/xcat/share/xcat/tools there and the tests die or silently skip. The CI takes a copy of the checkout before the build for this reason; see preserve_source_tree().

What does not belong here

Anything that needs an installed xCAT, a populated /install, a real service binary or a live daemon. Those go in ../integration and run on a management node through xcattest. Both suites run on every pull request -- the workflow installs xCAT on the runner and then runs the ci_test cases against it -- so putting a test in integration/ does not cost it CI coverage. What differs is what each suite is allowed to depend on, and that unit tests also run standalone from a bare checkout with no xCAT at all.

Shell-script unit tests belong in ../bats and run with BATS. Do not add Perl .t tests that grep shell source when the behavior can be exercised by sourcing a shell library or script and shadowing the external commands it calls.

The distinction matters because a test that needs an absent environment does not fail -- it calls plan skip_all and reports as skipped. A handful of those in a suite of several hundred assertions is easy to stop reading. Keeping the two kinds in separate directories means a skip in unit/ is a real signal rather than routine noise.

Guarding on a source file, on the other hand, is fine and common here:

plan skip_all => "compute.subiquity.tmpl not found" unless -f $tmpl_path;

That guard never fires when the tree is intact.