Skip to content

test(project): tell == 'false' from != 'true' on an evidence output (#967) - #970

Merged
JarryShaw merged 1 commit into
mainfrom
test/967-null-evidence-gate-row
Oct 1, 2026
Merged

JarryShaw merged 1 commit into
mainfrom
test/967-null-evidence-gate-row

Conversation

@JarryShaw

@JarryShaw JarryShaw commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What is the purpose of your pull request?

  • fix — corrects a defect
  • feat — adds a feature
  • perf — changes performance, not behaviour
  • refactor — changes neither behaviour nor performance
  • test — tests only
  • docs — documentation only
  • ci — workflows or build tooling
  • chore — anything else

Description of your pull request and other information

Closes #967. gate_rows grows one row per gated job holding that job's own evidence output at the
module's existing NULL, on a non-v ref with version_check at success, and asserts the gate
skips — the only shape on which == 'false' and != 'true' disagree, and != 'true' is the
dangerous direction: it runs where == 'false' skips, so unknown evidence becomes a publish attempt.

Applying == 'false' → != 'true' to a scratch copy of each real condition (the workflow file is
untouched), violations before → after this PR:

job evidence output evaluator text
tag PCAPKIT_CONDA_TAG_EXISTS 0 → 1 0 → 0
pypi PCAPKIT_PYPI_COMPLETE 0 → 1 0 → 0
conda PCAPKIT_CONDA_COMPLETE 0 → 1 0 → 0

One row per job, resolved from GATE_EVIDENCE[job]; the mutation joins MUTATIONS as evaluator-only
and fails on exactly that new row. github does not have the hole — version_check published no outputs at all holds every output at null with version_check.result still success, so its
evidence clause alone decides that row, now pinned by a test since it is a property of the row.
Reachability is stated rather than overclaimed: the row guards a future change, not a present bug —
neither curl-based check can emit an empty output under set -euo pipefail, but
PCAPKIT_TAG_EXISTS/PCAPKIT_CONDA_TAG_EXISTS come from mukunku/tag-exists-action@v1.7.0 and
tag's gate reads one. Nothing is weakened; row counts go 17/17/21 → 18/18/22.

…967)

`tests/project/test_release_gates.py` could not distinguish `<evidence> ==
'false'` from `<evidence> != 'true'`. Every row holding `version_check` at
`success` also held the evidence outputs at a literal `'true'`/`'false'`, where
the two spellings agree, and `text_gate_violations` only checks that the
output's *name* appears in the condition, never which comparison is made
against it. Measured against the module's own helpers on e1262ed, the
mutation was rejected by neither layer: 0 evaluator and 0 text violations for
each of `tag`, `pypi` and `conda`.

The spellings diverge only on an output that is neither -- `null` -- and for a
publishing job that is the dangerous direction: `!= 'true'` runs where
`== 'false'` skips, so evidence nothing established becomes a publish attempt.
So `release_context` grows an `unknown=` axis holding named outputs at the
module's existing `NULL`, and `gate_rows` grows one row per gated job holding
that job's own evidence at `null` on a non-`v` ref and expecting a skip. The
mutation joins `MUTATIONS` as rejected by the evaluator alone, and it fails on
exactly that one row: 1 evaluator, 0 text per job.

`github` does not have the hole. Its `version_check published no outputs at
all` row holds every output at `null` while leaving `version_check.result` at
`success`, so that gate's evidence clause alone decides the row -- a property
of the row rather than of the gate, and now pinned by a test of its own, since
rewriting the row to fail `version_check` too would take the discrimination
with it.

The new row guards a future change rather than a present bug, and `gate_rows`'
docstring says so: neither curl-based check can emit an empty output under
`set -euo pipefail`, but `PCAPKIT_TAG_EXISTS` and `PCAPKIT_CONDA_TAG_EXISTS`
come from `mukunku/tag-exists-action@v1.7.0` and `tag`'s gate reads the second.

Nothing is weakened: the text assertions and every existing row stay, and the
row counts in `test_the_tables_are_not_empty` go 17/17/21 -> 18/18/22. The
module passes 59 tests / 197 subtests under pytest, and `Ran 59 tests ... OK`
under plain unittest.
@JarryShaw JarryShaw added ci Pull requests that change CI or workflow configuration (ci: subject prefix) review: pending No verdict for the current head - never reviewed, or the head moved since the last one test Pull requests that add or correct tests (test: subject prefix) labels Oct 1, 2026
@JarryShaw

Copy link
Copy Markdown
Owner Author

GOOD TO GO at fe869656a — fable cross-review, first round clean. Different
model from the author (opus).

It found a second weakening the row closes, and I verified both sides. Beyond
== 'false' → != 'true' from #967, the mutant
(<evidence> == 'false' || <evidence> == '') — admitting an empty string — also
escaped the tables as #966 left them:

PRE-PR  (e1262ed4a):  tag/pypi/conda  evaluator=0  text=0   <- survived
POST-PR (fe869656a):  tag/pypi/conda  evaluator=1  text=0   <- caught
row counts: 17/17/21 -> 18/18/22

So this is a family of absent-evidence weakenings, not a single mutation.

Four mutants do survive both layers — contains(fromJSON('["false"]'), …), a
redundant always-true conjunct, == 'False', and == fromJSON('"false"'). All
four are semantically equivalent in Actions (array contains is loose-equality
membership; string comparison is case-insensitive, which the evaluator deliberately
mirrors). A truth-table layer cannot catch an equivalent mutant by construction, so
they are not held against this PR.

Also confirmed independently: no existing row was weakened — the first 17/17/21
rows are identical in name, context and expectation, the new one is a pure append,
and both polarities are still required. And a typo'd unknown= name fails loudly
rather than passing a vacuous row.

One incidental confirmation: a malformed mutant of mine made the evaluator raise
UnsupportedExpression rather than quietly return falsy — the fail-vacuously path
#966's review worried about is genuinely closed.

@JarryShaw JarryShaw added review: good-to-go Cross-review at the current head says ready; CI state is separate and removed review: pending No verdict for the current head - never reviewed, or the head moved since the last one labels Oct 1, 2026
@JarryShaw
JarryShaw merged commit ed3ddd0 into main Oct 1, 2026
63 checks passed
@JarryShaw
JarryShaw deleted the test/967-null-evidence-gate-row branch October 1, 2026 16:02
@JarryShaw JarryShaw removed the review: good-to-go Cross-review at the current head says ready; CI state is separate label Oct 1, 2026
@JarryShaw JarryShaw added this to the 1.5 milestone Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Pull requests that change CI or workflow configuration (ci: subject prefix) test Pull requests that add or correct tests (test: subject prefix)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

test: the release-gate tables cannot tell == 'false' from != 'true' on an evidence output

1 participant