Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 30 additions & 1 deletion docs/reference/workflows.md
Original file line number Diff line number Diff line change
Expand Up @@ -632,14 +632,43 @@ Steps can reference inputs and previous step outputs using `{{ expression }}` sy
| `context.run_id` | Current workflow run ID |
| `context.workflow_dir` | Resolved absolute path to the workflow source directory. Empty string for string-loaded workflows. |

Available filters: `default`, `join`, `contains`, `map`, `from_json`.
Available filters: `default`, `join`, `contains`, `map`, `from_json`, `to_json`, `upper`, `lower`, `split`, `length`.

| Filter | Example | Behavior |
| -------- | ------------------------------------------ | ----------------------------------------------------------------------------------------------- |
| `default`| `{{ val \| default('fb') }}` | Fallback for `None` or an empty string |
| `join` | `{{ list \| join(', ') }}` | Join list elements into a string |
| `contains`| `{{ text \| contains('sub') }}` | Substring or membership check |
| `map` | `{{ list \| map('attr') }}` | Extract an attribute from each item |
| `from_json`| `{{ out \| from_json }}` | Parse a JSON string into a typed value |
| `to_json`| `{{ obj \| to_json }}` | Serialize a value to a JSON string — the inverse of `from_json`; mapping keys must be strings |
| `upper` | `{{ text \| upper }}` | Uppercase a string |
| `lower` | `{{ text \| lower }}` | Lowercase a string |
| `split` | `{{ csv \| split(',') }}` | Split a string on a separator into a list of strings |
| `length` | `{{ items \| length }}` | Number of elements in a list, or characters in a string |

`default` falls back only for `None` and the empty string. Other falsy values — `0`, `false`, `[]`, `{}` — are passed through unchanged, so `{{ count | default(10) }}` still yields `0` for a zero count. Falsy is not the same as empty here.

Notes on the newer filters:

- **Types are validated, not coerced.** `upper` and `lower` accept strings only, `split` requires both a string value and a non-empty string separator, and `length` accepts lists and strings but rejects mappings. Anything else raises a `ValueError` naming the problem. Coercion is deliberately not performed: a type mismatch nearly always means the workflow is wired to the wrong variable, and a coerced result would hide that. A filter given the wrong number of arguments (`| upper('x')`, `| split` with no separator, `| split(',', 1)`) is reported as a known filter misused, which is distinct from an entirely unknown filter name: a call carrying more than one argument falls through to that same unsupported-form error rather than being evaluated as a single expression. The older filters (`join`, `map`, `contains`) are more permissive and unchanged: `join` stringifies unsupported values and `map`/`contains` return fallbacks rather than raising.
- **`to_json` output is deterministic.** Object keys are sorted and non-ASCII characters are left as-is rather than escaped, so the same value always serializes to the same bytes. That buys reproducibility only — it does **not** make the result safe to pass through a shell, because [interpolation adds no quoting or escaping](#interpolation-and-shell-safety). Do not interpolate unconstrained JSON into a `run` field.
Comment thread
Quratulain-bilal marked this conversation as resolved.
- **`to_json` rejects non-finite floats.** `NaN`, `Infinity`, and `-Infinity` raise a `ValueError` naming `to_json` instead of serializing to bare tokens. None of the three is valid JSON, so emitting them would hand a standards-compliant downstream parser a string it must reject.
- **`to_json` requires string mapping keys.** JSON objects have string keys, so a mapping with any other key type — `{1: "a"}`, `{True: "a"}` — raises a `ValueError` naming the key type instead. Left to `json.dumps`, an integer key would be coerced to `"1"` and collide with an existing `"1"` key, while mixed key types would fail `sort_keys` with an ordering `TypeError` reported only as "not JSON-serializable".
- **Trailing comparisons after a filter are rejected.** The parser splits on the top-level `|` before looking for operators, so `{{ items | length > 0 }}` raises rather than evaluating. The count does not exist until `length` runs, so there is no way to write that comparison; the supported branching form is the filter's own truthiness in a `condition:`, since `length` returns `0` for an empty input:

```yaml
condition: "{{ inputs.items | length }}" # 0 is False, any non-zero count is True
```

Example:

```yaml
condition: "{{ steps.test.output.exit_code == 0 }}"
args: "{{ inputs.spec }}"
message: "{{ status | default('pending') }}"
tag_count: "{{ inputs.tags | split(',') | length }}"
shell_flag: "{{ inputs.branch | upper }}"
```

### Interpolation and shell safety
Expand Down
193 changes: 180 additions & 13 deletions src/specify_cli/workflows/expressions.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@

import json
import re
from collections.abc import Callable
from contextvars import ContextVar
from typing import Any

Expand All @@ -23,6 +24,11 @@
"map",
"contains",
"from_json",
"upper",
"lower",
"split",
"length",
"to_json",
)


Expand Down Expand Up @@ -130,6 +136,162 @@ def _filter_from_json(value: Any) -> Any:
raise ValueError(f"from_json: invalid JSON: {exc}") from exc


def _filter_upper(value: Any) -> str:
"""Return *value* uppercased.

Raises ``ValueError`` on non-string input. Without the guard a non-string
value (an authoring mistake like ``| upper`` on a shell exit code) would
reach ``str.upper`` and raise a cryptic ``AttributeError`` that escapes the
evaluator and crashes the whole run, since the engine wraps neither
expression evaluation nor ``execute`` in a try/except. Silently coercing
with ``str(value).upper()`` is rejected for the same reason ``from_json``
does not coerce: a type mismatch here means the pipeline is wired to the
wrong variable, and rendering ``"0"`` for an int hides that.
"""
if not isinstance(value, str):
raise ValueError(f"upper: expected a string value, got {type(value).__name__}")
return value.upper()


def _filter_lower(value: Any) -> str:
"""Return *value* lowercased.

Raises ``ValueError`` on non-string input, for the same reason and with the
same trade-off as ``upper``: a type mismatch means the pipeline is wired to
the wrong variable, and coercing would hide it.
"""
if not isinstance(value, str):
raise ValueError(f"lower: expected a string value, got {type(value).__name__}")
return value.lower()


def _filter_split(value: Any, separator: str) -> list[str]:
"""Split *value* on *separator* into a list of strings.

The single-argument ``split(sep)`` form is the only one supported; there is
no maxsplit, because a partial split has no obvious meaning in a workflow
expression and an unused parameter is an authoring mistake worth reporting.

Raises ``ValueError`` when *value* is not a string, *separator* is not a
string, or *separator* is empty. Without those guards a non-string argument
reaches ``str.split`` and raises a cryptic ``TypeError``, and an empty
separator raises the bare ``ValueError: empty separator`` — neither names
the filter, and both escape the evaluator and crash the whole run, mirroring
the strict argument handling in ``join`` and ``map``. An empty separator has
no meaning anyway: ``str.split("")`` is an error in Python, so it is an
authoring mistake rather than a valid edge case.
"""
if not isinstance(value, str):
raise ValueError(f"split: expected a string value, got {type(value).__name__}")
if not isinstance(separator, str):
raise ValueError(
f"split: expected a string separator, got {type(separator).__name__}"
)
if separator == "":
raise ValueError("split: separator must not be empty")
return value.split(separator)
Comment thread
Quratulain-bilal marked this conversation as resolved.


def _filter_length(value: Any) -> int:
"""Return the length of a list or string.

Follows Jinja2's ``length``, which also counts characters in a string. The
two supported input types are exactly ``list`` and ``str``: dicts are
excluded even though ``len()`` accepts them, because a mapping's length is
rarely what a workflow author means by ``length`` and accepting it would
make ``{{ obj | length }}`` silently return a key count for one shape and a
value count for another. Other types (notably ``bool`` and ``None``) are
authoring mistakes and raise rather than coercing to 0.
"""
if isinstance(value, (list, str)):
return len(value)
raise ValueError(
"length: expected a list or string, got "
f"{type(value).__name__} (mappings are not supported)"
)


def _filter_to_json(value: Any) -> str:
"""Serialize *value* to a JSON string — the inverse of ``from_json``.

Serialization is pinned to ``sort_keys=True`` and ``ensure_ascii=False`` so
the output is byte-stable across runs, platforms, and dict insertion order.
It is why these flags are not left to the default: the default key order
varies with insertion order and escapes non-ASCII as ``\\uXXXX``, so the
same workflow would emit different bytes on different runs and hand
downstream tools mangled text. Determinism buys *reproducibility* — the same
value always serializes identically — and nothing more. It does not make the
result safe to pass through a shell: expression interpolation adds no
quoting or escaping, so JSON quotes and metacharacters are still interpreted
by whatever runs the ``run`` field. Interpolate unconstrained JSON into a
shell step only when you have constrained what it can contain; see the
"Interpolation and shell safety" section of ``docs/reference/workflows.md``.

Raises ``ValueError`` when *value* is not JSON-serializable, chained from
the underlying error so the offending type stays visible. ``allow_nan=False``
is what makes that true for non-finite floats: ``json.dumps`` would
otherwise emit bare ``NaN``/``Infinity``/``-Infinity``, none of which is
valid JSON, and hand downstream parsers a string they must reject.

Mapping keys must be strings, which is what JSON objects have anyway.
``_check_json_keys`` enforces that before ``json.dumps`` is reached, so a
non-string key is reported as the authoring mistake it is rather than
surfacing as an ordering ``TypeError`` from ``sort_keys=True`` (mixed key
types) or as silent ``1`` → ``"1"`` coercion that collides with an existing
``"1"`` key.
"""
_check_json_keys(value)
try:
return json.dumps(value, sort_keys=True, ensure_ascii=False, allow_nan=False)
except (TypeError, ValueError) as exc:
raise ValueError(f"to_json: value is not JSON-serializable: {exc}") from exc


def _check_json_keys(value: Any) -> None:
"""Raise ``ValueError`` if any mapping reachable from *value* has a
non-string key.

Walked iteratively with a ``seen`` set: a self-referential structure is
skipped rather than recursed into, so the circular reference stays for
``json.dumps`` to report with its own clearer message instead of the walk
exhausting the stack first.
"""
seen: set[int] = set()
stack: list[Any] = [value]
while stack:
item = stack.pop()
if isinstance(item, dict):
if id(item) in seen:
continue
seen.add(id(item))
for key, sub_value in item.items():
if not isinstance(key, str):
raise ValueError(
"to_json: mapping keys must be strings, got "
f"{type(key).__name__}: {key!r}"
)
stack.append(sub_value)
elif isinstance(item, (list, tuple)):
if id(item) in seen:
continue
seen.add(id(item))
stack.extend(item)


# Filters that take no arguments and tolerate no trailing tokens. Keyed by name
# so ``_apply_filter`` can recognize a mis-wired form of any of them from the
# leading filter name alone, instead of each needing its own branch. ``default``
# is deliberately absent: it is the one filter that is valid both bare and with
# an argument, so it is dispatched by both of the branches below.
_ZERO_ARG_FILTERS: dict[str, Callable[[Any], Any]] = {
"from_json": _filter_from_json,
"upper": _filter_upper,
"lower": _filter_lower,
"length": _filter_length,
"to_json": _filter_to_json,
}


# -- Expression resolution ------------------------------------------------

_EXPR_PATTERN = re.compile(r"\{\{(.+?)\}\}")
Expand Down Expand Up @@ -464,19 +626,21 @@ def _apply_filter(value: Any, filter_expr: str, namespace: dict[str, Any]) -> An
silently returning *value* unchanged: a passthrough would turn a mistyped
or unsupported filter into a wrong result with no signal.
"""
# `from_json` is strict: it takes no arguments and tolerates no trailing
# tokens. Match on the leading filter name and require the whole filter to
# be exactly `from_json`, so every mis-wired form (`from_json()`,
# `from_json('x')`, `from_json)`, `from_json extra`) fails loudly instead of
# silently falling through to the unknown-filter path.
# Zero-argument filters are strict: they take no arguments and tolerate no
# trailing tokens. Match on the leading filter name and require the whole
# filter to be exactly that name, so every mis-wired form (`from_json()`,
# `from_json('x')`, `from_json)`, `from_json extra`, and the same for
# `upper`/`lower`/`length`/`to_json`) fails loudly instead of silently
# falling through to the unknown-filter path.
leading = re.match(r"\w+", filter_expr)
if leading and leading.group(0) == "from_json":
if filter_expr != "from_json":
if leading and leading.group(0) in _ZERO_ARG_FILTERS:
fname = leading.group(0)
if filter_expr != fname:
raise ValueError(
"from_json: expected '| from_json' with no arguments or "
f"{fname}: expected '| {fname}' with no arguments or "
f"trailing tokens, got '| {filter_expr}'"
)
return _filter_from_json(value)
return _ZERO_ARG_FILTERS[fname](value)

# Parse filter name and argument. Use fullmatch (not match) so trailing
# tokens after the closing paren — e.g. a comparison/boolean operator that
Expand Down Expand Up @@ -519,6 +683,8 @@ def _apply_filter(value: Any, filter_expr: str, namespace: dict[str, Any]) -> An
return _filter_map(value, farg)
if fname == "contains":
return _filter_contains(value, farg)
if fname == "split":
return _filter_split(value, farg)
# Filter without args
if filter_expr == "default":
return _filter_default(value)
Expand All @@ -530,7 +696,8 @@ def _apply_filter(value: Any, filter_expr: str, namespace: dict[str, Any]) -> An
name = leading.group(0) if leading else filter_expr
expected = (
"expected one of default or default('x'), join('sep'), "
"map('attr'), contains('s'), or from_json"
"map('attr'), contains('s'), split('sep'), from_json, upper, "
"lower, length, or to_json"
)
if name in _REGISTERED_FILTERS:
raise ValueError(
Expand Down Expand Up @@ -586,7 +753,7 @@ def _evaluate_simple_expression(expr: str, namespace: dict[str, Any]) -> Any:
- Comparisons: ``==``, ``!=``, ``>``, ``<``, ``>=``, ``<=``
- Boolean operators: ``and``, ``or``, ``not``
- ``in``, ``not in``
- Pipe filters: ``| default('...')``, ``| join(', ')``, ``| contains('...')``, ``| from_json``, ``| map('...')``
- Pipe filters: ``| default('...')``, ``| join(', ')``, ``| contains('...')``, ``| from_json``, ``| map('...')``, ``| split(',')``, ``| upper``, ``| lower``, ``| length``, ``| to_json``
- String and numeric literals
"""
expr = expr.strip()
Expand Down Expand Up @@ -889,8 +1056,8 @@ def evaluate_condition(condition: str, context: Any) -> bool:
# strip that trailing newline matches neither branch and falls through to
# ``bool("false\n")`` -> True, silently taking an ``if`` step's ``then``
# branch (and keeping a ``while``/``do-while`` looping) on a step that
# printed "false". A workflow cannot strip it itself -- the registered
# filters are default/join/map/contains/from_json, there is no ``trim``.
# printed "false". A workflow cannot strip it itself -- no registered
# filter trims whitespace (there is no ``trim``).
# ``InitStep._resolve_bool`` and the catalog readers already strip before
# matching boolean text. ``bool(result)`` below still sees the raw string,
# so no non-boolean text changes truthiness.
Expand Down
4 changes: 2 additions & 2 deletions src/specify_cli/workflows/step/switch/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,8 @@ def execute(self, config: dict[str, Any], context: StepContext) -> StepResult:
# ``run: echo approve`` resolves to ``"approve\n"`` and matches no
# ``approve:`` case -- the switch silently falls through to ``default:``
# while still reporting COMPLETED. A workflow cannot strip it itself:
# the registered filters are default/join/map/contains/from_json, there
# is no ``trim``. ``evaluate_condition`` and ``InitStep._resolve_bool``
# no registered filter trims whitespace (there is no ``trim``).
# ``evaluate_condition`` and ``InitStep._resolve_bool``
# already strip before matching a resolved string against declared
# literals, and case keys are exactly such literals. ``expression_value``
# below still reports the raw value, so nothing downstream loses it.
Expand Down
Loading
Loading