Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 45 additions & 0 deletions CHANGELOG.rst
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,51 @@ Change Log
Unreleased
**********

Added
=====

* Added the static authorization schema, a versioned YAML format for declaring permissions,
permission categories, roles, and role extensions (ADR 0017). The permissions and roles that
``authz.policy`` defines are now also expressed as schema files under
``openedx_authz/authz/schema/``.
* Added the schema loading pipeline in ``openedx_authz/engine/schema/``, covering the discover,
load, validate and compile phases of the lifecycle (ADR 0018), plus render and apply in
``openedx_authz/engine/renderer.py``.
* Added schema discovery through the ``authz.schema`` entry-point group and the
``OPENEDX_AUTHZ_SCHEMA_DIRECTORIES`` setting, so applications can ship authorization definitions
with their code and operators can contribute them through deployment configuration (ADR 0019).
* Added the ``load_authz_schema`` management command, the single non-interactive deployment entry
point, with ``--dry-run`` to print the change report without writing, ``--force`` to allow
removing roles that still have assignments, and repeatable ``--dir`` for CI and local runs.
* Added ``role_extensions`` support: an application or deployment can add or remove permissions and
replace the display metadata or ``hidden`` flag of an existing static role without copying its
definition. ``priority`` resolves conflicts; an unresolvable equal-priority conflict stops the run
before any database change (ADR 0023).
* Added first-class tables for compiled definitions and their provenance, in migration
``0011_authz_schema_definitions``: permission categories, permission definitions, role
definitions, role-permission grants, schema sources, and one source-link table per definition kind
recording whether a contribution was a base definition or an extension (ADR 0025).
* Added source attribution at the role-permission grain, so contributions from different
applications to the same role remain distinguishable and queryable through
``origins_for_role``, ``origins_for_permission``, ``origins_for_category`` and
``origin_for_role_permission``.
* Added a change report before any write: the command lists the policy rows and the definitions that
would be added, updated or removed, including metadata-only edits that change no policy row
(ADR 0018 §6).

Notes
=====

* No authorization behavior changes in this release. Loading ``authz.policy`` works as before, and
the new pipeline runs only when ``load_authz_schema`` is invoked.
* Applying a schema is idempotent and preserves data the loader does not own: user assignments,
dynamic roles, legacy ``g2`` action-inheritance rows, and pre-existing policy rows that no schema
declares. Rows that already exist are adopted, gaining definition and source records rather than
being rewritten (ADR 0025 §6).
* Removing a static role that still has user assignments stops the deployment and reports the
assignments. ``--force`` removes the role together with its assignments and writes a
``RoleAssignmentAudit`` record for each one.

1.24.0 - 2026-09-14
*******************

Expand Down
41 changes: 29 additions & 12 deletions src/openedx_authz/engine/schema/loading.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,8 @@ class SchemaLoader:
"""

_UNKNOWN_DISTRIBUTION = "unknown"
_distribution_ambiguity_warned: set[tuple[str, str, str]] = set()
"""Tracks warned (package, resource_path, selected) combinations to deduplicate logs."""

def load(self, resources: list[DiscoveredResource]) -> list[SchemaDocument]:
"""Load every discovered resource into a :class:`SchemaDocument`.
Expand Down Expand Up @@ -111,7 +113,13 @@ def _resolve_distribution(cls, resource: DiscoveredResource) -> tuple[str, str]:
# pylint: disable=broad-exception-caught
except Exception: # noqa: BLE001 - defensive; metadata quirks across envs
mapping = {}
candidates = mapping.get(top_level) or []
# ``packages_distributions()`` can list the same distribution more than
# once for one top-level package (observed with editable installs and
# overlapping metadata). Those are not competing owners, so collapse to
# the distinct names before resolving — otherwise a lone real owner that
# happens to be listed twice looks "ambiguous" and the fallback path
# warns on every resource for a non-problem.
candidates = sorted(set(mapping.get(top_level) or []))
if not candidates:
return top_level, cls._UNKNOWN_DISTRIBUTION

Expand All @@ -131,7 +139,8 @@ def _select_owning_distribution(cls, candidates: list[str], resource: Discovered
it. If exactly zero or more than one distribution claims the file (or the
file lists are unavailable), we cannot know the true owner, so we return
the first candidate in sorted order — a stable choice across environments
— and log the ambiguity.
— and log the ambiguity once per distinct (package, resource_path, selected)
combination.
"""
if len(candidates) == 1:
return candidates[0]
Expand All @@ -142,16 +151,24 @@ def _select_owning_distribution(cls, candidates: list[str], resource: Discovered
return owners[0]

fallback = sorted(candidates)[0]
logger.warning(
"Could not uniquely resolve the distribution that ships a schema resource; selecting deterministically.",
extra={
"top_level": resource.package,
"resource_path": resource.resource_path,
"candidates": sorted(candidates),
"matched_owners": sorted(owners),
"selected": fallback,
},
)
# Deduplicate warnings by tracking (package, resource_path, selected) combinations
warn_key = (resource.package, resource.resource_path, fallback)
if warn_key not in cls._distribution_ambiguity_warned:
cls._distribution_ambiguity_warned.add(warn_key)
logger.info(
"Schema resource for package '%s' is claimed by multiple distributions %s; "
"deterministically selected '%s'.",
resource.package,
sorted(candidates),
fallback,
extra={
"package": resource.package,
"resource_path": resource.resource_path,
"candidates": sorted(candidates),
"matched_owners": sorted(owners),
"selected": fallback,
},
)
return fallback

@staticmethod
Expand Down
110 changes: 110 additions & 0 deletions src/openedx_authz/engine/schema/pipeline.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
"""End-to-end orchestration of the authz schema lifecycle (ADR 0018).

:class:`SchemaPipeline` wires the steps together:

discover -> load -> validate -> compile -> render -> (plan) -> apply

The Casbin-free steps (discover..compile) live in :mod:`openedx_authz.engine.schema`;
render/apply live in :mod:`openedx_authz.engine.renderer`. This orchestrator is
the single entry point used by the deployment management command and by tests.

Deployment runs discover-through-apply before the application serves traffic
(ADR 0018 §2). CI/local runs may stop after ``plan`` for a dry run, or pass
explicit resources.
"""

from __future__ import annotations

import logging

from openedx_authz.engine.renderer import (
ApplyResult,
ChangePlan,
PolicyRenderer,
SchemaApplier,
)
from openedx_authz.engine.schema.compilation import SchemaCompiler
from openedx_authz.engine.schema.discovery import SchemaDiscovery
from openedx_authz.engine.schema.exceptions import SchemaValidationError
from openedx_authz.engine.schema.loading import SchemaLoader
from openedx_authz.engine.schema.types import CompiledSchema
from openedx_authz.engine.schema.validation import SchemaValidator, ValidationIssue

logger = logging.getLogger(__name__)


class SchemaPipeline:
"""Runs the schema lifecycle from discovery through apply.

Components are injected for testability; each defaults to its standard
implementation.
"""

def __init__(
self,
*,
discovery: SchemaDiscovery | None = None,
loader: SchemaLoader | None = None,
validator: SchemaValidator | None = None,
compiler: SchemaCompiler | None = None,
renderer: PolicyRenderer | None = None,
applier: SchemaApplier | None = None,
):
self._discovery = discovery or SchemaDiscovery()
self._loader = loader or SchemaLoader()
self._validator = validator or SchemaValidator()
self._compiler = compiler or SchemaCompiler()
self._renderer = renderer or PolicyRenderer()
self._applier = applier or SchemaApplier()

def compile(self) -> CompiledSchema:
"""Run discover -> load -> validate -> compile and return the result.

Validation gates twice: once on the loaded documents, then again on the
compiled schema, because extensions and priority resolution can only be
checked after they are applied (ADR 0017 §4).

Raises:
SchemaValidationError: If either validation pass finds error-level
issues.
SchemaCompileError: On an unresolvable conflict.
"""
resources = self._discovery.discover()
documents = self._loader.load(resources)

self._gate(self._validator.validate(documents))
schema = self._compiler.compile(documents)
self._gate(self._validator.validate_compiled(schema))
Comment on lines +75 to +77

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What would happen if two commands with different schemas (let's say one it's out of date) are executed at the same time? I guess the atomic in the previous PR would take care of that?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the previous PR ensures that applying the actual change is atomic, and if two commands run at the same time, the last one finishing will win.

However there is nothing that would prevent this from happening, but given how this command is meant to be run (on deployment, usually via tutor), I don't see this happening easily.

What do you think?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see it more as two operators running the command from different shells at the same time, rather than necessarily two deployments, since the command can also be executed manually.

I don't think we need to prevent that from happening, but we should at least make it visible when another execution is already in progress.


return schema

def _gate(self, issues: list[ValidationIssue]) -> None:
"""Report every issue, then stop the run if any is error-level.

Warnings are logged and the run continues; errors are logged and raised
together so the deployment report lists all of them at once.
"""
for issue in issues:
log = logger.error if issue.is_error else logger.warning
log("authz schema %s: %s [%s]", issue.level, issue.message, issue.source_id or "-")
if self._validator.has_errors(issues):
raise SchemaValidationError([i for i in issues if i.is_error])

def plan(self) -> ChangePlan:
"""Run through render and produce the change report without writing.

Used for dry-run / CI review (ADR 0018 §6).
"""
schema = self.compile()
rendered = self._renderer.render(schema)
return self._applier.plan(rendered, schema)

def apply(self, *, force: bool = False) -> ApplyResult:
"""Run the full lifecycle and persist the result transactionally.

Args:
force: Allow removal of roles that still have assignments (ADR 0018).
"""
schema = self.compile()
rendered = self._renderer.render(schema)
return self._applier.apply(rendered, schema, force=force)
Loading
Loading