Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/actor-error-reporting.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"effect-machine": minor
---

Report contained actor failures to the Effect `ErrorReporter`s that were registered where the actor was spawned. Each runtime generation reports its complete defect cause once when it closes, including transition, spawn, task, background, and cleanup failures. Supervised restarts report each generation. A restart step that fails before a new generation exists reports with phase `restart`. This includes a restart recovery and a restart schedule that dies; an exhausted schedule does not report. A cold-start recovery that fails before the first generation runs reports with phase `recovery`, also when a stop interrupts the start. A recovery that throws instead of returning an Effect reports the same way. A schedule defect that settles together with schedule exhaustion still reports. A final output defect reports when the actor completes. Child actors report through the reporters of the handler that spawned them. Normal stops, final states, and pure interruption do not report. Each report annotates the reporting fiber with the actor ID, generation, and defect phase. The actor keeps its spawn-time reporters, so a later `start` or `stop` caller cannot replace them. Handler code inside the actor also runs with the spawn-time reporters, so `Effect.withErrorReporting` in a handler reaches them and no longer reaches reporters that were provided only around `start`. A failure that also reaches a caller boundary, such as `Effect.withErrorReporting`, an RPC server or an HTTP handler, can report again there. A reporter built with `ErrorReporter.make` skips the second report only when the defect is an object; a primitive defect or a raw reporter reports twice. A parent handler that re-raises a child failure, such as a failed child stop or a failed `self.spawn` start, reports that cause again as the parent's defect, with the same deduplication rules. A throwing reporter does not change the actor exit. Effect calls the reporters in order, so a throwing reporter can stop the reporters after it. Cluster entity runtimes report generation defects through the reporters in their allocation context. The generation of an entity report names one allocation, so it is `0` after each reactivation. Without registered reporters, behavior does not change.
8 changes: 8 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -203,6 +203,14 @@ const otherExit = yield * other.awaitExit; // ActorExit<OtherState>
- **Classifier** — `shouldRestart` optionally skips restart for specific defect types
- Entity-machine: cluster-supervised via `defectRetryPolicy`, NOT local supervision

## Error Reporting

- Lifecycle owners report to the Effect `ErrorReporter`s captured at spawn: generation closure reports its combined defect, terminal completion reports a final output defect, and the supervisor reports a restart step that fails before a new generation exists.
- `createActor` pins `CurrentErrorReporters` in the actor service context. Restarted generations allocate in the supervisor fiber, which also carries the start caller's context; do not let that caller replace the spawn-time reporters.
- Report before publishing the exit, and isolate reporter defects. A throwing reporter must not block exit settlement.
- Inside the actor runtime, do not report a cause again where it is only re-raised: stop failures belong to the generation, and implicit-system close failures belong to the child actors. A parent handler that re-raises a child failure is a new parent defect, so the parent generation reports it too. That overlap, like the overlap with caller boundaries, relies on `ErrorReporter.make` identity dedup and is documented as such.
- Wrap user callbacks that return an Effect (such as `lifecycle.recovery.resolve`) in `Effect.suspend` before attaching report hooks, so a synchronous throw becomes a reported defect.

## Child Actors

Spawn children from `.spawn()`/`.background()` handlers via `self.spawn(id, childMachine)`:
Expand Down
1 change: 1 addition & 0 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -313,6 +313,7 @@ An exit animation can keep a screen mounted after a final transition. Keep requi
8. **call vs send**: `send` = fire-and-forget, `call` = request-reply, `ask` = typed reply
9. **Non-Effect code**: Use `actor.client`. React and Solid should use Actor Atoms.
10. **ActorStoppedError**: Pending `call`/`ask` Deferreds settled on stop
11. **Error reporters at spawn**: provide `ErrorReporter.layer` where the actor is spawned. Lifecycle owners report each generation defect, restart-step failure, and final output defect to those reporters. Pure interruption does not report.

## Cluster / Entity Machines

Expand Down
25 changes: 25 additions & 0 deletions docs/persistence-and-supervision.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,31 @@ Machine.spawn(machine, {

See [`supervision.ts`](../examples/core/src/supervision.ts).

## Error reporting

The actor lifecycle reports contained failures to Effect `ErrorReporter`s. Register the reporters in the context that spawns the actor:

```ts
Machine.spawn(machine).pipe(Effect.provide(ErrorReporter.layer([reporter])));
```

- Each generation reports its complete defect once, when it closes. The cause includes transition, spawn, task, background, and cleanup failures of that generation.
- Each supervised restart reports its own generation. A restart step that fails before a new generation exists reports with phase `restart`. This includes a restart recovery that dies and a restart schedule that dies. An exhausted schedule does not report.
- A cold-start `lifecycle.recovery.resolve` that fails before the first generation runs reports with phase `recovery`. It also reports when a stop interrupts the start.
- A final output defect reports once when the actor completes.
- Child actors report their own failures. They capture the reporters of the parent handler that spawned them.
- Normal stops, final states, and pure interruption do not report.
- The actor keeps the reporters that it captured at spawn. A later `start` or `stop` caller cannot add or replace them.
- Handler code inside the actor also runs with the spawn-time reporters. `Effect.withErrorReporting` in a handler reaches them. It does not reach reporters that were provided only around `start`.
- Each report annotates the reporting fiber with `effect_machine.actor.id`, `effect_machine.actor.generation`, and `effect_machine.defect.phase`. Reporters read them from `References.CurrentLogAnnotations`.
- A failure that also reaches a caller can report again at the caller's boundary, for example `Effect.withErrorReporting`, an RPC server or an HTTP handler. Use `ErrorReporter.make`. It skips a cause or an error object that it already reported. A primitive defect, such as `Effect.die("boom")`, is not an object, so it reports twice. A raw reporter that is not built with `ErrorReporter.make` also reports twice.
- A parent handler that re-raises a child failure fails the parent too, so the parent reports that cause again as its own defect. Examples are a scoped child whose stop fails and a `self.spawn` whose start fails. `ErrorReporter.make` skips the second report when the defect is an object. A primitive defect or a raw reporter reports twice.
- A cold-start recovery that throws instead of returning an Effect reports like a recovery that dies.
- A reporter that throws does not change the actor exit. Effect calls the reporters of one set in order, so a throwing reporter can stop the reporters after it.
- A cluster entity report names one allocation. Its generation is `0` after each reactivation.

Inspection `@machine.error` events stay diagnostics. They do not replace reporting.

## Local and cluster durability

Local actors use lifecycle hooks. Entity machines use the cluster persistence adapter.
Expand Down
94 changes: 76 additions & 18 deletions src/actor.ts
Original file line number Diff line number Diff line change
Expand Up @@ -10,14 +10,17 @@ import {
Deferred,
Cause,
Effect,
ErrorReporter,
Exit,
Fiber,
Layer,
MutableHashMap,
Option,
PubSub,
Pull,
Queue,
Ref,
Result,
Schedule,
Scope,
Semaphore,
Expand All @@ -44,6 +47,7 @@ import { DuplicateActorError, ActorStoppedError } from "./errors.js";
import {
createRuntime,
notifyStateListeners,
reportLifecycleFailure,
type RuntimeLifecycleHooks,
type RuntimeQueuedEvent,
type RuntimeHandle,
Expand Down Expand Up @@ -693,6 +697,7 @@ function activateGeneration<S extends AnyState, O>(
/**
* Run the supervision loop for a supervised actor.
* Observes exit deferred, applies restart policy, resets cell resources on restart.
* Returns the terminal generation exit. The actor owner completes it.
* @internal
*/
const runSupervisionLoop = <
Expand All @@ -709,9 +714,10 @@ const runSupervisionLoop = <
) => Effect.Effect<RuntimeHandle<S, E>>;
lifecycle?: Lifecycle<S, E>;
onRestart?: (generation: number, exit: ActorExit<unknown>) => Effect.Effect<void>;
onTerminal: (exit: RuntimeExit<S>) => Effect.Effect<void>;
/** The reporters captured at spawn. */
errorReporters: ReadonlySet<ErrorReporter.ErrorReporter>;
},
) =>
): Effect.Effect<RuntimeExit<S> | undefined> =>
Effect.gen(function* () {
const step = yield* Schedule.toStepWithSleep(options.supervision.schedule);

Expand All @@ -723,24 +729,30 @@ const runSupervisionLoop = <
const generationExit = yield* Deferred.await(currentRuntime.exitDeferred);
yield* currentRuntime.awaitClosed;

if (generationExit._tag !== "Defect") {
yield* options.onTerminal(generationExit);
return;
}
if (generationExit._tag !== "Defect") return generationExit;

if (
options.supervision.shouldRestart !== undefined &&
!options.supervision.shouldRestart(generationExit)
) {
yield* options.onTerminal(generationExit);
return;
return generationExit;
}

const pull = step(generationExit);
const scheduleExit = yield* pull.pipe(Effect.exit);
if (scheduleExit._tag === "Failure") {
yield* options.onTerminal(generationExit);
return;
// An exhausted schedule halts its pull. A defect can settle with that halt, for example in a
// finalizer, so report whatever remains once the halt is removed.
const scheduleFailure = Pull.filterDone(scheduleExit.cause);
if (Result.isFailure(scheduleFailure)) {
yield* reportLifecycleFailure(scheduleFailure.failure, {
reporters: options.errorReporters,
actorId: cell.id,
generation: cell.generation.current,
phase: "restart",
});
}
return generationExit;
}

// Bump generation before restart — recovery.resolve sees the new generation
Expand Down Expand Up @@ -817,6 +829,7 @@ export const createActor = Effect.fn("effect-machine.actor.spawn")(function* <
) {
const lifecycle: Lifecycle<S, E> | undefined = options.lifecycle;
const capturedContext = yield* Effect.context<R>();
const errorReporters = yield* ErrorReporter.CurrentErrorReporters;

// Spawn is cold. The caller has already resolved machine input and hydration.
// Recovery runs during start, not allocate.
Expand All @@ -832,7 +845,12 @@ export const createActor = Effect.fn("effect-machine.actor.spawn")(function* <
const localInspector = options.inspect ?? ambientInspector;
const systemInspectors = systemInspectorsBySystem.get(system) ?? new Set<SystemInspector>();
const inspectorValue = makeInspectionDispatcher(localInspector, systemInspectors);
const serviceContext = Context.add(capturedContext, ActorInspection, inspectorValue);
// Pin the spawn-time reporters. Restarted generations allocate in the supervisor
// fiber, which also carries the start caller's context.
const serviceContext = capturedContext.pipe(
Context.add(ActorInspection, inspectorValue),
Context.add(ErrorReporter.CurrentErrorReporters, errorReporters),
);

// Actor-specific state
const childrenMap = new Map<string, ActorRef<AnyState, unknown>>();
Expand Down Expand Up @@ -929,6 +947,16 @@ export const createActor = Effect.fn("effect-machine.actor.spawn")(function* <
terminalCause = cause;
terminal = { _tag: "Defect", cause, phase: "cleanup" };
}
// Generations report their own failures, and children report theirs. Only the
// output conversion belongs to the actor owner.
if (Exit.isFailure(converted)) {
yield* reportLifecycleFailure(converted.cause, {
reporters: errorReporters,
actorId: id,
generation: generation.current,
phase: "cleanup",
});
}
yield* SubscriptionRef.set(lifecycleRef, terminal);
yield* Deferred.succeed(terminalExitDeferred, terminal);
if (terminalCause !== undefined) {
Expand Down Expand Up @@ -1141,11 +1169,26 @@ export const createActor = Effect.fn("effect-machine.actor.spawn")(function* <
});
// Run recovery if lifecycle.recovery exists AND not hydrated (hydrate takes precedence)
if (lifecycle?.recovery !== undefined && !isHydrated) {
const resolved = yield* lifecycle.recovery.resolve({
actorId: id,
generation: generation.current,
machineInitial: options.machineInitial,
});
const recovery = lifecycle.recovery;
// `suspend` turns a resolver that throws into a defect that `onError` can see.
const resolved = yield* Effect.suspend(() =>
recovery.resolve({
actorId: id,
generation: generation.current,
machineInitial: options.machineInitial,
}),
).pipe(
// No generation has run yet, so no generation closure can report this failure.
// `onError` reports even when a stop is interrupting the start.
Effect.onError((cause) =>
reportLifecycleFailure(cause, {
reporters: errorReporters,
actorId: id,
generation: generation.current,
phase: "recovery",
}),
),
);
if (stopRequested) return yield* Effect.interrupt;
if (Option.isSome(resolved)) {
// Update cell stateRef
Expand Down Expand Up @@ -1176,8 +1219,23 @@ export const createActor = Effect.fn("effect-machine.actor.spawn")(function* <
spawnGeneration,
lifecycle,
onRestart: options.onRestart,
onTerminal: completeTerminal,
}),
errorReporters,
}).pipe(
// A restart step can fail before a new generation exists to report it. `onError`
// reports even when a stop is interrupting the supervisor.
Effect.onError((cause) =>
reportLifecycleFailure(cause, {
reporters: errorReporters,
actorId: id,
generation: generation.current,
phase: "restart",
}),
),
Effect.flatMap((terminalExit) => {
if (terminalExit === undefined) return Effect.void;
return completeTerminal(terminalExit);
}),
),
);
} else {
// No supervision — wire terminal exit from the current generation
Expand Down
51 changes: 48 additions & 3 deletions src/internal/runtime.ts
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ import {
Cause,
Deferred,
Effect,
ErrorReporter,
Exit,
Fiber,
Queue,
Expand Down Expand Up @@ -126,6 +127,40 @@ const RuntimeExit = {
}),
};

/** Spawn-time reporters and lifecycle owner context for one failure report. */
interface LifecycleFailureReport {
readonly reporters: ReadonlySet<ErrorReporter.ErrorReporter>;
readonly actorId: string;
readonly generation: number;
/**
* `restart` marks a supervision step that failed before a new generation existed.
* `recovery` marks a cold-start recovery that failed before the first generation ran.
*/
readonly phase: DefectPhase | "restart" | "recovery";
}

/**
* Report a failure that an actor lifecycle owner settles. The reporters come
* from the spawn context, not from whichever fiber closes the actor.
* @internal
*/
export const reportLifecycleFailure: {
(report: LifecycleFailureReport): (cause: Cause.Cause<unknown>) => Effect.Effect<void>;
(cause: Cause.Cause<unknown>, report: LifecycleFailureReport): Effect.Effect<void>;
} = dual(2, (cause: Cause.Cause<unknown>, report: LifecycleFailureReport): Effect.Effect<void> => {
if (report.reporters.size === 0 || Cause.hasInterruptsOnly(cause)) return Effect.void;
return ErrorReporter.report(cause).pipe(
Effect.annotateLogs({
"effect_machine.actor.id": report.actorId,
"effect_machine.actor.generation": report.generation,
"effect_machine.defect.phase": report.phase,
}),
Effect.provideService(ErrorReporter.CurrentErrorReporters, report.reporters),
// Reporters are host callbacks. A throwing reporter must not stop lifecycle settlement.
Effect.ignoreCause,
);
});

/** @internal */
export interface RuntimeHandle<S, E> {
/** Enqueue a fire-and-forget event */
Expand Down Expand Up @@ -247,6 +282,7 @@ export const createRuntime = Effect.fn("effect-machine.runtime.create")(function

// Capture services at allocation so delayed start and stop retain them.
const services = yield* Effect.context<R>();
const errorReporters = yield* ErrorReporter.CurrentErrorReporters;
const fork = Effect.runForkWith(services);

const { stateRef, latestTransitionRef, stoppedRef, eventQueue, listeners } = config.cellResources;
Expand Down Expand Up @@ -424,15 +460,24 @@ export const createRuntime = Effect.fn("effect-machine.runtime.create")(function
} else {
closed = Exit.failCause(cleanupCause);
}
let finalExit = selectedReason;
if (Exit.isFailure(closed)) {
let cause = closed.cause;
if (selectedReason._tag === "Defect") {
cause = Cause.combine(closed.cause)(selectedReason.cause);
}
yield* setExit(RuntimeExit.Defect(cause, "cleanup"));
} else {
yield* setExit(selectedReason);
finalExit = RuntimeExit.Defect(cause, "cleanup");
}
// Report the complete generation failure before any waiter can observe the exit.
if (finalExit._tag === "Defect") {
yield* reportLifecycleFailure(finalExit.cause, {
reporters: errorReporters,
actorId,
generation,
phase: finalExit.phase,
});
}
yield* setExit(finalExit);
yield* Deferred.succeed(closedDeferred, closed);
if (Exit.isFailure(closed)) {
return yield* Effect.failCause(closed.cause).pipe(Effect.orDie);
Expand Down
Loading
Loading