diff --git a/evals/tutor-quality/semantic-case-review-matrix.md b/evals/tutor-quality/semantic-case-review-matrix.md new file mode 100644 index 00000000..fd5f63ea --- /dev/null +++ b/evals/tutor-quality/semantic-case-review-matrix.md @@ -0,0 +1,18 @@ +# Stage-2B Semantic Case Review Matrix + +The cases below are reserved Stage-2B intents, not executable Semantic Corpus +records. They become executable only after the exact sketch, frozen question, +answer binding, ordered history, declared application context, allowed +evidence, and human-reviewed interpretation are pinned. All initial observed +cases are known/exposed development material; none is Gold or held out. + +| Case | Required pinned inputs and review | Current status | +| --- | --- | --- | +| `TQ-SEM-002` — `millis()` / `delay()` | Exact timing sketch; exact preceding question and correct answer; complete bounded turn order; phase/difficulty; reviewed timing facts and a human-defined range for a genuinely new non-blocking reasoning demand. | Reserved intent only; not executable. | +| `TQ-SEM-003` — `lastReport` | Exact declaration, assignments, and use; preceding question/answer/history; application context and difficulty; reviewed distinction between stored timestamp and elapsed duration. | Reserved intent only; not executable. | +| `TQ-SEM-004` — `unsigned long`, `millis()`, overflow | Exact type declaration and uses; ordered timing questions and answers; phase/Strategy/difficulty; reviewed return/type/wraparound facts and acceptable transfer beyond prior reasoning. | Reserved intent only; not executable. | +| `TQ-SEM-005` — matrix `int` versus `byte` | Exact matrix declaration, indexing, and use; preceding question/answer/history; Course Content/Strategy context or explicit absence; reviewed platform type/range facts and interpretation of the actual learner misconception. | Reserved intent only; not executable. | + +No case details are inferred from these names. Adding any case or changing its +evidence/reference is a Semantic Corpus version change. Candidate-specific +exposure remains a separate `CaseExposureRecord`, never a corpus field. diff --git a/evals/tutor-quality/semantic-corpus.yaml b/evals/tutor-quality/semantic-corpus.yaml new file mode 100644 index 00000000..6ff09d9f --- /dev/null +++ b/evals/tutor-quality/semantic-corpus.yaml @@ -0,0 +1,77 @@ +schemaVersion: tutor-quality-semantic-corpus-v1 +corpusId: tutor-quality-semantic +corpusVersion: 1 +digest: 00f7848de47e08547b86a5da0210ded6b3eb1e83f31a0ba34f7d56ad4205dac8 +cases: + - id: TQ-SEM-001 + caseVersion: 1 + digest: 04ba6e31a87e4595fc11b72c5a1f9ac260fa22fcd0d7186a6dcd4729d3e6a0b9 + purpose: Check that a substantively correct INPUT_PULLUP answer is recognized without false correction. + role: development + sketch: + reference: semantic-fixtures/TQ-SEM-001-input-pullup.ino + digest: e0ab46c8cadd3bbf5cab80b0237d8303918e152023de1793b2d1f40ae698d9c5 + frozenPreTurnContext: + schemaVersion: tutor-quality-frozen-context-v1 + sketchRef: semantic-fixtures/TQ-SEM-001-input-pullup.ino + sketchDigest: e0ab46c8cadd3bbf5cab80b0237d8303918e152023de1793b2d1f40ae698d9c5 + question: Welche logische Bedingung muss erfüllt sein, damit die LED eingeschaltet wird? + learnerAnswer: + text: buttonPin, also an PIN2 muss GND anliegen! + category: fully-correct + bindsToQuestion: Welche logische Bedingung muss erfüllt sein, damit die LED eingeschaltet wird? + priorDialog: [] + courseContent: + kind: free-tutor + difficulty: 30 + digest: f89c3200af9ff90d7ab282b8e90f511b50c6c8593b20745bc51a005639340e0a + expectedAnswerInterpretation: When the button connects buttonPin to GND, INPUT_PULLUP makes the input read LOW; this sketch writes HIGH to ledPin when digitalRead(buttonPin) == LOW. + acceptableTutorResponses: + diagnoses: + - Recognizes the learner's core condition as correct, allowing substantively equivalent wording. + feedbackApproaches: + - May clarify the active-low input without claiming the learner's answer is wrong. + followUps: + - May ask a sketch-grounded question that adds a distinct reasoning demand. + answerCategory: fully-correct + knownFailurePatterns: + - correct-answer-rejected + assessableDimensions: + - factual-correctness + - sketch-code-grounding + - learner-answer-diagnosis + - precision + dimensionEvidenceLimitations: + instructional-usefulness: No learner history beyond this synthetic answer is represented. + scaffolding: No learner history beyond this synthetic answer is represented. + dialogic-progression: The case has no preceding dialog turn. + non-repetition: The case has no preceding dialog turn. + difficulty-appropriateness: The case scopes any judgment to its fixed synthetic input only. + evidenceSources: + - id: sketch + kind: repository/sketch + reference: semantic-fixtures/TQ-SEM-001-input-pullup.ino + digest: e0ab46c8cadd3bbf5cab80b0237d8303918e152023de1793b2d1f40ae698d9c5 + - id: input-pullup-reference-draft + kind: draft-factual-reference + reference: TQ-SEM-001-factual-reference + digest: e0d1f86e092f980a8adf4b771cfc673d943a5492129f95dc850409a3123b5af0 + permittedExternalKnowledge: false + factualReferenceBundle: + id: TQ-SEM-001-factual-reference + version: 1 + sourceKind: draft-factual-reference + reviewStatus: pending-review + digest: e0d1f86e092f980a8adf4b771cfc673d943a5492129f95dc850409a3123b5af0 + provenance: Draft facts derived from the pinned sketch and requiring reviewer confirmation of board and wiring assumptions. + limitations: + - Confirm the target board, button-to-GND wiring, and LED_BUILTIN active-high behavior before Gold use. + facts: + - id: internal-pullup + statement: pinMode(buttonPin, INPUT_PULLUP) enables the microcontroller's internal pull-up for buttonPin. + - id: grounded-input-low + statement: With the declared button-to-GND wiring, pressing the button makes digitalRead(buttonPin) return LOW. + - id: sketch-condition + statement: The sketch writes HIGH to ledPin, which is LED_BUILTIN, when digitalRead(buttonPin) == LOW. + humanReference: + status: pending-review diff --git a/evals/tutor-quality/semantic-fixtures/TQ-SEM-001-input-pullup.ino b/evals/tutor-quality/semantic-fixtures/TQ-SEM-001-input-pullup.ino new file mode 100644 index 00000000..9b5806cd --- /dev/null +++ b/evals/tutor-quality/semantic-fixtures/TQ-SEM-001-input-pullup.ino @@ -0,0 +1,14 @@ +const int buttonPin = 2; +const int ledPin = LED_BUILTIN; + +void setup() { + pinMode(buttonPin, INPUT_PULLUP); + pinMode(ledPin, OUTPUT); +} + +void loop() { + digitalWrite( + ledPin, + digitalRead(buttonPin) == LOW ? HIGH : LOW + ); +} diff --git a/server/services/tutor/evaluation/semantic/case-exposure.ts b/server/services/tutor/evaluation/semantic/case-exposure.ts new file mode 100644 index 00000000..ff79c92e --- /dev/null +++ b/server/services/tutor/evaluation/semantic/case-exposure.ts @@ -0,0 +1,234 @@ +import { canonicalSemanticDigest, canonicalSemanticJson, deepFreeze, isSha256Digest } from "./semantic-canonical"; +import type { CaseExposureStatus } from "./semantic-types"; + +export type CaseExposureAvailability = "available" | "unavailable" | "unknown"; +export type CaseExposureArtifact = "caseDefinition" | "tutorOutput" | "humanReference" | "judgeResult"; + +export interface CaseExposureSnapshot { + readonly recordedAt: string; + readonly availableArtifacts: Readonly>; + readonly targetedChange: { + readonly status: "informed" | "not-informed" | "unknown"; + readonly rationale?: string; + }; +} + +export interface CaseExposureRecordInput { + readonly candidateIdentity: string; + readonly corpusId: string; + readonly corpusVersion: number; + readonly caseId: string; + readonly semanticCaseDigest: string; + readonly history: readonly CaseExposureSnapshot[]; +} + +export interface CaseExposureRecord extends CaseExposureRecordInput { + readonly schemaVersion: "tutor-quality-case-exposure-v1"; + readonly recordVersion: number; + readonly exposureStatus: CaseExposureStatus; + readonly identity: string; + readonly digest: string; + readonly previousRecordDigest?: string; +} + +export interface CaseExposureBinding { + readonly candidateIdentity: string; + readonly corpusId: string; + readonly corpusVersion: number; + readonly caseId: string; + readonly semanticCaseDigest: string; +} + +export type CaseExposureRecordValidation = + | { readonly valid: true; readonly record: CaseExposureRecord } + | { readonly valid: false; readonly reason: string }; + +const ARTIFACTS: readonly CaseExposureArtifact[] = ["caseDefinition", "tutorOutput", "humanReference", "judgeResult"]; +const OUTCOME_ARTIFACTS: readonly CaseExposureArtifact[] = ["tutorOutput", "humanReference", "judgeResult"]; + +function recordIdentity(input: CaseExposureRecordInput, recordVersion: number, previousRecordDigest?: string): string { + return canonicalSemanticDigest({ + schemaVersion: "tutor-quality-case-exposure-v1", + candidateIdentity: input.candidateIdentity, + corpusId: input.corpusId, + corpusVersion: input.corpusVersion, + caseId: input.caseId, + semanticCaseDigest: input.semanticCaseDigest, + recordVersion, + ...(previousRecordDigest ? { previousRecordDigest } : {}), + }); +} + +function exposureStatus(snapshot: CaseExposureSnapshot): CaseExposureStatus { + if (snapshot.targetedChange.status === "informed") return "used-for-targeted-change"; + if (snapshot.targetedChange.status === "unknown" || ARTIFACTS.some((artifact) => snapshot.availableArtifacts[artifact] === "unknown")) return "unknown"; + if (OUTCOME_ARTIFACTS.some((artifact) => snapshot.availableArtifacts[artifact] === "available")) return "outcome-exposed"; + if (snapshot.availableArtifacts.caseDefinition === "available") return "case-known"; + return "unexposed"; +} + +function validateSnapshotObject(snapshot: CaseExposureSnapshot, index: number): void { + if (snapshot === null || typeof snapshot !== "object") throw new Error(`Exposure history[${index}] must be an object`); + if (Object.keys(snapshot).some((key) => !["recordedAt", "availableArtifacts", "targetedChange"].includes(key))) throw new Error(`Exposure history[${index}] contains an unsupported field`); +} + +function validateSnapshotTimestamp(snapshot: CaseExposureSnapshot, index: number, previous?: CaseExposureSnapshot): void { + if (typeof snapshot.recordedAt !== "string" || !/^\d{4}-\d\d-\d\dT\d\d:\d\d:\d\d(?:\.\d{3})?Z$/.test(snapshot.recordedAt) || Number.isNaN(Date.parse(snapshot.recordedAt))) { + throw new Error(`Exposure history[${index}] recordedAt must be an explicit UTC timestamp`); + } + if (previous && Date.parse(snapshot.recordedAt) <= Date.parse(previous.recordedAt)) throw new Error("Exposure history timestamps must be strictly increasing"); +} + +function validateSnapshotArtifacts(snapshot: CaseExposureSnapshot, index: number): void { + if (!snapshot.availableArtifacts || Object.keys(snapshot.availableArtifacts).length !== ARTIFACTS.length || ARTIFACTS.some((artifact) => Object.getOwnPropertyDescriptor(snapshot.availableArtifacts, artifact) === undefined || !["available", "unavailable", "unknown"].includes(snapshot.availableArtifacts[artifact]))) { + throw new Error(`Exposure history[${index}] must record availability for every artifact`); + } +} + +function validateSnapshotTargetedChange(snapshot: CaseExposureSnapshot, index: number): void { + if (!snapshot.targetedChange || !["informed", "not-informed", "unknown"].includes(snapshot.targetedChange.status)) { + throw new Error(`Exposure history[${index}] targeted-change state is invalid`); + } + if (Object.keys(snapshot.targetedChange).some((key) => !["status", "rationale"].includes(key))) throw new Error(`Exposure history[${index}] targeted-change state contains an unsupported field`); + if (snapshot.targetedChange.status === "informed" && (!snapshot.targetedChange.rationale || snapshot.targetedChange.rationale.trim().length === 0)) { + throw new Error(`Exposure history[${index}] requires the rationale for a targeted change`); + } + if (snapshot.targetedChange.status !== "informed" && snapshot.targetedChange.rationale !== undefined) { + throw new Error(`Exposure history[${index}] cannot attach a targeted-change rationale to a non-targeted state`); + } +} + +function validateSnapshotHistory(snapshot: CaseExposureSnapshot, previous?: CaseExposureSnapshot): void { + if (previous) { + for (const artifact of ARTIFACTS) { + if (previous.availableArtifacts[artifact] === "available" && snapshot.availableArtifacts[artifact] !== "available") { + throw new Error(`Exposure history cannot erase known availability of ${artifact}`); + } + } + if (previous.targetedChange.status === "informed" && snapshot.targetedChange.status !== "informed") { + throw new Error("Exposure history cannot erase a previously recorded targeted change"); + } + } +} + +function validateSnapshot(snapshot: CaseExposureSnapshot, index: number, previous?: CaseExposureSnapshot): void { + validateSnapshotObject(snapshot, index); + validateSnapshotTimestamp(snapshot, index, previous); + validateSnapshotArtifacts(snapshot, index); + validateSnapshotTargetedChange(snapshot, index); + validateSnapshotHistory(snapshot, previous); +} + +function validateInput(input: CaseExposureRecordInput, record = false): void { + const allowedKeys = record + ? ["schemaVersion", "candidateIdentity", "corpusId", "corpusVersion", "caseId", "semanticCaseDigest", "history", "recordVersion", "exposureStatus", "identity", "digest", "previousRecordDigest"] + : ["candidateIdentity", "corpusId", "corpusVersion", "caseId", "semanticCaseDigest", "history"]; + if (input === null || typeof input !== "object" || Object.keys(input).some((key) => !allowedKeys.includes(key))) throw new Error("CaseExposureRecord input contains an unsupported field"); + if (!isSha256Digest(input.candidateIdentity)) throw new Error("Candidate identity must be a lowercase SHA-256 digest"); + if (!isSha256Digest(input.semanticCaseDigest)) throw new Error("Semantic Case digest must be a lowercase SHA-256 digest"); + if (!/^[a-z0-9][a-z0-9-]{0,63}$/.test(input.corpusId)) throw new Error("Corpus ID is invalid"); + if (!Number.isInteger(input.corpusVersion) || input.corpusVersion < 1) throw new Error("Corpus version must be a positive integer"); + if (!/^[A-Z0-9][A-Z0-9-]{0,63}$/.test(input.caseId)) throw new Error("Semantic Case ID is invalid"); + if (!Array.isArray(input.history) || input.history.length === 0) throw new Error("Exposure history must contain at least one version snapshot"); + input.history.forEach((snapshot, index) => validateSnapshot(snapshot, index, input.history[index - 1])); + const versions = input.history.length; + if (versions > 10_000) throw new Error("Exposure history exceeds the bounded record limit"); +} + +function sameBinding(left: CaseExposureBinding, right: CaseExposureBinding): boolean { + return left.candidateIdentity === right.candidateIdentity + && left.corpusId === right.corpusId + && left.corpusVersion === right.corpusVersion + && left.caseId === right.caseId + && left.semanticCaseDigest === right.semanticCaseDigest; +} + +function recordDigestContent(record: Omit | CaseExposureRecord): Record { + const { digest: _digest, ...source } = record as CaseExposureRecord; + return source; +} + +function buildCaseExposureRecord(input: CaseExposureRecordInput, previousRecordDigest?: string): CaseExposureRecord { + const recordVersion = input.history.length; + const base = { + schemaVersion: "tutor-quality-case-exposure-v1" as const, + ...input, + history: input.history.map((snapshot) => ({ + ...snapshot, + availableArtifacts: { ...snapshot.availableArtifacts }, + targetedChange: { ...snapshot.targetedChange }, + })), + recordVersion, + exposureStatus: exposureStatus(input.history.at(-1)!), + identity: recordIdentity(input, recordVersion, previousRecordDigest), + ...(previousRecordDigest ? { previousRecordDigest } : {}), + }; + return { ...base, digest: canonicalSemanticDigest(base) }; +} + +function rebuildExposureHistoryPrefix(input: CaseExposureRecordInput): CaseExposureRecord { + let previous: CaseExposureRecord | undefined; + for (let end = 1; end <= input.history.length; end += 1) { + const prefix: CaseExposureRecordInput = { + candidateIdentity: input.candidateIdentity, + corpusId: input.corpusId, + corpusVersion: input.corpusVersion, + caseId: input.caseId, + semanticCaseDigest: input.semanticCaseDigest, + history: input.history.slice(0, end), + }; + previous = buildCaseExposureRecord(prefix, previous?.digest); + } + if (!previous) throw new Error("Exposure history prefix must not be empty"); + return previous; +} + +export function createCaseExposureRecord(input: CaseExposureRecordInput, previous?: CaseExposureRecord): CaseExposureRecord { + validateInput(input); + const recordVersion = input.history.length; + if (previous) { + if (!sameBinding(input, previous)) throw new Error("CaseExposureRecord revision cannot change Candidate or Semantic Case binding"); + const priorValidation = validateCaseExposureRecord(previous, input); + if (!priorValidation.valid) throw new Error(`CaseExposureRecord revision requires a valid prior record: ${priorValidation.reason}`); + if (recordVersion !== previous.recordVersion + 1) throw new Error("CaseExposureRecord revision must append exactly one new history snapshot"); + if (canonicalSemanticJson(input.history.slice(0, previous.history.length)) !== canonicalSemanticJson(previous.history)) { + throw new Error("CaseExposureRecord revision must append to prior history without rewriting it"); + } + } else if (recordVersion !== 1) { + throw new Error("Initial CaseExposureRecord must start at version 1"); + } + const previousRecordDigest = previous?.digest; + if (previousRecordDigest !== undefined && !isSha256Digest(previousRecordDigest)) throw new Error("Previous CaseExposureRecord digest is invalid"); + return deepFreeze(buildCaseExposureRecord(input, previousRecordDigest)); +} + +export function validateCaseExposureRecord(input: unknown, expected: CaseExposureBinding): CaseExposureRecordValidation { + try { + if (input === null || typeof input !== "object" || Array.isArray(input)) return { valid: false, reason: "record-must-be-object" }; + const record = input as CaseExposureRecord; + validateInput(record, true); + if (record.schemaVersion !== "tutor-quality-case-exposure-v1") return { valid: false, reason: "unsupported-record-schema" }; + if (record.recordVersion !== record.history.length) return { valid: false, reason: "record-version-history-mismatch" }; + if (!sameBinding(record, expected)) return { valid: false, reason: "candidate-or-case-binding-mismatch" }; + if (record.exposureStatus !== exposureStatus(record.history.at(-1)!)) return { valid: false, reason: "exposure-status-mismatch" }; + if (record.recordVersion === 1 ? record.previousRecordDigest !== undefined : !isSha256Digest(record.previousRecordDigest)) { + return { valid: false, reason: "previous-record-digest-history-mismatch" }; + } + if (record.recordVersion > 1) { + const prefix: CaseExposureRecordInput = { + candidateIdentity: record.candidateIdentity, + corpusId: record.corpusId, + corpusVersion: record.corpusVersion, + caseId: record.caseId, + semanticCaseDigest: record.semanticCaseDigest, + history: record.history.slice(0, -1), + }; + if (rebuildExposureHistoryPrefix(prefix).digest !== record.previousRecordDigest) return { valid: false, reason: "previous-record-digest-history-mismatch" }; + } + if (!isSha256Digest(record.identity) || record.identity !== recordIdentity(record, record.recordVersion, record.previousRecordDigest)) return { valid: false, reason: "record-identity-mismatch" }; + if (!isSha256Digest(record.digest) || record.digest !== canonicalSemanticDigest(recordDigestContent(record))) return { valid: false, reason: "record-digest-mismatch" }; + return { valid: true, record }; + } catch (error) { + return { valid: false, reason: error instanceof Error ? error.message : "invalid-case-exposure-record" }; + } +} diff --git a/server/services/tutor/evaluation/semantic/frozen-context.ts b/server/services/tutor/evaluation/semantic/frozen-context.ts new file mode 100644 index 00000000..fd16d36d --- /dev/null +++ b/server/services/tutor/evaluation/semantic/frozen-context.ts @@ -0,0 +1,85 @@ +import { tutorDialogTurnSchema, type TutorDialogTurn } from "@shared/tutor"; +import { z } from "zod"; +import { canonicalSemanticDigest, canonicalSemanticJson, deepFreeze, isSha256Digest } from "./semantic-canonical"; + +const courseContentContextSchema = z.discriminatedUnion("kind", [ + z.object({ kind: z.literal("free-tutor") }).strict(), + z.object({ + kind: z.literal("repository-course-content"), + reference: z.string().min(1), + revision: z.string().min(1), + digest: z.string().regex(/^[0-9a-f]{64}$/), + }).strict(), +]); + +const frozenContextSchema = z.object({ + schemaVersion: z.literal("tutor-quality-frozen-context-v1"), + sketchRef: z.string().min(1), + sketchDigest: z.string().regex(/^[0-9a-f]{64}$/), + question: z.string().min(1), + learnerAnswer: z.object({ + text: z.string().min(1), + category: z.enum(["fully-correct", "partially-correct", "typical-misconception", "terminology-confusion", "correct-poorly-phrased", "explicitly-unknown", "off-topic", "unexpectedly-strong"]), + bindsToQuestion: z.string().min(1), + }).strict(), + priorDialog: z.array(z.unknown()), + courseContent: courseContentContextSchema, + difficulty: z.number().int().min(1).max(100), + digest: z.string().regex(/^[0-9a-f]{64}$/).optional(), +}).strict(); + +export type FrozenPreTurnContextInput = Omit, "digest"> & { readonly digest?: string }; +export interface FrozenPreTurnContext extends Omit, "priorDialog" | "digest"> { + readonly priorDialog: readonly TutorDialogTurn[]; + readonly digest: string; +} + +export type FrozenPreTurnContextValidation = + | { readonly valid: true; readonly context: FrozenPreTurnContext } + | { readonly valid: false; readonly reason: string }; + +function parseInput(input: unknown): FrozenPreTurnContextInput { + const parsed = frozenContextSchema.safeParse(input); + if (!parsed.success) { + const details = parsed.error.issues.map(({ path, message }) => `${path.join(".")}: ${message}`).join("; "); + throw new Error(`Invalid Frozen Pre-Turn Context: ${details}`); + } + if (parsed.data.question !== parsed.data.learnerAnswer.bindsToQuestion) { + throw new Error("Invalid Frozen Pre-Turn Context: learner answer must bind to the exact preceding question"); + } + const priorDialog: TutorDialogTurn[] = []; + for (const [index, rawTurn] of parsed.data.priorDialog.entries()) { + const turn = tutorDialogTurnSchema.safeParse(rawTurn); + if (!turn.success || canonicalSemanticJson(turn.data) !== canonicalSemanticJson(rawTurn)) { + throw new Error(`Invalid Frozen Pre-Turn Context: priorDialog[${index}] is invalid or would require normalization`); + } + priorDialog.push(turn.data); + } + return { ...parsed.data, priorDialog } as FrozenPreTurnContextInput; +} + +function digestInput(input: FrozenPreTurnContextInput): Record { + const { digest: _digest, ...source } = input; + return source; +} + +export function createFrozenPreTurnContext(input: FrozenPreTurnContextInput): FrozenPreTurnContext { + const parsed = parseInput(input); + const digest = canonicalSemanticDigest(digestInput(parsed)); + if (parsed.digest !== undefined && parsed.digest !== digest) { + throw new Error("Invalid Frozen Pre-Turn Context: digest does not match its canonical content"); + } + if (!isSha256Digest(digest)) throw new Error("Invalid Frozen Pre-Turn Context: unable to produce a valid digest"); + return deepFreeze({ ...parsed, digest }) as FrozenPreTurnContext; +} + +export function validateFrozenPreTurnContext(input: unknown): FrozenPreTurnContextValidation { + try { + if (input === null || typeof input !== "object" || Array.isArray(input) || !isSha256Digest((input as { digest?: unknown }).digest)) { + return { valid: false, reason: "missing-or-invalid-context-digest" }; + } + return { valid: true, context: createFrozenPreTurnContext(input as FrozenPreTurnContextInput) }; + } catch (error) { + return { valid: false, reason: error instanceof Error ? error.message : "invalid-frozen-context" }; + } +} diff --git a/server/services/tutor/evaluation/semantic/semantic-canonical.ts b/server/services/tutor/evaluation/semantic/semantic-canonical.ts new file mode 100644 index 00000000..1b3d63c9 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/semantic-canonical.ts @@ -0,0 +1,54 @@ +import { createHash } from "node:crypto"; + +function compareUtf8(left: string, right: string): number { + return Buffer.compare(Buffer.from(left, "utf8"), Buffer.from(right, "utf8")); +} + +function serializeCanonical(value: unknown, active: Set): string { + if (value === null) return "null"; + if (typeof value === "string" || typeof value === "boolean") return JSON.stringify(value); + if (typeof value === "number") { + if (!Number.isFinite(value)) throw new TypeError("Canonical JSON does not allow non-finite numbers"); + return JSON.stringify(value); + } + if (value === undefined) return "null"; + if (typeof value !== "object") throw new TypeError("Canonical JSON accepts only JSON values"); + if (active.has(value)) throw new TypeError("Canonical JSON does not allow cyclic values"); + active.add(value); + try { + if (Array.isArray(value)) return `[${Array.from(value, (entry) => serializeCanonical(entry, active)).join(",")}]`; + const prototype = Object.getPrototypeOf(value); + if (prototype !== Object.prototype && prototype !== null) throw new TypeError("Canonical JSON accepts only plain objects"); + const object = value as Record; + const entries = Object.keys(object) + .filter((key) => object[key] !== undefined) + .sort(compareUtf8) + .map((key) => `${JSON.stringify(key)}:${serializeCanonical(object[key], active)}`); + return `{${entries.join(",")}}`; + } finally { + active.delete(value); + } +} + +export function canonicalSemanticJson(value: unknown): string { + return serializeCanonical(value, new Set()); +} + +export function semanticSha256(value: string | Uint8Array): string { + return createHash("sha256").update(value).digest("hex"); +} + +export function canonicalSemanticDigest(value: unknown): string { + return semanticSha256(canonicalSemanticJson(value)); +} + +export function isSha256Digest(value: unknown): value is string { + return typeof value === "string" && /^[0-9a-f]{64}$/.test(value); +} + +export function deepFreeze(value: T): T { + if (value === null || typeof value !== "object") return value; + for (const child of Object.values(value as Record)) deepFreeze(child); + if (!Object.isFrozen(value)) Object.freeze(value); + return value; +} diff --git a/server/services/tutor/evaluation/semantic/semantic-corpus.ts b/server/services/tutor/evaluation/semantic/semantic-corpus.ts new file mode 100644 index 00000000..1c6f7644 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/semantic-corpus.ts @@ -0,0 +1,370 @@ +import { z } from "zod"; +import type { TutorPlanningContentContext } from "../../tutor-planning"; +import { canonicalSemanticDigest, semanticSha256 } from "./semantic-canonical"; +import { validateFrozenPreTurnContext, type FrozenPreTurnContext } from "./frozen-context"; +import { + SEMANTIC_ANSWER_CATEGORIES, + SEMANTIC_CASE_ROLES, + SEMANTIC_EVIDENCE_SOURCE_KINDS, + SEMANTIC_RUBRIC_DIMENSIONS, + type SemanticAnswerCategory, + type SemanticCaseRole, + type SemanticEvidenceSource, + type SemanticRubricDimension, +} from "./semantic-types"; + +const digestSchema = z.string().regex(/^[0-9a-f]{64}$/); +const evidenceSourceSchema = z.object({ + id: z.string().min(1), + kind: z.enum(SEMANTIC_EVIDENCE_SOURCE_KINDS), + reference: z.string().min(1), + digest: digestSchema, +}).strict(); + +const factualReferenceBundleSchema = z.object({ + id: z.string().min(1), + version: z.number().int().min(1), + sourceKind: z.enum(["draft-factual-reference", "reviewed-factual-reference"]), + reviewStatus: z.enum(["pending-review", "reviewed"]), + provenance: z.string().min(1), + limitations: z.array(z.string().min(1)), + facts: z.array(z.object({ id: z.string().min(1), statement: z.string().min(1) }).strict()).min(1), + digest: digestSchema.optional(), +}).strict(); + +const acceptableResponsesSchema = z.object({ + diagnoses: z.array(z.string().min(1)).min(1), + feedbackApproaches: z.array(z.string().min(1)).min(1), + followUps: z.array(z.string().min(1)).min(1), +}).strict(); + +const semanticCaseSchema = z.object({ + id: z.string().regex(/^[A-Z0-9][A-Z0-9-]{0,63}$/), + caseVersion: z.number().int().min(1), + digest: digestSchema.optional(), + purpose: z.string().min(1), + role: z.enum(SEMANTIC_CASE_ROLES), + sketch: z.object({ reference: z.string().min(1), digest: digestSchema }).strict(), + frozenPreTurnContext: z.unknown(), + expectedAnswerInterpretation: z.string().min(1), + acceptableTutorResponses: acceptableResponsesSchema, + answerCategory: z.enum(SEMANTIC_ANSWER_CATEGORIES), + knownFailurePatterns: z.array(z.string().min(1)).min(1), + assessableDimensions: z.array(z.enum(SEMANTIC_RUBRIC_DIMENSIONS)).min(1), + dimensionEvidenceLimitations: z.record(z.string(), z.string().min(1)), + evidenceSources: z.array(evidenceSourceSchema).min(1), + permittedExternalKnowledge: z.boolean(), + factualReferenceBundle: factualReferenceBundleSchema, + humanReference: z.object({ status: z.literal("pending-review") }).strict(), +}).strict(); + +const corpusSchema = z.object({ + schemaVersion: z.literal("tutor-quality-semantic-corpus-v1"), + corpusId: z.string().regex(/^[a-z0-9][a-z0-9-]{0,63}$/), + corpusVersion: z.number().int().min(1), + digest: digestSchema.optional(), + cases: z.array(semanticCaseSchema).min(1), +}).strict(); + +export type SemanticCorpusReferences = { + readonly sketches: ReadonlyMap; + readonly courseContent: ReadonlyMap; +}; + +export interface SemanticFactualReferenceBundle { + readonly id: string; + readonly version: number; + readonly sourceKind: "draft-factual-reference" | "reviewed-factual-reference"; + readonly reviewStatus: "pending-review" | "reviewed"; + readonly provenance: string; + readonly limitations: readonly string[]; + readonly facts: readonly { readonly id: string; readonly statement: string }[]; + readonly digest: string; +} + +export interface SemanticCase { + readonly corpusId: string; + readonly corpusVersion: number; + readonly id: string; + readonly caseVersion: number; + readonly digest: string; + readonly purpose: string; + readonly role: SemanticCaseRole; + readonly sketch: { readonly reference: string; readonly digest: string }; + readonly frozenPreTurnContext: FrozenPreTurnContext; + readonly expectedAnswerInterpretation: string; + readonly acceptableTutorResponses: { + readonly diagnoses: readonly string[]; + readonly feedbackApproaches: readonly string[]; + readonly followUps: readonly string[]; + }; + readonly answerCategory: SemanticAnswerCategory; + readonly knownFailurePatterns: readonly string[]; + readonly assessableDimensions: readonly SemanticRubricDimension[]; + readonly dimensionEvidenceLimitations: Readonly>; + readonly evidenceSources: readonly SemanticEvidenceSource[]; + readonly permittedExternalKnowledge: boolean; + readonly factualReferenceBundle: SemanticFactualReferenceBundle; + readonly humanReference: { readonly status: "pending-review" | "reviewed" | "disputed" | "adjudicated" }; +} + +export interface SemanticCorpus { + readonly schemaVersion: "tutor-quality-semantic-corpus-v1"; + readonly corpusId: string; + readonly corpusVersion: number; + readonly cases: readonly SemanticCase[]; + readonly digest: string; +} + +export type SemanticCorpusEvolutionComparison = + | { readonly valid: true } + | { readonly valid: false; readonly reason: "version-not-increased" | "version-regressed" }; + +export function factualReferenceBundleDigest(bundle: Omit | Record): string { + const { digest: _digest, ...source } = bundle as Record; + const facts = Array.isArray(source.facts) + ? [...source.facts as Array>].sort((left, right) => compareUtf8(String(left.id), String(right.id))) + : source.facts; + const limitations = Array.isArray(source.limitations) ? [...source.limitations as string[]].sort(compareUtf8) : source.limitations; + return canonicalSemanticDigest({ ...source, ...(facts ? { facts } : {}), ...(limitations ? { limitations } : {}) }); +} + +function caseContentWithoutDigest(semanticCase: Omit | SemanticCase | Record): Record { + const { digest: _digest, corpusId: _corpusId, corpusVersion: _corpusVersion, ...source } = semanticCase as Record; + return source; +} + +function normalizedCase(semanticCase: SemanticCase): Record { + const source = caseContentWithoutDigest(semanticCase) as Omit; + return { + ...source, + acceptableTutorResponses: { + diagnoses: [...source.acceptableTutorResponses.diagnoses].sort(compareUtf8), + feedbackApproaches: [...source.acceptableTutorResponses.feedbackApproaches].sort(compareUtf8), + followUps: [...source.acceptableTutorResponses.followUps].sort(compareUtf8), + }, + knownFailurePatterns: [...source.knownFailurePatterns].sort(compareUtf8), + assessableDimensions: [...source.assessableDimensions].sort(compareUtf8), + evidenceSources: [...source.evidenceSources].sort((left, right) => compareUtf8(left.id, right.id)), + factualReferenceBundle: { + ...source.factualReferenceBundle, + digest: factualReferenceBundleDigest(source.factualReferenceBundle), + limitations: [...source.factualReferenceBundle.limitations].sort(compareUtf8), + facts: [...source.factualReferenceBundle.facts].sort((left, right) => compareUtf8(left.id, right.id)), + }, + }; +} + +function compareUtf8(left: string, right: string): number { + return Buffer.compare(Buffer.from(left, "utf8"), Buffer.from(right, "utf8")); +} + +function normalizedCaseSource(caseItem: Record): Record { + const source = caseContentWithoutDigest(caseItem); + const acceptable = source.acceptableTutorResponses as Record | undefined; + const bundle = source.factualReferenceBundle as Record | undefined; + return { + ...source, + ...(acceptable ? { + acceptableTutorResponses: Object.fromEntries(Object.entries(acceptable).map(([key, value]) => [ + key, + Array.isArray(value) ? [...value as string[]].sort(compareUtf8) : value, + ])), + } : {}), + ...(Array.isArray(source.knownFailurePatterns) ? { knownFailurePatterns: [...source.knownFailurePatterns as string[]].sort(compareUtf8) } : {}), + ...(Array.isArray(source.assessableDimensions) ? { assessableDimensions: [...source.assessableDimensions as string[]].sort(compareUtf8) } : {}), + ...(Array.isArray(source.evidenceSources) ? { + evidenceSources: [...source.evidenceSources as Array>].sort((left, right) => compareUtf8(String(left.id), String(right.id))), + } : {}), + ...(bundle ? { + factualReferenceBundle: { + ...bundle, + digest: factualReferenceBundleDigest(bundle), + ...(Array.isArray(bundle.limitations) ? { limitations: [...bundle.limitations as string[]].sort(compareUtf8) } : {}), + ...(Array.isArray(bundle.facts) ? { + facts: [...bundle.facts as Array>].sort((left, right) => compareUtf8(String(left.id), String(right.id))), + } : {}), + }, + } : {}), + }; +} + +export function semanticCorpusDigest(corpus: unknown): string { + if (corpus === null || typeof corpus !== "object" || Array.isArray(corpus)) throw new TypeError("Semantic Corpus must be an object"); + const source = corpus as Record; + if (!Array.isArray(source.cases)) throw new TypeError("Semantic Corpus cases must be an array"); + const cases = source.cases.map((item) => { + if (item === null || typeof item !== "object" || Array.isArray(item)) throw new TypeError("Semantic Corpus case must be an object"); + const caseItem = item as Record; + return normalizedCaseSource(caseItem); + }); + return canonicalSemanticDigest({ + schemaVersion: source.schemaVersion, + corpusId: source.corpusId, + corpusVersion: source.corpusVersion, + cases: cases.toSorted((left, right) => compareUtf8(String(left.id), String(right.id))), + }); +} + +function semanticCorpusContentDigest(corpus: SemanticCorpus): string { + return canonicalSemanticDigest({ + schemaVersion: corpus.schemaVersion, + corpusId: corpus.corpusId, + cases: [...corpus.cases].map(normalizedCase).sort((left, right) => compareUtf8(String(left.id), String(right.id))), + }); +} + +export function semanticCaseDigest(semanticCase: SemanticCase): string { + return canonicalSemanticDigest(normalizedCase(semanticCase)); +} + +function invalidCorpus(message: string): never { + throw new Error(`Invalid Stage-2B Semantic Corpus: ${message}`); +} + +function assertUnique(values: readonly string[], label: string): void { + if (new Set(values).size !== values.length) invalidCorpus(`${label} must be duplicate-free`); +} + +function checkSketchReferences(semanticCase: SemanticCase, references: SemanticCorpusReferences): void { + if (semanticCase.frozenPreTurnContext.sketchRef !== semanticCase.sketch.reference + || semanticCase.frozenPreTurnContext.sketchDigest !== semanticCase.sketch.digest) { + invalidCorpus(`${semanticCase.id} Frozen Pre-Turn Context must bind to the exact Semantic Case sketch and digest`); + } + const sketch = references.sketches.get(semanticCase.sketch.reference); + if (sketch === undefined) invalidCorpus(`${semanticCase.id} references missing sketch ${semanticCase.sketch.reference}`); + if (semanticSha256(sketch) !== semanticCase.sketch.digest) invalidCorpus(`${semanticCase.id} sketch digest does not match fixture bytes`); + const sourceSketch = semanticCase.evidenceSources.find(({ id }) => id === "sketch"); + if (sourceSketch?.kind !== "repository/sketch" || sourceSketch?.reference !== semanticCase.sketch.reference || sourceSketch?.digest !== semanticCase.sketch.digest) { + invalidCorpus(`${semanticCase.id} must identify its exact sketch as repository/sketch evidence`); + } +} + +function checkFactualReference(semanticCase: SemanticCase): void { + const factualSource = semanticCase.evidenceSources.find(({ kind }) => kind === semanticCase.factualReferenceBundle.sourceKind); + if (factualSource?.reference !== semanticCase.factualReferenceBundle.id || factualSource?.digest !== semanticCase.factualReferenceBundle.digest) { + invalidCorpus(`${semanticCase.id} factual reference source must bind to its bundle identity and digest`); + } + if (semanticCase.factualReferenceBundle.sourceKind === "draft-factual-reference" && semanticCase.factualReferenceBundle.reviewStatus !== "pending-review") { + invalidCorpus(`${semanticCase.id} draft factual references must remain pending review`); + } + if (semanticCase.factualReferenceBundle.sourceKind === "reviewed-factual-reference" && semanticCase.factualReferenceBundle.reviewStatus !== "reviewed") { + invalidCorpus(`${semanticCase.id} reviewed factual references must be marked reviewed`); + } +} + +function checkExternalEvidencePolicy(semanticCase: SemanticCase): void { + const hasExternalSource = semanticCase.evidenceSources.some(({ kind }) => kind === "permitted-external-knowledge"); + if (hasExternalSource !== semanticCase.permittedExternalKnowledge) invalidCorpus(`${semanticCase.id} external evidence and evidence policy disagree`); +} + +function checkCourseContentReferences(semanticCase: SemanticCase, references: SemanticCorpusReferences): void { + const courseSources = semanticCase.evidenceSources.filter(({ kind }) => kind === "repository/course-content"); + if (semanticCase.frozenPreTurnContext.courseContent.kind === "repository-course-content") { + if (courseSources.length !== 1 + || courseSources[0].reference !== semanticCase.frozenPreTurnContext.courseContent.reference + || courseSources[0].digest !== semanticCase.frozenPreTurnContext.courseContent.digest) { + invalidCorpus(`${semanticCase.id} repository Course Content evidence must bind to its exact context reference and digest`); + } + } else if (courseSources.length > 0) { + invalidCorpus(`${semanticCase.id} cannot claim repository Course Content evidence without a Course Content context`); + } + if (semanticCase.frozenPreTurnContext.courseContent.kind === "repository-course-content") { + const content = references.courseContent.get(semanticCase.frozenPreTurnContext.courseContent.reference); + if (!content) invalidCorpus(`${semanticCase.id} references missing Course Content ${semanticCase.frozenPreTurnContext.courseContent.reference}`); + if (content.revision !== semanticCase.frozenPreTurnContext.courseContent.revision) invalidCorpus(`${semanticCase.id} Course Content revision does not match its fixture`); + if (canonicalSemanticDigest(content) !== semanticCase.frozenPreTurnContext.courseContent.digest) invalidCorpus(`${semanticCase.id} Course Content digest does not match its fixture`); + } +} + +function checkCaseReferences(semanticCase: SemanticCase, references: SemanticCorpusReferences): void { + checkSketchReferences(semanticCase, references); + checkFactualReference(semanticCase); + checkExternalEvidencePolicy(semanticCase); + checkCourseContentReferences(semanticCase, references); +} + +export function parseSemanticCorpus(input: unknown, references: SemanticCorpusReferences): SemanticCorpus { + const parsed = corpusSchema.safeParse(input); + if (!parsed.success) invalidCorpus(parsed.error.issues.map(({ path, message }) => `${path.join(".")}: ${message}`).join("; ")); + const ids = parsed.data.cases.map(({ id }) => id); + assertUnique(ids, "case IDs"); + const cases = parsed.data.cases.map((rawCase) => { + const context = validateFrozenPreTurnContext(rawCase.frozenPreTurnContext); + if (!context.valid) invalidCorpus(`${rawCase.id} has invalid Frozen Pre-Turn Context: ${context.reason}`); + const bundleInput = rawCase.factualReferenceBundle; + const facts = [...bundleInput.facts].sort((left, right) => compareUtf8(left.id, right.id)); + assertUnique(facts.map(({ id }) => id), `${rawCase.id} factual-reference fact IDs`); + assertUnique(bundleInput.limitations, `${rawCase.id} factual-reference limitations`); + const bundleSource = { + id: bundleInput.id, + version: bundleInput.version, + sourceKind: bundleInput.sourceKind, + reviewStatus: bundleInput.reviewStatus, + provenance: bundleInput.provenance, + limitations: [...bundleInput.limitations].sort(compareUtf8), + facts, + }; + const bundleDigest = factualReferenceBundleDigest(bundleSource); + if (bundleInput.digest !== undefined && bundleInput.digest !== bundleDigest) invalidCorpus(`${rawCase.id} factual reference digest does not match bundle content`); + const dimensions = rawCase.assessableDimensions; + assertUnique(dimensions, `${rawCase.id} assessable dimensions`); + const limitationKeys = Object.keys(rawCase.dimensionEvidenceLimitations); + for (const key of limitationKeys) { + if (!(SEMANTIC_RUBRIC_DIMENSIONS as readonly string[]).includes(key)) invalidCorpus(`${rawCase.id} declares an unknown rubric dimension limitation ${key}`); + } + for (const dimension of SEMANTIC_RUBRIC_DIMENSIONS) { + if (!dimensions.includes(dimension) && !limitationKeys.includes(dimension)) invalidCorpus(`${rawCase.id} must explain why ${dimension} is not assessable`); + if (dimensions.includes(dimension) && limitationKeys.includes(dimension)) invalidCorpus(`${rawCase.id} cannot both assess and declare an evidence limitation for ${dimension}`); + } + assertUnique(rawCase.knownFailurePatterns, `${rawCase.id} known failure patterns`); + assertUnique(rawCase.acceptableTutorResponses.diagnoses, `${rawCase.id} acceptable diagnoses`); + assertUnique(rawCase.acceptableTutorResponses.feedbackApproaches, `${rawCase.id} acceptable feedback approaches`); + assertUnique(rawCase.acceptableTutorResponses.followUps, `${rawCase.id} acceptable follow-ups`); + assertUnique(rawCase.evidenceSources.map(({ id }) => id), `${rawCase.id} evidence source IDs`); + const { digest: declaredCaseDigest, ...rawCaseContent } = rawCase; + const semanticCaseWithoutDigest: Omit = { + ...rawCaseContent, + sketch: { ...rawCase.sketch }, + frozenPreTurnContext: context.context, + acceptableTutorResponses: { + diagnoses: [...rawCase.acceptableTutorResponses.diagnoses].sort(compareUtf8), + feedbackApproaches: [...rawCase.acceptableTutorResponses.feedbackApproaches].sort(compareUtf8), + followUps: [...rawCase.acceptableTutorResponses.followUps].sort(compareUtf8), + }, + knownFailurePatterns: [...rawCase.knownFailurePatterns].sort(compareUtf8), + assessableDimensions: [...dimensions].sort(compareUtf8), + dimensionEvidenceLimitations: { ...rawCase.dimensionEvidenceLimitations }, + evidenceSources: [...rawCase.evidenceSources].sort((left, right) => compareUtf8(left.id, right.id)), + factualReferenceBundle: { ...bundleSource, digest: bundleDigest }, + humanReference: { ...rawCase.humanReference }, + }; + const caseDigest = canonicalSemanticDigest(semanticCaseWithoutDigest); + if (declaredCaseDigest !== undefined && declaredCaseDigest !== caseDigest) invalidCorpus(`${rawCase.id} case digest does not match content`); + const semanticCase: SemanticCase = { + ...semanticCaseWithoutDigest, + corpusId: parsed.data.corpusId, + corpusVersion: parsed.data.corpusVersion, + digest: caseDigest, + }; + checkCaseReferences(semanticCase, references); + return semanticCase; + }).sort((left, right) => compareUtf8(left.id, right.id)); + + const corpusWithoutDigest = { + schemaVersion: parsed.data.schemaVersion, + corpusId: parsed.data.corpusId, + corpusVersion: parsed.data.corpusVersion, + cases, + } as const; + const digest = semanticCorpusDigest(corpusWithoutDigest); + if (parsed.data.digest !== undefined && parsed.data.digest !== digest) invalidCorpus("corpus digest does not match parsed content"); + return { ...corpusWithoutDigest, digest }; +} + +export function compareSemanticCorpusVersions(previous: SemanticCorpus, current: SemanticCorpus): SemanticCorpusEvolutionComparison { + if (current.corpusVersion < previous.corpusVersion) return { valid: false, reason: "version-regressed" }; + const changed = semanticCorpusContentDigest(previous) !== semanticCorpusContentDigest(current); + if (changed && current.corpusVersion <= previous.corpusVersion) return { valid: false, reason: "version-not-increased" }; + return { valid: true }; +} diff --git a/server/services/tutor/evaluation/semantic/semantic-types.ts b/server/services/tutor/evaluation/semantic/semantic-types.ts new file mode 100644 index 00000000..3eba7ea3 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/semantic-types.ts @@ -0,0 +1,50 @@ +export const SEMANTIC_CASE_ROLES = ["development", "calibration", "held-out-evaluation"] as const; +export type SemanticCaseRole = typeof SEMANTIC_CASE_ROLES[number]; + +export const SEMANTIC_RUBRIC_DIMENSIONS = [ + "factual-correctness", + "sketch-code-grounding", + "learner-answer-diagnosis", + "precision", + "instructional-usefulness", + "scaffolding", + "dialogic-progression", + "non-repetition", + "difficulty-appropriateness", +] as const; +export type SemanticRubricDimension = typeof SEMANTIC_RUBRIC_DIMENSIONS[number]; + +export const SEMANTIC_ANSWER_CATEGORIES = [ + "fully-correct", + "partially-correct", + "typical-misconception", + "terminology-confusion", + "correct-poorly-phrased", + "explicitly-unknown", + "off-topic", + "unexpectedly-strong", +] as const; +export type SemanticAnswerCategory = typeof SEMANTIC_ANSWER_CATEGORIES[number]; + +export const SEMANTIC_EVIDENCE_SOURCE_KINDS = [ + "repository/sketch", + "repository/course-content", + "draft-factual-reference", + "reviewed-factual-reference", + "permitted-external-knowledge", +] as const; +export type SemanticEvidenceSourceKind = typeof SEMANTIC_EVIDENCE_SOURCE_KINDS[number]; + +export type CaseExposureStatus = + | "unexposed" + | "case-known" + | "outcome-exposed" + | "used-for-targeted-change" + | "unknown"; + +export interface SemanticEvidenceSource { + readonly id: string; + readonly kind: SemanticEvidenceSourceKind; + readonly reference: string; + readonly digest: string; +} diff --git a/server/services/tutor/evaluation/semantic/stage2a-adapter.ts b/server/services/tutor/evaluation/semantic/stage2a-adapter.ts new file mode 100644 index 00000000..cf295228 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/stage2a-adapter.ts @@ -0,0 +1,136 @@ +import type { TutorQualityEvaluationScenario } from "../real-provider-evaluation"; +import type { TutorPlanningContentContext } from "../../tutor-planning"; +import { canonicalSemanticDigest, canonicalSemanticJson, semanticSha256 } from "./semantic-canonical"; +import { validateFrozenPreTurnContext } from "./frozen-context"; +import { semanticCaseDigest, semanticCorpusDigest, type SemanticCase, type SemanticCorpus } from "./semantic-corpus"; + +export interface SemanticFixtureResolver { + readonly sketches: ReadonlyMap; + readonly courseContent: ReadonlyMap; +} + +export interface AdapterGeneratedProvenance { + readonly kind: "adapter-generated"; + readonly semanticCorpusId: string; + readonly semanticCorpusVersion: number; + readonly semanticCaseId: string; + readonly semanticCaseDigest: string; + readonly frozenPreTurnContextDigest: string; + readonly scenarioId: string; + readonly stageAScenario: TutorQualityEvaluationScenario; + readonly stageAScenarioDigest: string; + readonly identity: string; +} + +export interface AdapterGeneratedScenario { + readonly scenario: TutorQualityEvaluationScenario; + readonly provenance: AdapterGeneratedProvenance; +} + +export interface ScenarioCompatibility { + readonly compatible: boolean; + readonly reason?: string; +} + +function makeExpectedTurn(semanticCase: SemanticCase): TutorQualityEvaluationScenario["turns"][number] { + const context = semanticCase.frozenPreTurnContext; + return { + kind: "dialog", + question: context.question, + answer: context.learnerAnswer.text, + bindsToQuestion: context.learnerAnswer.bindsToQuestion, + difficulty: context.difficulty, + history: context.priorDialog, + }; +} + +function stageAEffectiveTurn(turn: TutorQualityEvaluationScenario["turns"][number]): TutorQualityEvaluationScenario["turns"][number] { + // Stage 2A's historyFromSource() treats an omitted history as an empty list. + // Normalize only that runtime default; preserve every declared history entry and its order. + return turn.kind === "dialog" && turn.history === undefined ? { ...turn, history: [] } : turn; +} + +export function assessStageAScenarioCompatibility( + scenario: TutorQualityEvaluationScenario, + semanticCase: SemanticCase, +): ScenarioCompatibility { + const context = semanticCase.frozenPreTurnContext; + if (context.sketchRef !== semanticCase.sketch.reference || context.sketchDigest !== semanticCase.sketch.digest) { + return { compatible: false, reason: "frozen-context-sketch-reference-or-digest-mismatch" }; + } + if (scenario.sketchRef !== semanticCase.sketch.reference || semanticSha256(scenario.sketch) !== semanticCase.sketch.digest) { + return { compatible: false, reason: "sketch-reference-or-digest-mismatch" }; + } + if (scenario.turns.length !== 1 || canonicalSemanticJson(stageAEffectiveTurn(scenario.turns[0])) !== canonicalSemanticJson(makeExpectedTurn(semanticCase))) { + return { compatible: false, reason: "preceding-question-answer-history-or-difficulty-mismatch" }; + } + if (context.courseContent.kind === "free-tutor") { + if (scenario.courseContent !== undefined) return { compatible: false, reason: "course-content-present-for-free-tutor-case" }; + } else { + if (scenario.courseContent?.revision !== context.courseContent.revision) { + return { compatible: false, reason: "course-content-revision-mismatch" }; + } + if (canonicalSemanticDigest(scenario.courseContent) !== context.courseContent.digest) { + return { compatible: false, reason: "course-content-or-strategy-digest-mismatch" }; + } + } + return { compatible: true }; +} + +function provenanceIdentity(source: Omit): string { + return canonicalSemanticDigest(source); +} + +export function toStage2AScenario( + corpus: SemanticCorpus, + semanticCase: SemanticCase, + fixtures: SemanticFixtureResolver, +): AdapterGeneratedScenario { + if (semanticCorpusDigest(corpus) !== corpus.digest) throw new Error("Semantic Corpus digest does not match its canonical content"); + const contextValidation = validateFrozenPreTurnContext(semanticCase.frozenPreTurnContext); + if (!contextValidation.valid) throw new Error(`Invalid Frozen Pre-Turn Context: ${contextValidation.reason}`); + const corpusCase = corpus.cases.find(({ id }) => id === semanticCase.id); + if (semanticCase.corpusId !== corpus.corpusId || semanticCase.corpusVersion !== corpus.corpusVersion + || !corpusCase || semanticCaseDigest(corpusCase) !== semanticCaseDigest(semanticCase)) { + throw new Error("Semantic Case is not the immutable case content selected by this Semantic Corpus"); + } + const sketch = fixtures.sketches.get(semanticCase.sketch.reference); + if (sketch === undefined || semanticSha256(sketch) !== semanticCase.sketch.digest) { + throw new Error("Stage-2A adapter cannot resolve the exact pinned sketch fixture"); + } + + let courseContent: TutorPlanningContentContext | undefined; + if (contextValidation.context.courseContent.kind === "repository-course-content") { + const reference = contextValidation.context.courseContent; + courseContent = fixtures.courseContent.get(reference.reference); + if (courseContent?.revision !== reference.revision || canonicalSemanticDigest(courseContent) !== reference.digest) { + throw new Error("Stage-2A adapter cannot resolve the exact Course Content and Strategy context"); + } + } + + const scenario: TutorQualityEvaluationScenario = { + id: semanticCase.id, + corpusId: corpus.corpusId, + corpusVersion: corpus.corpusVersion, + sketchRef: semanticCase.sketch.reference, + sketch, + ...(courseContent ? { courseContent } : {}), + turns: [makeExpectedTurn(semanticCase)], + }; + const provenanceWithoutIdentity: Omit = { + kind: "adapter-generated", + semanticCorpusId: corpus.corpusId, + semanticCorpusVersion: corpus.corpusVersion, + semanticCaseId: semanticCase.id, + semanticCaseDigest: semanticCase.digest, + frozenPreTurnContextDigest: contextValidation.context.digest, + scenarioId: scenario.id, + stageAScenario: scenario, + stageAScenarioDigest: canonicalSemanticDigest(scenario), + }; + const provenance: AdapterGeneratedProvenance = { + ...provenanceWithoutIdentity, + identity: provenanceIdentity(provenanceWithoutIdentity), + }; + return { scenario, provenance }; +} diff --git a/server/services/tutor/evaluation/semantic/stage2a-transcript.ts b/server/services/tutor/evaluation/semantic/stage2a-transcript.ts new file mode 100644 index 00000000..98021a72 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/stage2a-transcript.ts @@ -0,0 +1,348 @@ +import { tutorContentResultSchema } from "@shared/tutor"; +import type { + TutorQualityEvaluationScenario, + TutorQualityExecutionStatus, + TutorQualityTranscript, +} from "../real-provider-evaluation"; +import { canonicalSemanticDigest, canonicalSemanticJson, isSha256Digest } from "./semantic-canonical"; +import { semanticCaseDigest, type SemanticCase } from "./semantic-corpus"; +import { assessStageAScenarioCompatibility, type AdapterGeneratedProvenance } from "./stage2a-adapter"; +import { stage2bTranscriptDigest } from "./transcript-reference"; + +export interface ExistingTranscriptCompatibilityMapping { + readonly kind: "existing-transcript-compatibility"; + readonly mappingVersion: 1; + readonly sourceId: string; + readonly sourceEvaluationIdentity: string; + readonly sourceScenarioId: string; + readonly sourceScenario: TutorQualityEvaluationScenario; + readonly sourceScenarioDigest: string; + readonly semanticCorpusId: string; + readonly semanticCorpusVersion: number; + readonly semanticCaseId: string; + readonly semanticCaseDigest: string; + readonly frozenPreTurnContextDigest: string; + readonly compatibilityDigest: string; + readonly identity: string; +} + +export type Stage2ATranscriptProvenance = AdapterGeneratedProvenance | ExistingTranscriptCompatibilityMapping; + +export type Stage2ATranscriptValidation = + | { + readonly valid: true; + readonly transcript: TutorQualityTranscript; + readonly transcriptDigest: string; + readonly compatibilityMapping: Stage2ATranscriptProvenance; + readonly semanticEligibility: "eligible" | "not-evaluated"; + } + | { readonly valid: false; readonly reason: string }; + +function isRecord(value: unknown): value is Record { + return value !== null && typeof value === "object" && !Array.isArray(value); +} + +function hasOnlyKeys(value: Record, allowed: readonly string[]): boolean { + return Object.keys(value).every((key) => allowed.includes(key)); +} + +const LEARNING_PHASES = new Set(["LEARN", "DEEPEN", "EXPAND"]); + +function isExpectedScenario(value: unknown): boolean { + if (value === undefined) return true; + if (!isRecord(value) || !hasOnlyKeys(value, ["topicId", "topicIdAbsent", "learningPhase", "stateUnchanged", "questionNotRepeat"])) return false; + if (["topicId", "topicIdAbsent"].some((key) => value[key] !== undefined && !isNonEmptyString(value[key]))) return false; + const phase = value.learningPhase; + if (phase !== undefined && (typeof phase !== "string" || !LEARNING_PHASES.has(phase))) return false; + if (value.stateUnchanged !== undefined && typeof value.stateUnchanged !== "boolean") return false; + return value.questionNotRepeat === undefined || value.questionNotRepeat === "exact-or-heuristic"; +} + +function isStageAScenarioObject(value: unknown): value is TutorQualityEvaluationScenario { + if (!isRecord(value) || !hasOnlyKeys(value, ["id", "corpusId", "corpusVersion", "sketchRef", "sketch", "courseContent", "turns", "expected"])) return false; + if (!isNonEmptyString(value.id) || !isNonEmptyString(value.corpusId) || !Number.isInteger(value.corpusVersion) + || !isNonEmptyString(value.sketchRef) || typeof value.sketch !== "string" || !Array.isArray(value.turns)) return false; + return isExpectedScenario(value.expected) && (value.courseContent === undefined || isRecord(value.courseContent)); +} + +function isExecutionStatus(value: unknown): value is TutorQualityExecutionStatus { + return value === "completed" || value === "invalid" || value === "technical-failure" || value === "not-run"; +} + +interface TranscriptCallCounts { + readonly total: number; + readonly modelListCalls: number; + readonly generationCalls: number; +} + +const LAST_C0_CONTROL_CODE = 0x1f; +const DELETE_CONTROL_CODE = 0x7f; + +function containsAsciiControlCharacter(value: string): boolean { + for (const character of value) { + const codePoint = character.codePointAt(0); + if (codePoint !== undefined && (codePoint <= LAST_C0_CONTROL_CODE || codePoint === DELETE_CONTROL_CODE)) return true; + } + return false; +} + +function isCallCounts(value: unknown): value is TranscriptCallCounts { + return isRecord(value) + && ["total", "modelListCalls", "generationCalls"].every((key) => Number.isInteger(value[key]) && Number(value[key]) >= 0) + && Number(value.total) === Number(value.modelListCalls) + Number(value.generationCalls); +} + +function isDeterministicChecks(value: unknown): boolean { + return Array.isArray(value) && value.every((check) => isRecord(check) + && hasOnlyKeys(check, ["name", "passed", "details"]) + && isNonEmptyString(check.name) + && typeof check.passed === "boolean" + && (check.details === undefined || typeof check.details === "string")); +} + +function isInvariantViolations(value: unknown, turnCount: number): boolean { + const sources = new Set(["raw-provider", "final-tutor", "state", "scenario"]); + return Array.isArray(value) && value.every((violation) => isRecord(violation) + && hasOnlyKeys(violation, ["code", "source", "turnIndex", "details"]) + && isNonEmptyString(violation.code) + && sources.has(String(violation.source)) + && (violation.turnIndex === undefined || (Number.isInteger(violation.turnIndex) && Number(violation.turnIndex) >= 0 && Number(violation.turnIndex) < turnCount)) + && (violation.details === undefined || typeof violation.details === "string")); +} + +function isNonEmptyString(value: unknown): value is string { + return typeof value === "string" && value.trim().length > 0; +} + +function isTranscriptEnvelope(value: unknown): value is Record { + return isRecord(value) + && value.schemaVersion === "tutor-quality-transcript-v1" + && isSha256Digest(value.evaluationIdentity) + && typeof value.runId === "string" + && value.runId.length > 0 + && isExecutionStatus(value.executionStatus) + && Array.isArray(value.turns) + && isDeterministicChecks(value.deterministicChecks) + && isRecord(value.scenario) + && isRecord(value.metadata); +} + +function isTranscriptScenario(value: unknown, executionStatus: TutorQualityExecutionStatus, turnCount: number): value is Record & { readonly syntheticTurns: unknown[] } { + if (!isRecord(value) || !Array.isArray(value.syntheticTurns)) return false; + if (!isNonEmptyString(value.id) || !isNonEmptyString(value.corpusId) + || !Number.isInteger(value.corpusVersion) || Number(value.corpusVersion) < 1) return false; + if (!isNonEmptyString(value.sketchRef) || typeof value.sketch !== "string") return false; + if (executionStatus === "completed" && turnCount !== value.syntheticTurns.length) return false; + return turnCount <= value.syntheticTurns.length; +} + +function isTranscriptMetadata(metadata: unknown, scenario: Record): metadata is Record { + if (!isRecord(metadata) || metadata.corpusId !== scenario.corpusId || metadata.corpusVersion !== scenario.corpusVersion) return false; + if (!isNonEmptyString(metadata.providerId) || !isNonEmptyString(metadata.requestedModel) + || !Array.isArray(metadata.returnedModels) || !metadata.returnedModels.every(isNonEmptyString)) return false; + if (!isRecord(metadata.promptRevision) || !isNonEmptyString(metadata.promptRevision.id) + || !isSha256Digest(metadata.promptRevision.templateDigest) || !isNonEmptyString(metadata.courseContentRevision)) return false; + if (typeof metadata.gitSha !== "string" || !/^[0-9a-f]{40}$/.test(metadata.gitSha) || metadata.gitState !== "clean") return false; + if (!Number.isInteger(metadata.sampleIndex) || Number(metadata.sampleIndex) < 0 + || !Number.isInteger(metadata.sampleCount) || Number(metadata.sampleCount) < 1) return false; + return Number.isFinite(metadata.sampleDurationMs) && Number(metadata.sampleDurationMs) >= 0 + && isCallCounts(metadata.providerCalls) + && Number.isInteger(metadata.maxCalls) && Number(metadata.maxCalls) >= 1; +} + +function isTranscriptTurn( + turn: unknown, + index: number, + syntheticTurn: unknown, + executionStatus: TutorQualityExecutionStatus, +): turn is Record & { readonly providerCalls: TranscriptCallCounts } { + if (!isRecord(turn) || !isRecord(turn.input)) return false; + if (Number(turn.index) !== index || !Number.isFinite(turn.durationMs) || Number(turn.durationMs) < 0) return false; + if (!isDeterministicChecks(turn.deterministicChecks) || !isCallCounts(turn.providerCalls)) return false; + const finalTutorResult = turn.finalTutorResult; + if (finalTutorResult !== undefined && (!isRecord(finalTutorResult) || !tutorContentResultSchema.safeParse(finalTutorResult).success)) return false; + if (turn.returnedModel !== undefined && typeof turn.returnedModel !== "string") return false; + if (executionStatus === "completed" && finalTutorResult === undefined) return false; + return canonicalSemanticJson(turn.input) === canonicalSemanticJson(syntheticTurn); +} + +function transcriptTurnCallTotals( + turns: unknown[], + syntheticTurns: unknown[], + executionStatus: TutorQualityExecutionStatus, +): TranscriptCallCounts | undefined { + if (executionStatus === "completed" && turns.length !== syntheticTurns.length) return undefined; + if (turns.length > syntheticTurns.length) return undefined; + const totals = { total: 0, modelListCalls: 0, generationCalls: 0 }; + for (const [index, turn] of turns.entries()) { + if (!isTranscriptTurn(turn, index, syntheticTurns[index], executionStatus)) return undefined; + totals.total += Number(turn.providerCalls.total); + totals.modelListCalls += Number(turn.providerCalls.modelListCalls); + totals.generationCalls += Number(turn.providerCalls.generationCalls); + } + return totals; +} + +function isTranscriptMinimum(value: unknown): value is TutorQualityTranscript { + try { + if (!isTranscriptEnvelope(value)) return false; + const scenario = value.scenario as Record; + const metadata = value.metadata as Record; + const turns = value.turns as unknown[]; + const executionStatus = value.executionStatus as TutorQualityExecutionStatus; + if (!isTranscriptScenario(scenario, executionStatus, turns.length)) return false; + if (!isInvariantViolations(value.invariantViolations, scenario.syntheticTurns.length)) return false; + if (!isTranscriptMetadata(metadata, scenario)) return false; + const turnCalls = transcriptTurnCallTotals(turns, scenario.syntheticTurns, executionStatus); + return turnCalls !== undefined + && isCallCounts(metadata.providerCalls) + && canonicalSemanticJson(turnCalls) === canonicalSemanticJson(metadata.providerCalls); + } catch { + return false; + } +} + +function transcriptMatchesScenario(transcript: TutorQualityTranscript, scenario: TutorQualityEvaluationScenario): boolean { + const scenarioRecord = transcript.scenario; + if (scenarioRecord.id !== scenario.id || scenarioRecord.corpusId !== scenario.corpusId || scenarioRecord.corpusVersion !== scenario.corpusVersion) return false; + if (scenarioRecord.sketchRef !== scenario.sketchRef || scenarioRecord.sketch !== scenario.sketch) return false; + if (canonicalSemanticJson(scenarioRecord.syntheticTurns) !== canonicalSemanticJson(scenario.turns)) return false; + if (transcript.metadata.courseContentRevision !== (scenario.courseContent?.revision ?? "free-tutor")) return false; + if (scenario.courseContent?.progressionState === undefined) return transcript.stateBefore === undefined; + return transcript.stateBefore !== undefined && canonicalSemanticJson(transcript.stateBefore) === canonicalSemanticJson(scenario.courseContent.progressionState); +} + +function existingMappingContent(input: Omit): Record { + return input; +} + +function existingMappingIdentity(input: Omit): string { + return canonicalSemanticDigest(existingMappingContent(input)); +} + +function expectedMappingCompatibilityDigest( + transcript: TutorQualityTranscript, + sourceScenario: TutorQualityEvaluationScenario, + semanticCase: SemanticCase, +): string { + return canonicalSemanticDigest({ + sourceEvaluationIdentity: transcript.evaluationIdentity, + sourceScenario: sourceScenario, + semanticCaseId: semanticCase.id, + semanticCaseDigest: semanticCaseDigest(semanticCase), + frozenPreTurnContextDigest: semanticCase.frozenPreTurnContext.digest, + }); +} + +export function createExistingTranscriptCompatibilityMapping(input: { + readonly sourceId: string; + readonly transcript: TutorQualityTranscript; + readonly sourceScenario: TutorQualityEvaluationScenario; + readonly semanticCase: SemanticCase; +}): ExistingTranscriptCompatibilityMapping { + const { sourceId, transcript, sourceScenario, semanticCase } = input; + if (!sourceId.trim()) throw new Error("Existing-transcript mapping requires a stable source identity"); + if (sourceId.length > 512 || containsAsciiControlCharacter(sourceId)) throw new Error("Existing-transcript source identity must be bounded and contain no control characters"); + if (!isStageAScenarioObject(sourceScenario)) throw new Error("Existing Stage-A source scenario contains invalid or unsupported fields"); + if (!isTranscriptMinimum(transcript)) throw new Error("Existing Stage-2A transcript schema or identity is invalid"); + const compatibility = assessStageAScenarioCompatibility(sourceScenario, semanticCase); + if (!compatibility.compatible) throw new Error(`Existing Stage-2A transcript is not compatible: ${compatibility.reason}`); + if (!transcriptMatchesScenario(transcript, sourceScenario)) throw new Error("Existing Stage-2A transcript does not match its declared source scenario"); + const sourceScenarioDigest = canonicalSemanticDigest(sourceScenario); + const mappingWithoutIdentity: Omit = { + kind: "existing-transcript-compatibility", + mappingVersion: 1, + sourceId, + sourceEvaluationIdentity: transcript.evaluationIdentity, + sourceScenarioId: sourceScenario.id, + sourceScenario, + sourceScenarioDigest, + semanticCaseId: semanticCase.id, + semanticCaseDigest: semanticCaseDigest(semanticCase), + frozenPreTurnContextDigest: semanticCase.frozenPreTurnContext.digest, + compatibilityDigest: expectedMappingCompatibilityDigest(transcript, sourceScenario, semanticCase), + semanticCorpusId: semanticCase.corpusId, + semanticCorpusVersion: semanticCase.corpusVersion, + }; + return { ...mappingWithoutIdentity, identity: existingMappingIdentity(mappingWithoutIdentity) }; +} + +function validateAdapterProvenance( + provenance: AdapterGeneratedProvenance, + transcript: TutorQualityTranscript, + semanticCase: SemanticCase, +): string | undefined { + if (!isStageAScenarioObject(provenance.stageAScenario)) return "adapter-stage-a-scenario-schema-invalid"; + if (provenance.semanticCaseId !== semanticCase.id || provenance.semanticCaseDigest !== semanticCaseDigest(semanticCase)) return "adapter-case-identity-mismatch"; + if (provenance.frozenPreTurnContextDigest !== semanticCase.frozenPreTurnContext.digest) return "adapter-frozen-context-mismatch"; + if (provenance.scenarioId !== semanticCase.id || provenance.stageAScenario.id !== semanticCase.id) return "adapter-scenario-id-mismatch"; + if (provenance.stageAScenario.corpusId !== provenance.semanticCorpusId || provenance.stageAScenario.corpusVersion !== provenance.semanticCorpusVersion) return "adapter-corpus-source-identity-mismatch"; + if (provenance.semanticCorpusId !== semanticCase.corpusId || provenance.semanticCorpusVersion !== semanticCase.corpusVersion) return "adapter-corpus-identity-invalid"; + if (provenance.stageAScenarioDigest !== canonicalSemanticDigest(provenance.stageAScenario)) return "adapter-scenario-digest-mismatch"; + const identitySource = { ...provenance } as Record; + delete identitySource.identity; + if (provenance.identity !== canonicalSemanticDigest(identitySource)) return "adapter-mapping-identity-mismatch"; + const compatible = assessStageAScenarioCompatibility(provenance.stageAScenario, semanticCase); + if (!compatible.compatible) return compatible.reason; + if (!transcriptMatchesScenario(transcript, provenance.stageAScenario)) return "transcript-does-not-match-adapter-scenario"; + return undefined; +} + +function validateExistingProvenance( + provenance: ExistingTranscriptCompatibilityMapping, + transcript: TutorQualityTranscript, + semanticCase: SemanticCase, +): string | undefined { + if (!provenance.sourceId.trim() || provenance.sourceId.length > 512 || containsAsciiControlCharacter(provenance.sourceId)) return "existing-transcript-source-identity-missing-or-invalid"; + if (!isStageAScenarioObject(provenance.sourceScenario)) return "existing-transcript-source-scenario-invalid"; + if (provenance.sourceEvaluationIdentity !== transcript.evaluationIdentity) return "existing-transcript-evaluation-identity-mismatch"; + if (provenance.sourceScenarioId !== transcript.scenario.id || provenance.sourceScenario.id !== transcript.scenario.id) return "existing-transcript-source-scenario-mismatch"; + if (provenance.semanticCaseId !== semanticCase.id || provenance.semanticCaseDigest !== semanticCaseDigest(semanticCase)) return "existing-transcript-semantic-case-mismatch"; + if (provenance.semanticCorpusId !== semanticCase.corpusId || provenance.semanticCorpusVersion !== semanticCase.corpusVersion) return "existing-transcript-semantic-corpus-mismatch"; + if (provenance.frozenPreTurnContextDigest !== semanticCase.frozenPreTurnContext.digest) return "existing-transcript-context-mismatch"; + if (provenance.sourceScenarioDigest !== canonicalSemanticDigest(provenance.sourceScenario)) return "existing-transcript-source-scenario-digest-mismatch"; + const compatibility = assessStageAScenarioCompatibility(provenance.sourceScenario, semanticCase); + if (!compatibility.compatible) return compatibility.reason; + if (!transcriptMatchesScenario(transcript, provenance.sourceScenario)) return "existing-transcript-content-does-not-match-source-scenario"; + if (provenance.compatibilityDigest !== expectedMappingCompatibilityDigest(transcript, provenance.sourceScenario, semanticCase)) return "existing-transcript-compatibility-proof-mismatch"; + const mappingSource = { ...provenance } as Record; + delete mappingSource.identity; + if (provenance.identity !== existingMappingIdentity(mappingSource as Omit)) return "existing-transcript-mapping-identity-mismatch"; + return undefined; +} + +export function validateStage2ATranscript( + input: unknown, + semanticCase: SemanticCase, + provenance?: AdapterGeneratedProvenance | ExistingTranscriptCompatibilityMapping, +): Stage2ATranscriptValidation { + if (!isTranscriptMinimum(input)) return { valid: false, reason: "stage-a-transcript-schema-or-identity-invalid" }; + if (!provenance) return { valid: false, reason: "explicit-stage-b-transcript-mapping-required" }; + const transcript = input; + const mappingKeys = provenance.kind === "adapter-generated" + ? ["kind", "semanticCorpusId", "semanticCorpusVersion", "semanticCaseId", "semanticCaseDigest", "frozenPreTurnContextDigest", "scenarioId", "stageAScenario", "stageAScenarioDigest", "identity"] + : ["kind", "mappingVersion", "sourceId", "sourceEvaluationIdentity", "sourceScenarioId", "sourceScenario", "sourceScenarioDigest", "semanticCorpusId", "semanticCorpusVersion", "semanticCaseId", "semanticCaseDigest", "frozenPreTurnContextDigest", "compatibilityDigest", "identity"]; + if (!isRecord(provenance) || !hasOnlyKeys(provenance, mappingKeys)) return { valid: false, reason: "stage-b-transcript-mapping-schema-invalid" }; + let reason: string | undefined; + if (provenance.kind === "adapter-generated") { + reason = validateAdapterProvenance(provenance, transcript, semanticCase); + } else if (provenance.kind === "existing-transcript-compatibility") { + reason = validateExistingProvenance(provenance, transcript, semanticCase); + } else { + reason = "unsupported-stage-b-transcript-mapping"; + } + if (reason) return { valid: false, reason }; + try { + return { + valid: true, + transcript, + transcriptDigest: stage2bTranscriptDigest(transcript), + compatibilityMapping: provenance, + semanticEligibility: transcript.executionStatus === "completed" ? "eligible" : "not-evaluated", + }; + } catch { + return { valid: false, reason: "stage-a-transcript-cannot-be-canonicalized" }; + } +} + +export { stage2bTranscriptDigest } from "./transcript-reference"; diff --git a/server/services/tutor/evaluation/semantic/transcript-reference.ts b/server/services/tutor/evaluation/semantic/transcript-reference.ts new file mode 100644 index 00000000..cbe087b7 --- /dev/null +++ b/server/services/tutor/evaluation/semantic/transcript-reference.ts @@ -0,0 +1,49 @@ +import type { TutorQualityTranscript } from "../real-provider-evaluation"; +import { canonicalSemanticDigest, canonicalSemanticJson, isSha256Digest } from "./semantic-canonical"; +import type { Stage2ATranscriptProvenance } from "./stage2a-transcript"; + +export interface Stage2ATranscriptReference { + readonly schemaVersion: "tutor-quality-stage2b-transcript-reference-v1"; + readonly stageAEvaluationIdentity: string; + readonly transcriptDigest: string; + readonly compatibilityMappingIdentity: string; + readonly compatibilityMappingKind: Stage2ATranscriptProvenance["kind"]; + readonly identity: string; +} + +export function canonicalStage2ATranscriptJson(transcript: TutorQualityTranscript): string { + return canonicalSemanticJson(transcript); +} + +/** Stage 2B's own canonical digest of an immutable Stage-2A transcript. */ +export function stage2bTranscriptDigest(transcript: TutorQualityTranscript): string { + return canonicalSemanticDigest(transcript); +} + +export function createStage2ATranscriptReference( + transcript: TutorQualityTranscript, + transcriptDigest: string, + compatibilityMapping: Stage2ATranscriptProvenance, +): Stage2ATranscriptReference { + if (!isSha256Digest(transcript.evaluationIdentity)) throw new Error("Stage-2A evaluationIdentity is missing or invalid"); + if (!isSha256Digest(transcriptDigest) || transcriptDigest !== stage2bTranscriptDigest(transcript)) { + throw new Error("Stage-B transcriptDigest does not match the immutable Stage-2A transcript"); + } + if (!compatibilityMapping || !isSha256Digest(compatibilityMapping.identity)) throw new Error("Stage-B compatibility mapping identity is invalid"); + const { identity: _mappingIdentity, ...mappingContent } = compatibilityMapping; + if (compatibilityMapping.identity !== canonicalSemanticDigest(mappingContent)) throw new Error("Stage-B compatibility mapping digest does not match its content"); + if (compatibilityMapping.kind === "existing-transcript-compatibility" + && compatibilityMapping.sourceEvaluationIdentity !== transcript.evaluationIdentity) { + throw new Error("Existing Stage-A transcript identity does not match its compatibility mapping"); + } + const identitySource = { + schemaVersion: "tutor-quality-stage2b-transcript-reference-v1" as const, + stageAEvaluationIdentity: transcript.evaluationIdentity, + transcriptDigest, + compatibilityMappingIdentity: compatibilityMapping.identity, + compatibilityMappingKind: compatibilityMapping.kind, + }; + return { ...identitySource, identity: canonicalSemanticDigest(identitySource) }; +} + +export { canonicalSemanticJson as canonicalStage2BJson } from "./semantic-canonical"; diff --git a/tests/server/services/tutor/evaluation/semantic/case-exposure.test.ts b/tests/server/services/tutor/evaluation/semantic/case-exposure.test.ts new file mode 100644 index 00000000..7831a8f3 --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/case-exposure.test.ts @@ -0,0 +1,197 @@ +import { describe, expect, it } from "vitest"; +import { canonicalSemanticDigest } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-canonical"; +import { + createCaseExposureRecord, + validateCaseExposureRecord, + type CaseExposureSnapshot, +} from "../../../../../../server/services/tutor/evaluation/semantic/case-exposure"; +import { parseSemanticCorpus } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-corpus"; +import { createFrozenPreTurnContext } from "../../../../../../server/services/tutor/evaluation/semantic/frozen-context"; +import { + validFrozenPreTurnContextInput, + validSemanticCorpusReferences, + validSemanticCorpusSource, +} from "./semantic-test-fixtures"; + +const binding = { + candidateIdentity: "a".repeat(64), + corpusId: "tutor-quality-semantic", + corpusVersion: 1, + caseId: "TQ-SEM-001", + semanticCaseDigest: "b".repeat(64), +}; + +function snapshot( + availableArtifacts: CaseExposureSnapshot["availableArtifacts"], + targetedChange: CaseExposureSnapshot["targetedChange"] = { status: "not-informed" }, + recordedAt = "2026-09-27T12:00:00.000Z", +): CaseExposureSnapshot { + return { recordedAt, availableArtifacts, targetedChange }; +} + +describe("Candidate-specific CaseExposureRecord", () => { + it.each([ + ["unexposed", { caseDefinition: "unavailable", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }], + ["case-known", { caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }], + ["outcome-exposed", { caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }], + ["used-for-targeted-change", { caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "informed", rationale: "The observed correction error directly motivated this Candidate change." }], + ["unknown", { caseDefinition: "unknown", tutorOutput: "unknown", humanReference: "unknown", judgeResult: "unknown" }, { status: "unknown" }], + ] as const)("derives the SSOT exposure state %s from its evidence", (expected, availableArtifacts, targetedChange) => { + const record = createCaseExposureRecord({ + ...binding, + history: [snapshot(availableArtifacts, targetedChange)], + }); + + expect(record.exposureStatus).toBe(expected); + expect(validateCaseExposureRecord(record, binding)).toMatchObject({ valid: true }); + expect(record.identity).toMatch(/^[0-9a-f]{64}$/); + expect(record.digest).toMatch(/^[0-9a-f]{64}$/); + }); + + it("binds exposure to one Candidate and the exact case version", () => { + const record = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + + expect(validateCaseExposureRecord(record, { ...binding, candidateIdentity: "c".repeat(64) })).toMatchObject({ valid: false }); + expect(validateCaseExposureRecord(record, { ...binding, corpusVersion: 2 })).toMatchObject({ valid: false }); + }); + + it("appends a new immutable revision without rewriting prior exposure history", () => { + const first = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + const second = createCaseExposureRecord({ + ...binding, + history: [ + ...first.history, + snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z"), + ], + }, first); + + expect(second.recordVersion).toBe(2); + expect(second.previousRecordDigest).toBe(first.digest); + expect(second.exposureStatus).toBe("outcome-exposed"); + expect(first.recordVersion).toBe(1); + expect(validateCaseExposureRecord(second, binding)).toMatchObject({ valid: true }); + expect(first.exposureStatus).toBe("case-known"); + expect(first.history).toHaveLength(1); + expect(Object.isFrozen(first)).toBe(true); + }); + + it("validates exposure chains beyond two revisions", () => { + const first = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + const second = createCaseExposureRecord({ + ...binding, + history: [ + ...first.history, + snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z"), + ], + }, first); + const third = createCaseExposureRecord({ + ...binding, + history: [ + ...second.history, + snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "available", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-29T10:00:00.000Z"), + ], + }, second); + + expect(validateCaseExposureRecord(third, binding)).toMatchObject({ valid: true }); + }); + + it("verifies the previous-record digest against the actual history prefix", () => { + const first = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + const second = createCaseExposureRecord({ + ...binding, + history: [ + ...first.history, + snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z"), + ], + }, first); + const { digest: _originalDigest, ...secondWithoutDigest } = second; + const forgedBase = { ...secondWithoutDigest, previousRecordDigest: "f".repeat(64) }; + const forgedIdentity = canonicalSemanticDigest({ + schemaVersion: forgedBase.schemaVersion, + candidateIdentity: forgedBase.candidateIdentity, + corpusId: forgedBase.corpusId, + corpusVersion: forgedBase.corpusVersion, + caseId: forgedBase.caseId, + semanticCaseDigest: forgedBase.semanticCaseDigest, + recordVersion: forgedBase.recordVersion, + previousRecordDigest: forgedBase.previousRecordDigest, + }); + const forgedWithoutDigest = { ...forgedBase, identity: forgedIdentity }; + const forged = { ...forgedWithoutDigest, digest: canonicalSemanticDigest(forgedWithoutDigest) }; + + expect(validateCaseExposureRecord(forged, binding)).toMatchObject({ valid: false, reason: "previous-record-digest-history-mismatch" }); + }); + + it("requires a verifiable previous-record link after the initial exposure version", () => { + const first = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + const second = createCaseExposureRecord({ + ...binding, + history: [ + ...first.history, + snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z"), + ], + }, first); + const { previousRecordDigest: _previousRecordDigest, digest: _originalDigest, ...withoutPrevious } = second; + const withoutPreviousIdentity = canonicalSemanticDigest({ + schemaVersion: withoutPrevious.schemaVersion, + candidateIdentity: withoutPrevious.candidateIdentity, + corpusId: withoutPrevious.corpusId, + corpusVersion: withoutPrevious.corpusVersion, + caseId: withoutPrevious.caseId, + semanticCaseDigest: withoutPrevious.semanticCaseDigest, + recordVersion: withoutPrevious.recordVersion, + }); + const withoutPreviousBase = { ...withoutPrevious, identity: withoutPreviousIdentity }; + const forged = { ...withoutPreviousBase, digest: canonicalSemanticDigest(withoutPreviousBase) }; + + expect(validateCaseExposureRecord(forged, binding)).toMatchObject({ valid: false, reason: "previous-record-digest-history-mismatch" }); + }); + + it("rejects a revision that drops prior events or changes Candidate/case ownership", () => { + const first = createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + + expect(() => createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "unavailable", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }, first)).toThrow(/append|history/i); + expect(() => createCaseExposureRecord({ + ...binding, + candidateIdentity: "c".repeat(64), + history: [...first.history, snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z")], + }, first)).toThrow(/candidate|binding/i); + expect(() => createCaseExposureRecord({ + ...binding, + history: [...first.history, snapshot({ caseDefinition: "available", tutorOutput: "available", humanReference: "unavailable", judgeResult: "unavailable" }, { status: "not-informed" }, "2026-09-28T10:00:00.000Z")], + }, { ...first, digest: "f".repeat(64) })).toThrow(/valid prior record/i); + }); + + it("does not change the Semantic Corpus digest when Candidate exposure is added", () => { + const corpus = parseSemanticCorpus(validSemanticCorpusSource(createFrozenPreTurnContext(validFrozenPreTurnContextInput())), validSemanticCorpusReferences()); + const digestBefore = corpus.digest; + + createCaseExposureRecord({ + ...binding, + history: [snapshot({ caseDefinition: "available", tutorOutput: "unavailable", humanReference: "unavailable", judgeResult: "unavailable" })], + }); + + expect(corpus.digest).toBe(digestBefore); + }); +}); diff --git a/tests/server/services/tutor/evaluation/semantic/frozen-context.test.ts b/tests/server/services/tutor/evaluation/semantic/frozen-context.test.ts new file mode 100644 index 00000000..7421bde1 --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/frozen-context.test.ts @@ -0,0 +1,75 @@ +import { describe, expect, it } from "vitest"; +import { + createFrozenPreTurnContext, + validateFrozenPreTurnContext, +} from "../../../../../../server/services/tutor/evaluation/semantic/frozen-context"; +import { deepFreeze } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-canonical"; +import { INPUT_PULLUP_QUESTION, validFrozenPreTurnContextInput } from "./semantic-test-fixtures"; + +describe("Frozen Pre-Turn Context", () => { + it("binds the synthetic answer to the exact declared preceding question", () => { + const context = createFrozenPreTurnContext(validFrozenPreTurnContextInput()); + + expect(validateFrozenPreTurnContext(context)).toMatchObject({ valid: true, context: { digest: context.digest } }); + expect(context.learnerAnswer.bindsToQuestion).toBe(INPUT_PULLUP_QUESTION); + }); + + it("rejects an answer silently rebound to another question", () => { + const context = createFrozenPreTurnContext(validFrozenPreTurnContextInput()); + const invalid = { ...context, learnerAnswer: { ...context.learnerAnswer, bindsToQuestion: "Welche Farbe hat die LED?" } }; + + expect(validateFrozenPreTurnContext(invalid)).toMatchObject({ valid: false }); + }); + + it("preserves ordered prior dialog and detects every material context change", () => { + const context = createFrozenPreTurnContext({ + ...validFrozenPreTurnContextInput(), + priorDialog: [ + { question: "Erste Frage?", answer: "Erste Antwort.", responseStyle: "normal" }, + { question: "Zweite Frage?", answer: "Zweite Antwort.", responseStyle: "normal" }, + ], + }); + const reordered = createFrozenPreTurnContext({ ...context, priorDialog: [...context.priorDialog].reverse(), digest: undefined }); + const changedDifficulty = createFrozenPreTurnContext({ ...context, difficulty: 31, digest: undefined }); + const changedSketch = createFrozenPreTurnContext({ ...context, sketchDigest: "f".repeat(64), digest: undefined }); + + expect(reordered.digest).not.toBe(context.digest); + expect(changedDifficulty.digest).not.toBe(context.digest); + expect(changedSketch.digest).not.toBe(context.digest); + expect(validateFrozenPreTurnContext({ ...context, digest: "0".repeat(64) })).toMatchObject({ valid: false }); + }); + + it("rejects divergent frozen contexts instead of repairing them", () => { + const left = createFrozenPreTurnContext(validFrozenPreTurnContextInput()); + const right = createFrozenPreTurnContext({ ...validFrozenPreTurnContextInput(), question: "A different question?", learnerAnswer: { ...validFrozenPreTurnContextInput().learnerAnswer, bindsToQuestion: "A different question?" } }); + + expect(left.digest).not.toBe(right.digest); + expect(validateFrozenPreTurnContext(right)).toMatchObject({ valid: true }); + expect(validateFrozenPreTurnContext({ ...right, digest: left.digest })).toMatchObject({ valid: false }); + }); + + it("recursively freezes nested values even when a prior dialog turn is already shallow-frozen", () => { + const turn = Object.freeze({ + question: "Which topic was mastered?", + answer: "Timing.", + responseStyle: "normal" as const, + masteredTopicIds: ["timing"], + }); + const context = createFrozenPreTurnContext({ + ...validFrozenPreTurnContextInput(), + priorDialog: [turn], + }); + + expect(Object.isFrozen(context.priorDialog[0])).toBe(true); + expect(Object.isFrozen(context.priorDialog[0]?.masteredTopicIds)).toBe(true); + }); + + it("recursively freezes descendants of any already-frozen object", () => { + const nested = [] as string[]; + const frozenParent = Object.freeze({ nested }); + + deepFreeze(frozenParent); + + expect(Object.isFrozen(nested)).toBe(true); + }); +}); diff --git a/tests/server/services/tutor/evaluation/semantic/semantic-corpus.test.ts b/tests/server/services/tutor/evaluation/semantic/semantic-corpus.test.ts new file mode 100644 index 00000000..9ed2d004 --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/semantic-corpus.test.ts @@ -0,0 +1,139 @@ +import { readFileSync } from "node:fs"; +import { fileURLToPath } from "node:url"; +import { parse as parseYaml } from "yaml"; +import { describe, expect, it } from "vitest"; +import { + compareSemanticCorpusVersions, + parseSemanticCorpus, + semanticCorpusDigest, +} from "../../../../../../server/services/tutor/evaluation/semantic/semantic-corpus"; +import { semanticSha256 } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-canonical"; +import { createFrozenPreTurnContext } from "../../../../../../server/services/tutor/evaluation/semantic/frozen-context"; +import { + INPUT_PULLUP_SKETCH_DIGEST, + INPUT_PULLUP_SKETCH_REF, + validFrozenPreTurnContextInput, + validSemanticCorpusReferences, + validSemanticCorpusSource, +} from "./semantic-test-fixtures"; + +function source() { + return validSemanticCorpusSource(createFrozenPreTurnContext(validFrozenPreTurnContextInput())); +} + +describe("Stage-2B Semantic Corpus", () => { + it("loads the single complete executable TQ-SEM-001 development case", () => { + const yamlPath = fileURLToPath(new URL("../../../../../../evals/tutor-quality/semantic-corpus.yaml", import.meta.url)); + const sketchPath = fileURLToPath(new URL("../../../../../../evals/tutor-quality/semantic-fixtures/TQ-SEM-001-input-pullup.ino", import.meta.url)); + const sketch = readFileSync(sketchPath, "utf8"); + const references = validSemanticCorpusReferences(); + references.sketches.set(INPUT_PULLUP_SKETCH_REF, sketch); + const corpus = parseSemanticCorpus(parseYaml(readFileSync(yamlPath, "utf8")), references); + + expect(semanticSha256(sketch)).toBe(INPUT_PULLUP_SKETCH_DIGEST); + expect(corpus.corpusId).toBe("tutor-quality-semantic"); + expect(corpus.corpusVersion).toBe(1); + expect(corpus.cases.map(({ id }) => id)).toEqual(["TQ-SEM-001"]); + expect(corpus.cases[0]).toMatchObject({ + role: "development", + sketch: { reference: INPUT_PULLUP_SKETCH_REF, digest: INPUT_PULLUP_SKETCH_DIGEST }, + frozenPreTurnContext: { + question: "Welche logische Bedingung muss erfüllt sein, damit die LED eingeschaltet wird?", + learnerAnswer: { + text: "buttonPin, also an PIN2 muss GND anliegen!", + bindsToQuestion: "Welche logische Bedingung muss erfüllt sein, damit die LED eingeschaltet wird?", + category: "fully-correct", + }, + courseContent: { kind: "free-tutor" }, + }, + humanReference: { status: "pending-review" }, + }); + }); + + it("rejects duplicate case IDs, unsupported roles, missing evidence, and broken answer bindings", () => { + const input = source(); + const semanticCase = input.cases[0]!; + + expect(() => parseSemanticCorpus({ ...input, cases: [semanticCase, semanticCase] }, validSemanticCorpusReferences())).toThrow(/duplicate/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, role: "held-out" }] }, validSemanticCorpusReferences())).toThrow(/role/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, humanReference: { status: "reviewed" } }] }, validSemanticCorpusReferences())).toThrow(/humanReference|reviewed/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, sketch: { ...semanticCase.sketch, reference: "missing.ino" } }] }, validSemanticCorpusReferences())).toThrow(/sketch/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, frozenPreTurnContext: { ...semanticCase.frozenPreTurnContext, learnerAnswer: { ...semanticCase.frozenPreTurnContext.learnerAnswer, bindsToQuestion: "another question" } } }] }, validSemanticCorpusReferences())).toThrow(/bind/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, frozenPreTurnContext: { ...semanticCase.frozenPreTurnContext, sketchDigest: "d".repeat(64) } }] }, validSemanticCorpusReferences())).toThrow(/sketch|digest/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, evidenceSources: [{ id: "unsupported", kind: "judge-opinion", reference: "x", digest: "d" }] }] }, validSemanticCorpusReferences())).toThrow(/evidence|source/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, assessableDimensions: ["factual-correctness", "factual-correctness"] }] }, validSemanticCorpusReferences())).toThrow(/duplicate/i); + expect(() => parseSemanticCorpus({ ...input, cases: [{ ...semanticCase, dimensionEvidenceLimitations: { ...semanticCase.dimensionEvidenceLimitations, inventedDimension: "unknown rubric dimension" } }] }, validSemanticCorpusReferences())).toThrow(/unknown rubric dimension/i); + }); + + it("requires a semantic corpus version increase when case evidence or role changes", () => { + const previous = parseSemanticCorpus(source(), validSemanticCorpusReferences()); + const altered = source(); + altered.cases[0]!.role = "calibration"; + const currentSameVersion = parseSemanticCorpus(altered, validSemanticCorpusReferences()); + const currentBumpedVersion = parseSemanticCorpus({ ...altered, corpusVersion: 2 }, validSemanticCorpusReferences()); + + expect(compareSemanticCorpusVersions(previous, currentSameVersion)).toMatchObject({ valid: false, reason: "version-not-increased" }); + expect(compareSemanticCorpusVersions(previous, currentBumpedVersion)).toMatchObject({ valid: true }); + }); + + it("rejects corpus-version rollback even when parsed content is otherwise identical", () => { + const previous = parseSemanticCorpus({ ...source(), corpusVersion: 2 }, validSemanticCorpusReferences()); + const rolledBack = parseSemanticCorpus(source(), validSemanticCorpusReferences()); + + expect(compareSemanticCorpusVersions(previous, rolledBack)).toMatchObject({ valid: false, reason: "version-regressed" }); + }); + + it("does not put Candidate exposure into the static corpus digest", () => { + const corpus = parseSemanticCorpus(source(), validSemanticCorpusReferences()); + const before = semanticCorpusDigest(corpus); + + expect(semanticCorpusDigest(corpus)).toBe(before); + expect(JSON.stringify(corpus)).not.toContain("candidateIdentity"); + expect(JSON.stringify(corpus)).not.toContain("targetedChange"); + }); + + it("canonicalizes unordered case sets and linked evidence records without separating bindings", () => { + const original = source(); + const originalCase = original.cases[0]!; + const reordered = { + ...original, + cases: [{ + ...originalCase, + acceptableTutorResponses: { + diagnoses: [...originalCase.acceptableTutorResponses.diagnoses].reverse(), + feedbackApproaches: [...originalCase.acceptableTutorResponses.feedbackApproaches].reverse(), + followUps: [...originalCase.acceptableTutorResponses.followUps].reverse(), + }, + knownFailurePatterns: [...originalCase.knownFailurePatterns].reverse(), + assessableDimensions: [...originalCase.assessableDimensions].reverse(), + evidenceSources: [...originalCase.evidenceSources].reverse(), + factualReferenceBundle: { + ...originalCase.factualReferenceBundle, + limitations: [...originalCase.factualReferenceBundle.limitations].reverse(), + facts: [...originalCase.factualReferenceBundle.facts].reverse(), + }, + }], + }; + + expect(semanticCorpusDigest(reordered)).toBe(semanticCorpusDigest(original)); + const relinked = structuredClone(reordered); + const sources = relinked.cases[0]!.evidenceSources; + [sources[0]!.digest, sources[1]!.digest] = [sources[1]!.digest, sources[0]!.digest]; + expect(semanticCorpusDigest(relinked)).not.toBe(semanticCorpusDigest(original)); + }); + + it("canonicalizes case-list order when multiple cases are present", () => { + const oneCase = source(); + const secondCase = { ...oneCase.cases[0]!, id: "TQ-SEM-001-SECOND" }; + const twoCases = { ...oneCase, cases: [oneCase.cases[0]!, secondCase] }; + + expect(semanticCorpusDigest(twoCases)).toBe(semanticCorpusDigest({ ...twoCases, cases: [...twoCases.cases].reverse() })); + }); + + it("rejects a sketch fixture whose bytes do not match the pinned digest", () => { + const references = validSemanticCorpusReferences(); + references.sketches.set(INPUT_PULLUP_SKETCH_REF, "different sketch"); + + expect(() => parseSemanticCorpus(source(), references)).toThrow(/digest/i); + }); +}); diff --git a/tests/server/services/tutor/evaluation/semantic/semantic-test-fixtures.ts b/tests/server/services/tutor/evaluation/semantic/semantic-test-fixtures.ts new file mode 100644 index 00000000..3e0eaccb --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/semantic-test-fixtures.ts @@ -0,0 +1,188 @@ +import { createHash } from "node:crypto"; +import type { TutorQualityEvaluationScenario, TutorQualityTranscript } from "../../../../../../server/services/tutor/evaluation/real-provider-evaluation"; +import type { TutorPlanningContentContext } from "../../../../../../server/services/tutor/tutor-planning"; +import { factualReferenceBundleDigest } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-corpus"; + +export const INPUT_PULLUP_SKETCH_REF = "semantic-fixtures/TQ-SEM-001-input-pullup.ino"; +export const INPUT_PULLUP_SKETCH = `const int buttonPin = 2; +const int ledPin = LED_BUILTIN; + +void setup() { + pinMode(buttonPin, INPUT_PULLUP); + pinMode(ledPin, OUTPUT); +} + +void loop() { + digitalWrite( + ledPin, + digitalRead(buttonPin) == LOW ? HIGH : LOW + ); +} +`; +export const INPUT_PULLUP_SKETCH_DIGEST = createHash("sha256").update(INPUT_PULLUP_SKETCH, "utf8").digest("hex"); +export const INPUT_PULLUP_QUESTION = "Welche logische Bedingung muss erfüllt sein, damit die LED eingeschaltet wird?"; +export const INPUT_PULLUP_ANSWER = "buttonPin, also an PIN2 muss GND anliegen!"; + +export function validFrozenPreTurnContextInput() { + return { + schemaVersion: "tutor-quality-frozen-context-v1", + sketchRef: INPUT_PULLUP_SKETCH_REF, + sketchDigest: INPUT_PULLUP_SKETCH_DIGEST, + question: INPUT_PULLUP_QUESTION, + learnerAnswer: { + text: INPUT_PULLUP_ANSWER, + category: "fully-correct", + bindsToQuestion: INPUT_PULLUP_QUESTION, + }, + priorDialog: [], + courseContent: { kind: "free-tutor" }, + difficulty: 30, + }; +} + +export function validSemanticCaseSource(frozenPreTurnContext: Record) { + const factualReferenceBundle = { + id: "TQ-SEM-001-factual-reference", + version: 1, + sourceKind: "draft-factual-reference", + reviewStatus: "pending-review", + provenance: "Draft facts derived from the pinned sketch and requiring reviewer confirmation of board and wiring assumptions.", + limitations: ["Confirm the target board, button-to-GND wiring, and LED_BUILTIN active-high behavior before Gold use."], + facts: [ + { id: "internal-pullup", statement: "pinMode(buttonPin, INPUT_PULLUP) enables the microcontroller's internal pull-up for buttonPin." }, + { id: "grounded-input-low", statement: "With the declared button-to-GND wiring, pressing the button makes digitalRead(buttonPin) return LOW." }, + { id: "sketch-condition", statement: "The sketch writes HIGH to ledPin, which is LED_BUILTIN, when digitalRead(buttonPin) == LOW." }, + ], + }; + const factualDigest = factualReferenceBundleDigest(factualReferenceBundle); + return { + id: "TQ-SEM-001", + caseVersion: 1, + purpose: "Check that a substantively correct INPUT_PULLUP answer is recognized without false correction.", + role: "development", + sketch: { + reference: INPUT_PULLUP_SKETCH_REF, + digest: INPUT_PULLUP_SKETCH_DIGEST, + }, + frozenPreTurnContext, + expectedAnswerInterpretation: "When the button connects buttonPin to GND, INPUT_PULLUP makes the input read LOW; this sketch writes HIGH to ledPin when digitalRead(buttonPin) == LOW.", + acceptableTutorResponses: { + diagnoses: ["Recognizes the learner's core condition as correct, allowing substantively equivalent wording."], + feedbackApproaches: ["May clarify the active-low input without claiming the learner's answer is wrong."], + followUps: ["May ask a sketch-grounded question that adds a distinct reasoning demand."], + }, + answerCategory: "fully-correct", + knownFailurePatterns: ["correct-answer-rejected"], + assessableDimensions: ["factual-correctness", "sketch-code-grounding", "learner-answer-diagnosis", "precision"], + dimensionEvidenceLimitations: { + "instructional-usefulness": "No learner history beyond this synthetic answer is represented.", + scaffolding: "No learner history beyond this synthetic answer is represented.", + "dialogic-progression": "The case has no preceding dialog turn.", + "non-repetition": "The case has no preceding dialog turn.", + "difficulty-appropriateness": "The case scopes any judgment to its fixed synthetic input only.", + }, + evidenceSources: [ + { + id: "sketch", + kind: "repository/sketch", + reference: INPUT_PULLUP_SKETCH_REF, + digest: INPUT_PULLUP_SKETCH_DIGEST, + }, + { + id: "input-pullup-reference-draft", + kind: "draft-factual-reference", + reference: factualReferenceBundle.id, + digest: factualDigest, + }, + ], + permittedExternalKnowledge: false, + factualReferenceBundle, + humanReference: { status: "pending-review" }, + }; +} + +export function validSemanticCorpusSource(frozenPreTurnContext: Record) { + return { + schemaVersion: "tutor-quality-semantic-corpus-v1", + corpusId: "tutor-quality-semantic", + corpusVersion: 1, + cases: [validSemanticCaseSource(frozenPreTurnContext)], + }; +} + +export function validSemanticCorpusReferences() { + return { + sketches: new Map([[INPUT_PULLUP_SKETCH_REF, INPUT_PULLUP_SKETCH]]), + courseContent: new Map(), + }; +} + +export function validStageAScenario(overrides: Partial = {}): TutorQualityEvaluationScenario { + return { + id: "TQ-SEM-001", + corpusId: "tutor-quality-semantic", + corpusVersion: 1, + sketchRef: INPUT_PULLUP_SKETCH_REF, + sketch: INPUT_PULLUP_SKETCH, + turns: [{ + kind: "dialog", + question: INPUT_PULLUP_QUESTION, + answer: INPUT_PULLUP_ANSWER, + bindsToQuestion: INPUT_PULLUP_QUESTION, + difficulty: 30, + history: [], + }], + ...overrides, + }; +} + +export function validStageATranscript( + scenario: TutorQualityEvaluationScenario = validStageAScenario(), + overrides: Partial = {}, +): TutorQualityTranscript { + const turn = scenario.turns[0]!; + return { + schemaVersion: "tutor-quality-transcript-v1", + runId: "tq2a-20260928T000000Z-test-run", + evaluationIdentity: "a".repeat(64), + metadata: { + providerId: "fixture-provider", + requestedModel: "fixed-model-v1", + returnedModels: ["fixed-model-v1"], + promptRevision: { id: "tutor-prompt-v1", templateDigest: "b".repeat(64) }, + courseContentRevision: scenario.courseContent?.revision ?? "free-tutor", + corpusId: scenario.corpusId, + corpusVersion: scenario.corpusVersion, + gitSha: "c".repeat(40), + gitState: "clean", + sampleIndex: 0, + sampleCount: 1, + sampleStartedAt: "2026-09-28T00:00:00.000Z", + sampleDurationMs: 1, + providerCalls: { total: 2, modelListCalls: 1, generationCalls: 1 }, + maxCalls: 10, + }, + scenario: { + id: scenario.id, + corpusId: scenario.corpusId, + corpusVersion: scenario.corpusVersion, + sketchRef: scenario.sketchRef, + sketch: scenario.sketch, + syntheticTurns: scenario.turns, + }, + turns: [{ + index: 0, + input: turn, + startedAt: "2026-09-28T00:00:00.000Z", + durationMs: 1, + providerCalls: { total: 2, modelListCalls: 1, generationCalls: 1 }, + finalTutorResult: { feedback: "Richtig erkannt.", question: "Was bewirkt INPUT_PULLUP?" }, + returnedModel: "fixed-model-v1", + deterministicChecks: [], + }], + deterministicChecks: [], + executionStatus: "completed", + invariantViolations: [], + ...overrides, + }; +} diff --git a/tests/server/services/tutor/evaluation/semantic/stage2a-adapter.test.ts b/tests/server/services/tutor/evaluation/semantic/stage2a-adapter.test.ts new file mode 100644 index 00000000..9a6b8ebb --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/stage2a-adapter.test.ts @@ -0,0 +1,82 @@ +import { describe, expect, it } from "vitest"; +import { parseSemanticCorpus } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-corpus"; +import { createFrozenPreTurnContext } from "../../../../../../server/services/tutor/evaluation/semantic/frozen-context"; +import { toStage2AScenario } from "../../../../../../server/services/tutor/evaluation/semantic/stage2a-adapter"; +import { + INPUT_PULLUP_ANSWER, + INPUT_PULLUP_QUESTION, + INPUT_PULLUP_SKETCH, + INPUT_PULLUP_SKETCH_REF, + validFrozenPreTurnContextInput, + validSemanticCorpusReferences, + validSemanticCorpusSource, +} from "./semantic-test-fixtures"; + +function parsedCorpus() { + return parseSemanticCorpus(validSemanticCorpusSource(createFrozenPreTurnContext(validFrozenPreTurnContextInput())), validSemanticCorpusReferences()); +} + +describe("Stage-2A Case Adapter", () => { + it("translates TQ-SEM-001 to the exact existing Stage-A scenario and dialog turn contract", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + + expect(adapted.scenario).toEqual({ + id: "TQ-SEM-001", + corpusId: "tutor-quality-semantic", + corpusVersion: 1, + sketchRef: INPUT_PULLUP_SKETCH_REF, + sketch: INPUT_PULLUP_SKETCH, + turns: [{ + kind: "dialog", + question: INPUT_PULLUP_QUESTION, + answer: INPUT_PULLUP_ANSWER, + bindsToQuestion: INPUT_PULLUP_QUESTION, + difficulty: 30, + history: [], + }], + }); + expect(adapted.provenance).toMatchObject({ + kind: "adapter-generated", + semanticCorpusId: "tutor-quality-semantic", + semanticCorpusVersion: 1, + semanticCaseId: "TQ-SEM-001", + scenarioId: "TQ-SEM-001", + frozenPreTurnContextDigest: semanticCase.frozenPreTurnContext.digest, + }); + }); + + it("keeps rubric, human labels, exposure, candidate, Judge, and Stage-B mapping out of Stage-A input", () => { + const adapted = toStage2AScenario(parsedCorpus(), parsedCorpus().cases[0]!, validSemanticCorpusReferences()); + const scenarioJson = JSON.stringify(adapted.scenario); + + expect(adapted.scenario).not.toHaveProperty("rubric"); + expect(adapted.scenario).not.toHaveProperty("humanReference"); + expect(adapted.scenario).not.toHaveProperty("candidateIdentity"); + expect(adapted.scenario).not.toHaveProperty("exposure"); + expect(adapted.scenario).not.toHaveProperty("judge"); + expect(adapted.scenario).not.toHaveProperty("mappingIdentity"); + expect(scenarioJson).not.toContain("correct-answer-rejected"); + expect(scenarioJson).not.toContain("pending-review"); + expect(adapted.scenario).not.toHaveProperty("expected"); + }); + + it("rejects a changed sketch or invalid Frozen Context before returning Stage-A inputs", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const references = validSemanticCorpusReferences(); + references.sketches.set(INPUT_PULLUP_SKETCH_REF, "changed bytes"); + + expect(() => toStage2AScenario(corpus, semanticCase, references)).toThrow(/sketch|digest/i); + expect(() => toStage2AScenario(corpus, { + ...semanticCase, + frozenPreTurnContext: { ...semanticCase.frozenPreTurnContext, learnerAnswer: { ...semanticCase.frozenPreTurnContext.learnerAnswer, bindsToQuestion: "another question" } }, + }, validSemanticCorpusReferences())).toThrow(/context|bind/i); + }); + + it("rejects a forged Semantic Corpus digest before returning a Stage-A scenario", () => { + const corpus = parsedCorpus(); + expect(() => toStage2AScenario({ ...corpus, digest: "f".repeat(64) }, corpus.cases[0]!, validSemanticCorpusReferences())).toThrow(/corpus digest/i); + }); +}); diff --git a/tests/server/services/tutor/evaluation/semantic/stage2a-transcript.test.ts b/tests/server/services/tutor/evaluation/semantic/stage2a-transcript.test.ts new file mode 100644 index 00000000..f7ae7ed3 --- /dev/null +++ b/tests/server/services/tutor/evaluation/semantic/stage2a-transcript.test.ts @@ -0,0 +1,285 @@ +import { describe, expect, it } from "vitest"; +import { parseSemanticCorpus } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-corpus"; +import { createFrozenPreTurnContext } from "../../../../../../server/services/tutor/evaluation/semantic/frozen-context"; +import { toStage2AScenario } from "../../../../../../server/services/tutor/evaluation/semantic/stage2a-adapter"; +import { + createExistingTranscriptCompatibilityMapping, + stage2bTranscriptDigest, + validateStage2ATranscript, +} from "../../../../../../server/services/tutor/evaluation/semantic/stage2a-transcript"; +import { + canonicalStage2ATranscriptJson, + createStage2ATranscriptReference, +} from "../../../../../../server/services/tutor/evaluation/semantic/transcript-reference"; +import { canonicalSemanticDigest } from "../../../../../../server/services/tutor/evaluation/semantic/semantic-canonical"; +import { + validFrozenPreTurnContextInput, + validSemanticCorpusReferences, + validSemanticCorpusSource, + validStageAScenario, + validStageATranscript, +} from "./semantic-test-fixtures"; + +function parsedCorpus() { + return parseSemanticCorpus(validSemanticCorpusSource(createFrozenPreTurnContext(validFrozenPreTurnContextInput())), validSemanticCorpusReferences()); +} + +describe("Stage-2A transcript compatibility and Stage-B digest", () => { + it("validates an adapter-generated transcript without changing its Stage-A identity or invariant array", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const invariantViolations = [{ code: "raw-response-repaired", source: "raw-provider" as const, turnIndex: 0 }]; + const transcript = validStageATranscript(adapted.scenario, { invariantViolations }); + const transcriptBytes = JSON.stringify(transcript); + const evaluationIdentity = transcript.evaluationIdentity; + + const result = validateStage2ATranscript(transcript, semanticCase, adapted.provenance); + + expect(result).toMatchObject({ valid: true, semanticEligibility: "eligible" }); + if (!result.valid) throw new Error("expected compatible Stage-A transcript"); + expect(result.transcriptDigest).toBe(stage2bTranscriptDigest(transcript)); + expect(result.transcript.evaluationIdentity).toBe(evaluationIdentity); + expect(result.transcript.invariantViolations).toBe(invariantViolations); + expect(JSON.stringify(transcript)).toBe(transcriptBytes); + }); + + it("accepts different legacy case/scenario IDs only after explicit full-context compatibility mapping", () => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ id: "legacy-input-pullup", corpusId: "legacy-anchor-corpus", corpusVersion: 3 }); + const transcript = validStageATranscript(sourceScenario); + const mapping = createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/run-123/transcript-legacy-input-pullup-0.json", + transcript, + sourceScenario, + semanticCase, + }); + const bytesBefore = JSON.stringify(transcript); + + const result = validateStage2ATranscript(transcript, semanticCase, mapping); + + expect(result).toMatchObject({ valid: true, semanticEligibility: "eligible" }); + if (!result.valid) throw new Error("expected explicit compatible mapping"); + const reference = createStage2ATranscriptReference(transcript, result.transcriptDigest, result.compatibilityMapping); + expect(reference.stageAEvaluationIdentity).toBe("a".repeat(64)); + expect(reference.transcriptDigest).toBe(result.transcriptDigest); + expect(reference.compatibilityMappingIdentity).toBe(mapping.identity); + expect(transcript.scenario.id).toBe("legacy-input-pullup"); + expect(transcript.evaluationIdentity).toBe("a".repeat(64)); + expect(JSON.stringify(transcript)).toBe(bytesBefore); + }); + + it.each([0x00, 0x1f, 0x7f])("rejects ASCII control character with code point %i in existing-transcript source identities", (code) => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ id: "legacy-input-pullup", corpusId: "legacy-anchor-corpus", corpusVersion: 3 }); + const transcript = validStageATranscript(sourceScenario); + const validMapping = createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-identity", + transcript, + sourceScenario, + semanticCase, + }); + const sourceId = `stage-a/source${String.fromCharCode(code)}identity`; + + expect(() => createExistingTranscriptCompatibilityMapping({ + sourceId, + transcript, + sourceScenario, + semanticCase, + })).toThrow(/control characters/i); + + const { identity: _identity, ...mappingContent } = { ...validMapping, sourceId }; + const invalidMapping = { ...mappingContent, identity: canonicalSemanticDigest(mappingContent) }; + expect(validateStage2ATranscript(transcript, semanticCase, invalidMapping)).toMatchObject({ + valid: false, + reason: "existing-transcript-source-identity-missing-or-invalid", + }); + }); + + it("rejects an expected learning phase that only string-coerces to a supported value", () => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ + id: "legacy-input-pullup", + corpusId: "legacy-anchor-corpus", + corpusVersion: 3, + expected: { learningPhase: ["LEARN"] as unknown as "LEARN" }, + }); + const transcript = validStageATranscript(sourceScenario); + + expect(() => createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-identity", + transcript, + sourceScenario, + semanticCase, + })).toThrow(/invalid or unsupported fields/i); + }); + + it.each(["LEARN", "DEEPEN", "EXPAND"] as const)("accepts the exact Stage-A learning phase %s", (learningPhase) => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ + id: "legacy-input-pullup", + corpusId: "legacy-anchor-corpus", + corpusVersion: 3, + expected: { learningPhase }, + }); + const transcript = validStageATranscript(sourceScenario); + const mapping = createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-identity", + transcript, + sourceScenario, + semanticCase, + }); + + expect(validateStage2ATranscript(transcript, semanticCase, mapping)).toMatchObject({ valid: true }); + }); + + it("treats an omitted legacy dialog history as Stage-A's empty-history default", () => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ + id: "legacy-no-history-field", + corpusId: "legacy-anchor-corpus", + corpusVersion: 3, + turns: [{ + kind: "dialog", + question: semanticCase.frozenPreTurnContext.question, + answer: semanticCase.frozenPreTurnContext.learnerAnswer.text, + bindsToQuestion: semanticCase.frozenPreTurnContext.question, + difficulty: semanticCase.frozenPreTurnContext.difficulty, + }], + }); + const transcript = validStageATranscript(sourceScenario); + const mapping = createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-with-default-history", + transcript, + sourceScenario, + semanticCase, + }); + + expect(sourceScenario.turns[0]).not.toHaveProperty("history"); + expect(validateStage2ATranscript(transcript, semanticCase, mapping)).toMatchObject({ valid: true, semanticEligibility: "eligible" }); + }); + + it("rejects missing mappings and any unproven question, answer, sketch, order, or context change", () => { + const semanticCase = parsedCorpus().cases[0]!; + const sourceScenario = validStageAScenario({ id: "legacy-input-pullup", corpusId: "legacy-anchor-corpus", corpusVersion: 3 }); + const transcript = validStageATranscript(sourceScenario); + const mapping = createExistingTranscriptCompatibilityMapping({ sourceId: "stage-a/source-identity", transcript, sourceScenario, semanticCase }); + + expect(validateStage2ATranscript(transcript, semanticCase, undefined)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ ...transcript, scenario: { ...transcript.scenario, sketch: "other sketch" } }, semanticCase, mapping)).toMatchObject({ valid: false }); + expect(() => createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-identity", + transcript, + sourceScenario: { ...sourceScenario, turns: [{ kind: "dialog", question: "Other?", answer: "Answer", bindsToQuestion: "Other?", difficulty: 30, history: [] }] }, + semanticCase, + })).toThrow(/compatib|context|question/i); + expect(() => createExistingTranscriptCompatibilityMapping({ + sourceId: "stage-a/source-identity", + transcript, + sourceScenario: { ...sourceScenario, unexpected: "must not enter a Stage-B mapping" } as typeof sourceScenario, + semanticCase, + })).toThrow(/unsupported fields/i); + expect(validateStage2ATranscript(transcript, semanticCase, { ...mapping, unexpected: "must be rejected" })).toMatchObject({ valid: false, reason: "stage-b-transcript-mapping-schema-invalid" }); + }); + + it("keeps a compatible non-completed Stage-A sample not-evaluated and preserves invariant violations", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const invariantViolations = [{ code: "not-run", source: "scenario" as const }]; + const transcript = validStageATranscript(adapted.scenario, { executionStatus: "technical-failure", invariantViolations }); + + const result = validateStage2ATranscript(transcript, semanticCase, adapted.provenance); + + expect(result).toMatchObject({ valid: true, semanticEligibility: "not-evaluated" }); + if (!result.valid) throw new Error("expected structurally compatible Stage-A transcript"); + expect(result.transcript.invariantViolations).toBe(invariantViolations); + }); + + it("admits a not-run Stage-A sample with no executed turns for reliability accounting only", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const transcript = validStageATranscript(adapted.scenario, { + executionStatus: "not-run", + turns: [], + deterministicChecks: [], + metadata: { + ...validStageATranscript(adapted.scenario).metadata, + providerCalls: { total: 0, modelListCalls: 0, generationCalls: 0 }, + }, + }); + + const result = validateStage2ATranscript(transcript, semanticCase, adapted.provenance); + + expect(result).toMatchObject({ valid: true, semanticEligibility: "not-evaluated" }); + }); + + it("canonicalizes object keys in UTF-8 order while preserving array order and not mutating the Stage-A object", () => { + const transcript = validStageATranscript(); + const withUnicodeKeys = { + ...transcript, + metadata: { ...transcript.metadata, "ä": "umlaut", z: "ascii" }, + } as typeof transcript; + const before = JSON.stringify(withUnicodeKeys); + const canonical = canonicalStage2ATranscriptJson(withUnicodeKeys); + const reorderedKeys = { + ...withUnicodeKeys, + metadata: { ...transcript.metadata, z: "ascii", "ä": "umlaut" }, + } as typeof transcript; + + expect(canonical).toContain('"z":"ascii","ä":"umlaut"'); + expect(stage2bTranscriptDigest(withUnicodeKeys)).toBe(stage2bTranscriptDigest(reorderedKeys)); + const ordered = { + ...transcript, + turns: [ + transcript.turns[0]!, + { ...transcript.turns[0]!, index: 1, input: { ...transcript.turns[0]!.input, question: "Second question", answer: "Second answer", bindsToQuestion: "Second question" } }, + ], + } as typeof transcript; + expect(stage2bTranscriptDigest(ordered)).not.toBe(stage2bTranscriptDigest({ ...ordered, turns: [...ordered.turns].reverse() })); + expect(JSON.stringify(withUnicodeKeys)).toBe(before); + }); + + it("rejects a transcript with a changed question-to-answer binding or a missing Stage-A identity", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const transcript = validStageATranscript(adapted.scenario); + + expect(validateStage2ATranscript({ ...transcript, turns: [{ ...transcript.turns[0]!, input: { ...transcript.turns[0]!.input, bindsToQuestion: "different" } }] }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ ...transcript, evaluationIdentity: "" }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ ...transcript, metadata: { ...transcript.metadata, promptRevision: undefined } }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + }); + + it("rejects malformed invariant and deterministic-check entries, not only malformed arrays", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const transcript = validStageATranscript(adapted.scenario); + + expect(validateStage2ATranscript({ ...transcript, invariantViolations: [null] }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ ...transcript, invariantViolations: [{ code: "", source: "invented" }] }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ ...transcript, deterministicChecks: [null] }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ + ...transcript, + deterministicChecks: [{ name: "", passed: "yes" }], + }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + expect(validateStage2ATranscript({ + ...transcript, + turns: [{ ...transcript.turns[0]!, deterministicChecks: [null] }], + }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + }); + + it("rejects a completed transcript whose normalized final Tutor result is not valid Tutor content", () => { + const corpus = parsedCorpus(); + const semanticCase = corpus.cases[0]!; + const adapted = toStage2AScenario(corpus, semanticCase, validSemanticCorpusReferences()); + const transcript = validStageATranscript(adapted.scenario); + + expect(validateStage2ATranscript({ + ...transcript, + turns: [{ ...transcript.turns[0]!, finalTutorResult: {} }], + }, semanticCase, adapted.provenance)).toMatchObject({ valid: false }); + }); +});