Companion concept to the semantic score contracts (#2812 Phase 4, #2908 precedent), extracted from the repowise assessment (temporal-crucible research notes on #2982).
Their governance structure, adapted to ours: a calibration contract is the empirical twin of a semantic contract — a signal may carry weight in a gated risk formula only if it demonstrates defect-lift beyond file/function size on event ground truth (the L2-logistic-with-size-control recipe; temporal-crucible's JSONL dataset is the corpus). Markers that fail are floored or confined to non-defect pillars — publicly, the way repowise floors error_handling/low_cohesion and prints why.
Sobering context that motivates this: repowise's own head-to-head reports partial correlation beyond size of −0.148 after 26 calibrated markers (CodeScene −0.137) — the entire field's defect signal beyond LOC is a sliver, and honest formulas should reflect which inputs actually earn their weights. Structure: tiered permission (defect-scoring / capped pillars / advisory-never-deducts), per-category caps, and "only the number carrying published claims may move."
Blocked on: temporal-crucible repo #2 (single-project calibration would overfit curl), and the commensurability audit remains the semantic-side guard. This issue is the design placeholder so the idea has an owner.
Companion concept to the semantic score contracts (#2812 Phase 4, #2908 precedent), extracted from the repowise assessment (temporal-crucible research notes on #2982).
Their governance structure, adapted to ours: a calibration contract is the empirical twin of a semantic contract — a signal may carry weight in a gated risk formula only if it demonstrates defect-lift beyond file/function size on event ground truth (the L2-logistic-with-size-control recipe; temporal-crucible's JSONL dataset is the corpus). Markers that fail are floored or confined to non-defect pillars — publicly, the way repowise floors
error_handling/low_cohesionand prints why.Sobering context that motivates this: repowise's own head-to-head reports partial correlation beyond size of −0.148 after 26 calibrated markers (CodeScene −0.137) — the entire field's defect signal beyond LOC is a sliver, and honest formulas should reflect which inputs actually earn their weights. Structure: tiered permission (defect-scoring / capped pillars / advisory-never-deducts), per-category caps, and "only the number carrying published claims may move."
Blocked on: temporal-crucible repo #2 (single-project calibration would overfit curl), and the commensurability audit remains the semantic-side guard. This issue is the design placeholder so the idea has an owner.