Insights F1: incomplete oracle, retained operational run

The human oracle was not established. All correctness comparisons remain unverified. The selected local runtime had no usable SQL provider route: it selected a provider without a configured key. These counts describe retained operational outcomes, not deployed-model accuracy.

MeasureCount
Questions30
Families15
Repeats3
Workers2
Submitted attempts90
Returned answers0
Clarification responses, counted as refused30
Provider errors60
Expected answers available0
Human verified answers0
Unverified comparisons90
Observed generation attempts60
Attempts with provenance60

Observed refusal rate: 33.33%. Match and mismatch rates: unavailable; there are zero valid oracle comparisons. No semantic defect classes can be ranked. The operational failure groups are provider errors and clarification responses.

FamilyQuestions
F102
F182
F252
F012
F082
F262
F092
F342
F172
F142
F132
F192
F052
F212
F072
Prompt routeAny material presentDenominatorPresence rateKnowledge-entry material present
default6060100.00%60
resolver0600.00%0
catalog6060100.00%0
exemplar0600.00%0

Presence is measured among observed generation attempts. Catalog material in this run consists of skill bundles, identified by content hash. It does not establish that a relevant individual definition appeared inside the bundle. Clarifications occurred before SQL generation and have no generation prompt. Relevant-definition coverage is unverified; no human oracle identified the necessary per-question definitions.

Compiled resolver filters: 0. Unknown leg labels in the retained run: 15. A subsequent instrumentation-only fix records trusted connection identity for empty schemas; retained attempts were not rerun or relabelled.

Latency statisticMilliseconds
minimum1080.968
median1235.812
p95_nearest_rank1976.701
maximum2537.528

Latency is client submission through terminal job and answer fetch. It includes polling and represents errors/clarifications, not successful answer latency.

Limits: no human-authored oracle SQL was supplied, and model authorship was prohibited. No expected SQL was written or executed. Ambiguity and alternate readings remain unassessed. Goal C had nine supplemental real-source queries, but their human authorship was not established. Both real sources passed read-only reachability checks. The local serving process had no BI sources; the harness used a read-only export of production source metadata with isolated SSH forwards. It inherited the local process provider environment and disabled startup reconciliation/index initialization. No synthetic data fixture was used. The local provider route was diagnosed after the run, and no attempts were rerun. The run cannot support claims about the operational production generator. The private pack accepts owner decisions locally; approvals do not retroactively supply SQL execution evidence.

Review: arithmetic and current public-artifact privacy passed. The reviewer initially echoed private scan patterns into its inspection output; the local console was removed, but retraction from the inspection context is unverified. The required review privacy boundary therefore failed. A categorical-export validation blocker was fixed and regression tested after that single review; the fix was not reviewed again by a model.

Manifest SHA-256: `8e9c0c21e400aed7a1736f475aea41810bbce0f1164f5ffead8095217f704694`

Attempts manifest SHA-256: `797400b8cea517f057686846424076cc6818b63ec0cfef1c2e2acc97d7010d7b`

Measured checkpoint: `541a463b`. Production and the main working tree were not changed.