Insights F1B: thirty-question oracle and measured product run

Thirty expected answers were independently authored by the evaluator and executed against real read-only sources. Owner ratification is pending. No product-generated SQL was used to author the expected answers; no second model checked the SQL.

MeasureCount
Questions30
Families15
Repeats3
Workers2
Scored conversations90
Expected answers computed30
Questions flagged for alternate readings10
Alternate readings computed13
Owner-ratified answers0
Matched answers0
Mismatched answers41
Refused after follow-through3
Errors46
Initial clarifications30

Match rate: 0.00%; mismatch rate: 45.56%; clarification rate: 33.33%. These rates use all ninety conversations as their denominator. Errors and refusals are neither matches nor mismatches.

Final outcomeCount
clarified_then_match0
clarified_then_mismatch17
error46
match0
mismatch24
refused3
Outcome after clarificationCount
clarified_then_mismatch17
error12
refused1
Defect among returned mismatchesCount
wrong_grain9
wrong_join_or_key3
definition_absent18
wrong_window8
definition_present_not_applied3
Relevant definition coverage per conversationCount
all18
no_generation_prompt1
none24
partial47
Prompt routePresentGeneration promptsRate
default104104100.00%
resolver01040.00%
catalog104104100.00%
exemplar01040.00%

Captured full prompts: 104/104. Compiled resolver filter entries: 0. Route presence counts any recorded material. Relevant coverage checks the required rules in the recorded prompt, including the skill bundle; bundle presence alone does not prove rule presence.

Generation prompts by source connection: ct: 18, gold: 86, unknown: 0.

FamilyQuestions
F012
F052
F072
F082
F092
F102
F132
F142
F172
F182
F192
F212
F252
F262
F342

Conversation latency median: 14.435 seconds; p95: 27.557 seconds (nearest-rank). Latency covers active first-turn submission through the final answer, including clarification follow-through and product retries. It includes failed outcomes.

Runtime qualification: loading the private launcher environment alone did not enable the SQL provider. The deployed checkpoint lacked the z.ai SQL transport even though Compose was configured for it. The evaluation branch adds that transport and explicitly selects z.ai; prompt hashes remain unchanged. The smoke completed and returned served model glm-5.3. Verifier off, defaults on, unchanged SQL prompt construction. This is a measured branch run with a disclosed transport repair, not proof of the unmodified deployed checkpoint.

Clock: expected SQL is anchored to the original F1 source preflight on 13 September 2026. Product conversations use the live clock on that date; the sources are live read-only replicas, not a frozen historical snapshot. Owner ambiguity choices may change the approved reading; primary and alternate expected results remain private.

Oracle audit: one primary answer was corrected from the ratified high-value rule during the run; its original seal is preserved. Two attribution alternatives were expanded to include every requested fact, and one explicit-reading false ambiguity was removed. Those changes used written rules and fresh read-only execution, without product SQL as their source.

Clarification harness receipt: three denominator-control submissions were rejected by request validation before the second turn was admitted. The binding shape was corrected and only those unadmitted follow-through turns were completed; initial questions and completed answers were not rerun. Their latency sums the active turns and excludes the administrative pause.

Arithmetic-only review: PASS. The public artifact uses an allowlisted aggregate projection and passes the email, at-sign, client-name and numeric-identifier scans. The arithmetic reviewer received only aggregate counts and performance timings; its output passed the same private-pattern scan.

Manifest SHA-256: 8e9c0c21e400aed7a1736f475aea41810bbce0f1164f5ffead8095217f704694

Oracle seal SHA-256: 05eef6df30cb108c1daf51c728de0851982f5afb7dfb041a3f2d3f20cea0dd24

Attempts manifest SHA-256: 58b38b8567895726797f8f8c70bcee94dad159b0ff2638b8c2007084192a4e8d