Thirty expected answers were independently authored by the evaluator and executed against real read-only sources. Owner ratification is pending. No product-generated SQL was used to author the expected answers; no second model checked the SQL.
| Measure | Count |
|---|---|
| Questions | 30 |
| Families | 15 |
| Repeats | 3 |
| Workers | 2 |
| Scored conversations | 90 |
| Expected answers computed | 30 |
| Questions flagged for alternate readings | 10 |
| Alternate readings computed | 13 |
| Owner-ratified answers | 0 |
| Matched answers | 0 |
| Mismatched answers | 41 |
| Refused after follow-through | 3 |
| Errors | 46 |
| Initial clarifications | 30 |
Match rate: 0.00%; mismatch rate: 45.56%; clarification rate: 33.33%. These rates use all ninety conversations as their denominator. Errors and refusals are neither matches nor mismatches.
| Final outcome | Count |
|---|---|
| clarified_then_match | 0 |
| clarified_then_mismatch | 17 |
| error | 46 |
| match | 0 |
| mismatch | 24 |
| refused | 3 |
| Outcome after clarification | Count |
|---|---|
| clarified_then_mismatch | 17 |
| error | 12 |
| refused | 1 |
| Defect among returned mismatches | Count |
|---|---|
| wrong_grain | 9 |
| wrong_join_or_key | 3 |
| definition_absent | 18 |
| wrong_window | 8 |
| definition_present_not_applied | 3 |
| Relevant definition coverage per conversation | Count |
|---|---|
| all | 18 |
| no_generation_prompt | 1 |
| none | 24 |
| partial | 47 |
| Prompt route | Present | Generation prompts | Rate |
|---|---|---|---|
| default | 104 | 104 | 100.00% |
| resolver | 0 | 104 | 0.00% |
| catalog | 104 | 104 | 100.00% |
| exemplar | 0 | 104 | 0.00% |
Captured full prompts: 104/104. Compiled resolver filter entries: 0. Route presence counts any recorded material. Relevant coverage checks the required rules in the recorded prompt, including the skill bundle; bundle presence alone does not prove rule presence.
Generation prompts by source connection: ct: 18, gold: 86, unknown: 0.
| Family | Questions |
|---|---|
| F01 | 2 |
| F05 | 2 |
| F07 | 2 |
| F08 | 2 |
| F09 | 2 |
| F10 | 2 |
| F13 | 2 |
| F14 | 2 |
| F17 | 2 |
| F18 | 2 |
| F19 | 2 |
| F21 | 2 |
| F25 | 2 |
| F26 | 2 |
| F34 | 2 |
Conversation latency median: 14.435 seconds; p95: 27.557 seconds (nearest-rank). Latency covers active first-turn submission through the final answer, including clarification follow-through and product retries. It includes failed outcomes.
Runtime qualification: loading the private launcher environment alone did not enable the SQL provider. The deployed checkpoint lacked the z.ai SQL transport even though Compose was configured for it. The evaluation branch adds that transport and explicitly selects z.ai; prompt hashes remain unchanged. The smoke completed and returned served model glm-5.3. Verifier off, defaults on, unchanged SQL prompt construction. This is a measured branch run with a disclosed transport repair, not proof of the unmodified deployed checkpoint.
Clock: expected SQL is anchored to the original F1 source preflight on 13 September 2026. Product conversations use the live clock on that date; the sources are live read-only replicas, not a frozen historical snapshot. Owner ambiguity choices may change the approved reading; primary and alternate expected results remain private.
Oracle audit: one primary answer was corrected from the ratified high-value rule during the run; its original seal is preserved. Two attribution alternatives were expanded to include every requested fact, and one explicit-reading false ambiguity was removed. Those changes used written rules and fresh read-only execution, without product SQL as their source.
Clarification harness receipt: three denominator-control submissions were rejected by request validation before the second turn was admitted. The binding shape was corrected and only those unadmitted follow-through turns were completed; initial questions and completed answers were not rerun. Their latency sums the active turns and excludes the administrative pause.
Arithmetic-only review: PASS. The public artifact uses an allowlisted aggregate projection and passes the email, at-sign, client-name and numeric-identifier scans. The arithmetic reviewer received only aggregate counts and performance timings; its output passed the same private-pattern scan.
Manifest SHA-256: 8e9c0c21e400aed7a1736f475aea41810bbce0f1164f5ffead8095217f704694
Oracle seal SHA-256: 05eef6df30cb108c1daf51c728de0851982f5afb7dfb041a3f2d3f20cea0dd24
Attempts manifest SHA-256: 58b38b8567895726797f8f8c70bcee94dad159b0ff2638b8c2007084192a4e8d