VOL. 03· EVALUATION REPORT· SEP 2026

Frontier vs. open weights,
on real questions from players.

Questions that players asked in the ASL Discord rules channel, answered by a frontier model and an open-weights model, and scored against the rulebook.

01
Findings

The frontier model is better — at about 70× the price.

1. gpt-6 answered 99% correctly (72 of 73). deepseek-v4-flash, an open-weights model, answered 93% (68 of 73).

2. deepseek-v4-flash cost a measured 0.34¢ per question. gpt-6 is estimated at ~25¢ — about 70× more.

3. deepseek-v4-flash was slow: 67s per question, against ~25s for gpt-6. The time goes into its tool loop — it made about 5 rulebook lookups per question, and each one is another round trip that re-sends a growing context. When it answered without a lookup it took 15s.

02
Model comparison

Accuracy, cost, latency
side by side.

TABLE 01 — questions from the ASL Discord rules channel · scored against rulebook-cited answers · human-reviewed

Model Passed Accuracy Cost / q Avg time Date run
gpt-6
frontier
72 / 73 99% ~25¢† ~25s† Sep 16 VIEW →
deepseek-v4-flash
open weights
68 / 73 93% 0.34¢ 67s Sep 16 VIEW →

The questions — each one comes from a thread in the ASL Discord rules channel (2025–26), reworded into a single self-contained question. The expected answer was written against the rulebook and the official Q&A, with section citations. Several expected answers were corrected during review, when a model's citations showed the rulebook said otherwise.

Scoring — Claude Fable 5.1 judged every answer assertion by assertion against the expected answer; borderline verdicts were settled by human review. Open any run to read the questions, both answers and the verdicts.

Cost / q & Avg time — deepseek-v4-flash ran through the same pipeline as live chat, so its eval cost is measured: what OpenRouter billed for the run, divided by 73 questions.

† gpt-6 was run through the Codex CLI on a subscription, with a local search over the same rulebook text and the same calculators, so no per-question token usage or timing was recorded. Its cost is an estimate: the production prompt (~22k input tokens, ~0.7k output) at the GPT-6 list price of $10 / $50 per million tokens. At the token volume deepseek's tool loop used on this set it would be ~87¢. Its time is a run average taken from the Codex session logs, and local rulebook search is instant, so it would be slower through the production pipeline. The two models did not run an identical retrieval setup, so read the accuracy gap as indicative.

Sample size — 73 questions. A 4-question gap is a signal, not a measurement to the decimal.

03
Production usage

Cost and latency,
in flight.

Daily averages from live chat · text-only baselines, image variants dashed, agentic (tool-using) variants dotted

Cost / question ($)
Response time (s)