
2026/09/22 22:01
Show HN: JevBench,型付き決定モデルのための再現可能なベンチマーク
RSS: https://news.ycombinator.com/rss
要約▶
本文
JevBench v1.3.0JevBench v1.3.0 · our own benchmarkJev-class modelsJevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.Version 1.3.0 measures 52 systems on the unchanged 534 decisions, including 220 hard ones, and ranks them by the JevBench Score. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.Scored 21 Sept 2026 · protocol jevbench::v1.2 · 72 easy + 96 standard + 146 judge + 220 hard decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 20fce8e6e4f0… · v1.0 resultsJevBench v1.3.0 · 534 decisions per systemJevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.0402SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · $0.022 est.3djev (Maisa, diffusion-gemma)†73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.4Winnow-12B Q8†71.2I 82 · C 72 · S 82 · K 53 · $0.014 est.66.7%31.3%34.2%—3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 sour CPUby Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked0.14.6none (label only)59.958.7~$0.024 est.47.2%16.7%31.5%7.7%1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU† Notes on 39 marked systems — how each was run† djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).† Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.† reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.† jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.† decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.† decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.† open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.† SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.† ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.† Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.† reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.† LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.† kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible $0.037 est.5reflex 4B†70.3I 80 · C 75 · S 68 · K 60 · $2.669 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction APIby Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked1.113.5none (label only)52.865.3$0.022 est.6jqv†68.6I 79 · C 79 · S 75 · K 47 · $0.0033 est.100.0%99.0%97.3%70.5%0.39 s rawp95 0.45 s rawproduction APIPartial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked24.867.492.161.30.0$0.056 est.7decision-machine-1†68.3I 62 · C 70 · S 93 · K 54 · $0.0358decider-35b-a3b†67.6I 80 · C 72 · S 81 · K 45 · $0.0010 est.27.8%30.2%21.9%31.8%0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 sour RunPod GPUHonorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.by mrmps (@michael_chomsky)classifier.dev†fast tierhonorable mention · not ranked83.685.177.987.684.3$0.067 est.9open-alternative-jev (Qwen3.5-4B)†67.0I 64 · C 63 · S 83 · K 60 · $0.0039 est.83.3%47.9%50.0%33.2%0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 sour CPU45by MixedbreadMixedbread mxbai-rerank-base-v2†0.86.783.187.567.9$0.01244.4%33.3%26.7%40.0%0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 sour RunPod GPU46by BAAIBAAI bge-reranker-v2-m3†0.76.383.889.573.4$0.007743.1%36.5%8.9%36.8%0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 sour RunPod GPU47by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base†0.34.676.890.669.6$0.01033.3%39.6%30.1%33.6%0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 sour RunPod GPU48by AltSlate LabsCerto v1†0.00.082.094.0100.0$0.022 est.10system-one-open66.6I 70 · C 57 · S 77 · K 65 · $0.0039 est.90.3%51.0%43.8%37.7%0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 sour CPU44by FastinoGLiNER2.5 small†Fastino, 74M13.825.647.277.882.4$0.015 est.11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · $0.0073 est.100.0%49.0%53.4%36.4%1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU43by FastinoGLiNER2.5 multi†Fastino, 287M16.627.756.167.882.4$0.066 est.12SimpleJev Qwen3.8-27B†66.3I 85 · C 81 · S 71 · K 39 · $0.0037 est.97.2%66.7%45.9%36.4%0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 sour CPU42by Kotoba Labsopen-jev-deberta-v3-largelocal CPU23.131.966.466.074.0$0.104 est.13ZeroEntropy zerank-2†66.0I 63 · C 76 · S 79 · K 50 · $0.04714GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.24215openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · $0.025 est.97.2%68.8%40.4%38.2%0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU41by FastinoGLiNER2†Fastino, gliner2.5-base24.035.623.771.883.1$0.131 est.16Qwen3-Reranker-4B†63.8I 64 · C 67 · S 79 · K 49 · $0.05017reflex-27b†63.3I 86 · C 86 · S 67 · K 32 · $0.0077 est.98.6%62.5%61.0%36.4%1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 sour CPU40by Aditya (isHeSatoshi)smalljev semantic-v9†27.435.158.979.857.9$0.181 est.18LitJev†62.7I 82 · C 84 · S 67 · K 34 · $0.0063 est.95.8%52.1%71.2%30.9%0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 sour RunPod GPU39by FastinoGLiNER2 large†29.640.124.361.773.3$0.163 est.19kev 0.6B†62.5I 52 · C 51 · S 76 · K 76 · $0.0037 est.86.1%65.6%61.0%38.2%0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 sour CPU38by Jared Palmerkev 0.5B†33.238.247.477.076.1$0.0063 est.20SimpleJev Qwen3.6-35B-A3B†62.5I 80 · C 67 · S 75 · K 38 · $0.0039 est.86.1%67.7%56.2%37.7%0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 sour CPU37by Hemant (heman10x)openJev Verdict†heman10x, ModernBERT-base 151M38.139.851.376.783.1$0.116 est.21djev†62.4I 81 · C 93 · S 75 · K 27 · $0.0066 est.87.5%62.5%71.2%33.2%0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 sour RunPod GPU36by Hemant (heman10x)openJev Verdict 1.4†38.938.674.178.182.4$0.274 est.22jev-local†61.8I 71 · C 69 · S 69 · K 43 · $0.249 est.100.0%79.2%88.4%42.7%0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU35by Deepan WadhwaOpenDecision†ModernBERT-large zero-shot40.640.856.179.975.3$0.077 est.23decider-2b†61.7I 61 · C 47 · S 83 · K 61 · $0.0029 est.94.4%72.9%69.2%34.1%0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 sour CPU34by Zefan Cai (@Zefan_Cai)Open-Jev 2B†Zefan Cai51.361.055.173.528.1$0.020 est.24Bespoke Nimble 9B†60.5I 78 · C 65 · S 79 · K 33 · $0.0060 est.100.0%76.0%61.6%37.7%0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 sour CPU33by Convai InnovationsLaya†Convai Innovations, ModernBERT-large 421M54.445.862.571.186.2$0.166 est.25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.26426OpenJev†60.0I 88 · C 70 · S 76 · K 28 · $0.089 est.100.0%90.6%91.8%50.0%0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU32by Logan Markewichjeff†Logan Markewich, GLiFormer 400M54.446.964.663.576.6$0.255 est.27kev 4B†59.7I 65 · C 42 · S 76 · K 62 · $0.249 est.100.0%90.6%81.5%60.9%0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 sour RunPod GPU31by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke54.870.336.884.441.5$0.019 est.28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.59429kev 8B†56.4I 69 · C 44 · S 75 · K 44 · $0.073 est.100.0%92.7%90.4%47.3%0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 sour RunPod GPU30by Zefan Cai (@Zefan_Cai)Open-Jev 9B†Zefan Cai55.071.263.372.028.1$0.073 est.30Open-Jev 9B†55.0I 71 · C 63 · S 72 · K 28 · $0.019 est.100.0%91.7%85.6%42.3%0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 sour RunPod GPU28by DeepSeekDeepSeek V4.1 Flashthinking default57.594.396.771.616.8$0.59498.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API29by Jared Palmerkev 8B†research preview56.469.444.274.944.0$0.249 est.31system-one54.8I 70 · C 37 · S 84 · K 41 · $0.255 est.100.0%100.0%94.5%78.2%0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 sour RunPod GPU27by Jared Palmerkev 4B†research preview59.764.842.075.761.8$0.089 est.32jeff†54.4I 47 · C 65 · S 63 · K 77 · $0.166 est.100.0%94.8%89.0%65.5%0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 sour RunPod GPU25by GoogleGemini 3.1 Flash-Lite60.185.668.181.827.4$0.264100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API26by razorback16OpenJev†thinking, BF1660.088.069.676.127.8$0.0060 est.33Laya†54.4I 46 · C 62 · S 71 · K 86 · $0.020 est.100.0%85.4%77.4%47.3%0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 sour RunPod GPU24by Bespoke LabsBespoke Nimble 9B†60.577.965.378.733.4$0.0029 est.34Open-Jev 2B†51.3I 61 · C 55 · S 73 · K 28 · $0.077 est.100.0%84.4%89.0%59.1%1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 sour RunPod GPU23by Mapikadecider-2b†61.761.246.683.261.0$0.249 est.35OpenDecision†40.6I 41 · C 56 · S 80 · K 75 · $0.274 est.95.8%99.0%80.1%77.7%0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU22by us (GitHub)jev-local†Qwen3.5-9B61.870.868.769.243.3$0.0066 est.36openJev Verdict 1.4†38.9I 39 · C 74 · S 78 · K 82 · $0.116 est.100.0%93.8%93.2%66.4%0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 sauthor's demo server21by David Villalon / Maisadjev†thinking62.480.892.775.226.9$0.0039 est.37openJev Verdict†38.1I 40 · C 51 · S 77 · K 83 · $0.0063 est.100.0%81.3%66.4%40.0%0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 sour RunPod GPU20by Featherless AISimpleJev Qwen3.6-35B-A3B†62.579.567.175.038.1$0.0037 est.38kev 0.5B†33.2I 38 · C 47 · S 77 · K 76 · $0.163 est.100.0%97.9%88.4%73.2%2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 sour RunPod GPU19by Jared Palmerkev 0.6B†research preview62.551.951.175.676.1$0.0063 est.39GLiNER2 large†29.6I 40 · C 24 · S 62 · K 73 · $0.181 est.100.0%95.8%95.9%75.9%1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 sour RunPod GPU18by Zhengxu YuLitJev†Qwen3.8-27B62.782.483.566.733.6$0.0077 est.40smalljev semantic-v9†27.4I 35 · C 59 · S 80 · K 58 · $0.131 est.100.0%95.8%95.2%71.4%0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server16by QwenQwen3-Reranker-4B†63.864.067.078.749.2$0.050100.0%79.2%87.7%50.0%0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 sour RunPod GPU17by kshetrajna12reflex-27b†Qwen3.8-27B63.385.886.267.532.3$0.025 est.41GLiNER2†24.0I 36 · C 24 · S 72 · K 83 · $0.104 est.100.0%96.9%93.2%75.0%1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 sauthor's demo server13by ZeroEntropyZeroEntropy zerank-2†66.063.076.579.049.8$0.047100.0%79.2%88.4%47.3%0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 sour RunPod GPU14by OpenAIGPT-5.6 Lunalow reasoning effort65.995.389.877.528.5$0.242100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API15by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang65.383.477.477.136.5$0.0037 est.42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · $0.066 est.100.0%95.8%91.1%65.5%0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU12by Featherless AISimpleJev Qwen3.8-27B†66.384.781.171.239.5$0.0073 est.43GLiNER2.5 multi†16.6I 28 · C 56 · S 68 · K 82 · $0.015 est.100.0%93.8%87.7%49.1%0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server11by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1666.479.264.883.245.5$0.0039 est.44GLiNER2.5 small†13.8I 26 · C 47 · S 78 · K 82 · $0.022 est.100.0%84.4%74.7%56.8%0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU10by mithalounisystem-one-openGemma 4 E2B LoRA on an L466.669.556.777.064.8$0.0039 est.45Mixedbread mxbai-rerank-base-v2†0.8I 7 · C 83 · S 88 · K 68 · $0.01246BAAI bge-reranker-v2-m3†0.7I 6 · C 84 · S 90 · K 73 · $0.007747Alibaba GTE Reranker ModernBERT-base†0.3I 5 · C 77 · S 91 · K 70 · $0.01048Certo v1†0.0I 0 · C 82 · S 94 · K 100 · $0.067 est.100.0%96.9%91.1%65.5%0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 sour RunPod GPU9by IkerMoelopen-alternative-jev†Qwen3.5-4B, IkerMoel67.064.063.283.559.6$0.0010 est.classifier.dev (fast tier)† (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · $0.056 est.100.0%95.8%92.5%64.5%0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 sour RunPod GPU7by milliseconds.ai (Baptiste Laget)decision-machine-1†milliseconds.ai68.362.170.492.953.7$0.035100.0%76.0%89.7%46.8%0.17 s rawp95 0.30 s rawproduction API8by Mapikadecider-35b-a3b†67.679.671.580.845.3$0.0033 est.Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · $0.022 est.100.0%94.8%97.3%63.2%1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 sour RunPod GPU6by hjmurmur (Octalab)jqv†Qwen3-32B zero-shot68.679.379.074.647.5$2.669 est.Needle 3, options as tools (partial run)1.1I 14 · C – · S 53 · K 65 · $0.037 est.100.0%96.9%91.1%70.9%0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 sour RunPod GPU5by kshetrajna12reflex 4B†70.380.175.268.059.7$0.014 est.Needle 3 (partial run)0.1I 5 · C – · S 60 · K 59 · $0.022 est.100.0%97.9%95.2%59.5%0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU3by Maisa (David Villalón)djev†Maisa, diffusion-gemma73.082.765.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API4by Eldan RingWinnow-12B Q8†71.282.072.082.352.9$0.024 est.Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)Jev (TypeSafe, closed)Jev rebuild (open, or open source planned)Instruction model, JSON schemaSmall tool-calling modelService built on JevZero-shot classifier (not a Jev rebuild)Closed decision model (API only, not Jev)Shown, not ranked — honorable mention (runs another entrant's model) · partial run⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.Legend and notes est. = no measured bill; priced like a large inference provider (how costs are estimated).ann. = the provider’s announced price, not yet charged.Names link to each project.A label-only system has no calibration (–, counted as 0).djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone
server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone
server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.Laya: The English checkpoint (repo root), run on our CPU through its own /v1/systemone
package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible laya
server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.Weighting: Intelligence : Calibration : Speed : CostOfficial defaultCustomIntelligence25 %Calibration25 %Speed25 %Cost25 %The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score. Explore by task difficultyAll tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.What the run says (JevBench Score)Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.1 — 1.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.Axes, tiers, latency and costSort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.💲 $ per 1,000 decisions, not per 1,000 tokens — one decision ≈ 950 input tokens.Rank#SystemEndpoint1by TypeSafe AIJev 1.13.074.485.782.783.352.0$0.040100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ73.179.072.683.759.5/v1/systemone
/v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.† SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.† djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.† jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.† decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.† Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.† OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.† kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.† kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.† Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.† jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.† Laya: The English checkpoint (repo root), run on our CPU through its own laya package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.† Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.† OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.† openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.† openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.† kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.† GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.† smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.† GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).† GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.† GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.† Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.† BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.† Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.† Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.† classifier.dev: Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.Honorable mentions — services built on another entrant's modelA service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.classifier.devno rank83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4Runs on Jev (TypeSafe).Intelligence85.1Calibration77.9Speed87.6Cost84.3$ per 1,000 decisions~$0.0033 est.Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.Why it is not ranked, what its price assumes, and what we foundclassifier.dev is not its own model. Its own pages say so: "The fast tier is Jev, TypeSafe's decision model" (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with "model": "jev-1.13.0" — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.Only the fast tier was measured. The smart tier's escalation was never run, so nothing here scores it.Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: "The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key" (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench's whole questions.Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev's 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev's own explanation for differences of this kind is batching ("The fast tier is Jev, packed a thousand to a request"); on their own two test sets they measured the same difference as noise.A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/aboutWhich public tasks did each system get right?This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped.Show 231 public task outcomes across 52 systemsTaskJev 1.13.0SemIfdjevWinnow-12B Q8reflex 4Bjqvdecision-machine-1decider-35b-a3bopen-alternative-jevsystem-one-openOpenJevSimpleJev Qwen3.8-27BZeroEntropy zerank-2GPT-5.6 Lunaopenjev-sglangQwen3-Reranker-4Breflex-27bLitJevkev 0.6BSimpleJev Qwen3.6-35B-A3Bdjevjev-localdecider-2bBespoke Nimble 9BGemini 3.1 Flash-LiteOpenJevkev 4BDeepSeek V4.1 Flashkev 8BOpen-Jev 9Bsystem-onejeffLayaOpen-Jev 2BOpenDecisionopenJev Verdict 1.4openJev Verdictkev 0.5BGLiNER2 largesmalljev semantic-v9GLiNER2open-jev-deberta-v3-largeGLiNER2.5 multiGLiNER2.5 smallMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-baseCerto v1classifier.devQwen3.8 27BNeedle 3, options as toolsNeedle 3Easy · 48 of 72 decisions publicEasy · 48 of 72 public48/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4842/4842/4841/4846/4848/4847/4847/4848/4844/4841/4822/4823/4817/4812/4848/4848/4832/4823/48easy-intent-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓✓✓×easy-intent-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓✓easy-intent-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓××easy-intent-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓×easy-intent-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓××easy-intent-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓×easy-intent-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓×✓✓✓×easy-intent-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓easy-intent-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓✓✓×easy-intent-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓××✓✓✓✓×easy-intent-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓×easy-intent-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓easy-fact-00noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓easy-fact-01noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××✓✓✓×✓××✓×✓×✓✓✓✓easy-fact-02noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓easy-fact-03noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓××✓×✓×✓✓✓×easy-fact-04noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×✓easy-fact-05noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓✓✓✓✓×✓×✓✓××easy-fact-06noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓easy-fact-07noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓✓✓✓×✓×✓✓××easy-fact-08noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓easy-fact-09noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓×✓✓×✓×✓✓××easy-fact-10noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓××easy-fact-11noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓✓✓✓×✓×✓✓××easy-extraction-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓×✓easy-extraction-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×✓easy-extraction-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓×✓easy-extraction-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓easy-extraction-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓×✓easy-extraction-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓easy-extraction-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×××✓✓✓✓easy-extraction-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓easy-extraction-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓easy-extraction-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓easy-extraction-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓×✓easy-extraction-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓easy-tool_selection-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓easy-tool_selection-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓×easy-tool_selection-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓×easy-tool_selection-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×easy-tool_selection-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×easy-tool_selection-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×easy-tool_selection-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓×easy-tool_selection-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓easy-tool_selection-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓××easy-tool_selection-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×easy-tool_selection-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×easy-tool_selection-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓×Medium (standard) · 72 of 96 decisions publicMedium (standard) · 72 of 96 public71/7271/7271/7269/7268/7269/7254/7270/7260/7267/7270/7270/7257/7270/7268/7254/7269/7271/7258/7267/7271/7260/7261/7267/7271/7272/7264/7271/7267/7265/7264/7254/7250/7255/7243/7250/7245/7235/7242/7249/7246/7231/7232/7230/7224/7226/7226/7224/7271/7271/7219/7212/72original-policy-01-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓×✓✓✓✓××original-policy-01-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓××✓✓✓✓✓××✓×✓×✓✓××original-policy-02-0noul✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓×✓✓✓✓××✓✓×✓original-policy-02-1noul✓×✓✓×✓×✓×✓✓✓×✓✓××✓×✓✓××✓✓✓✓✓✓✓×✓✓✓×✓✓××✓××✓✓✓✓××✓✓××original-policy-03-0noul✓✓✓✓✓✓×✓✓×✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓×××✓✓××✓×××✓×✓××✓✓××original-policy-03-1noul✓✓✓✓✓××✓×✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓××××××××✓××✓✓××original-policy-04-0noul✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓✓✓✓××✓×✓×✓✓××original-policy-04-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓××✓✓✓×✓××✓×✓✓✓✓××original-policy-05-0noul✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓✓✓✓××✓×✓✓✓✓××original-policy-05-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓×✓✓✓✓××original-policy-06-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓××✓✓×✓××✓✓××original-policy-06-1noul×✓✓×✓✓✓✓✓✓✓×✓××✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓×✓✓×✓×××✓××original-intent-01-0choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓×✓××✓×✓×✓✓✓✓×original-intent-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓✓×××✓✓✓×✓×✓××✓✓×✓original-intent-02-0choice✓✓✓✓✓✓×✓×✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓×✓×✓××✓×××××××××××××✓✓✓×original-intent-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×××✓✓✓✓×✓××××✓✓✓×original-intent-03-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓××××✓✓×××××✓✓××original-intent-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×original-intent-04-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×original-intent-04-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓××✓✓×✓✓×✓××××✓✓✓×original-intent-05-0choice✓✓✓✓✓✓×✓×✓✓✓×✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓✓✓××××✓✓✓××✓×✓×✓×✓×✓✓××original-intent-05-1choice✓✓✓✓✓✓×✓×✓✓✓×✓✓×✓✓✓✓✓×✓×✓✓✓✓✓✓✓××××✓✓✓✓×✓×✓×✓×✓×✓✓××original-intent-06-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓×✓×✓××✓××××✓✓✓×original-intent-06-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓××✓×✓×✓✓××××✓✓✓×original-ordinal-01-0score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓××✓✓××original-ordinal-01-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓×✓✓✓××original-ordinal-02-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××××✓✓××original-ordinal-02-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××××✓✓××original-ordinal-03-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓×✓×××✓✓✓××original-ordinal-03-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓×✓✓×✓×××✓✓✓××original-ordinal-04-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓×✓×✓✓××original-ordinal-04-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓✓×✓×✓✓××original-ordinal-05-0score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓!✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓××××✓×✓✓✓××original-ordinal-05-1score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓×✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓×✓××✓×✓✓✓××original-ordinal-06-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓××original-ordinal-06-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓××original-extraction-01-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓×✓✓✓××××××✓×××××××✓✓✓××original-extraction-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×××××✓×✓×××××✓✓✓×✓original-extraction-02-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×✓✓✓×✓✓××original-extraction-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××original-extraction-03-0choice✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓×✓×✓××××××✓✓×✓original-extraction-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×✓××××✓✓××original-extraction-04-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓××✓✓✓×××××✓✓××original-extraction-04-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓×✓×✓××××✓✓✓✓original-extraction-05-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××original-extraction-05-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××original-extraction-06-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓××××✓✓✓✓original-extraction-06-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓××××××✓✓×✓original-adequacy-01-0noul✓✓✓×✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓×✓✓×××✓✓✓×✓✓✓✓✓original-adequacy-01-1noul✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××✓×✓✓×✓✓✓✓✓original-adequacy-02-0noul✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓×✓✓✓✓✓✓✓×✓×✓✓××✓×××✓×✓✓✓××××✓✓✓✓××original-adequacy-02-1noul✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×××✓✓✓✓✓✓original-adequacy-03-0noul✓✓✓✓×××✓×✓✓✓×✓✓×✓✓××✓××××✓×✓×✓××××××××××✓×××××✓×✓✓××original-adequacy-03-1noul✓✓×✓✓✓×✓×✓×✓×✓×××✓××✓×××✓✓×✓×✓××✓×××××××✓×××××✓×✓✓××original-adequacy-04-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓××✓××××✓✓×✓✓✓××original-adequacy-04-1noul✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓××✓××××✓✓×✓✓✓××original-adequacy-05-0noul✓✓✓✓✓✓✓××✓✓✓×✓××✓✓××✓×××✓✓×✓××××✓××××✓×✓✓✓××××✓×✓✓××original-adequacy-05-1noul✓✓✓✓×××✓×✓✓✓×✓✓✓✓✓×✓✓××✓✓✓×✓✓×××✓✓×××××✓✓×××××✓×✓✓××original-adequacy-06-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓×××✓×✓✓×✓✓✓✓✓original-adequacy-06-1noul✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓×××✓✓×✓✓××✓✓✓✓×✓✓✓✓✓original-routing-01-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓××✓✓✓×original-routing-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓××××××××✓✓××original-routing-02-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓××××✓×✓✓××original-routing-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓××××✓×✓✓××original-routing-03-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓××original-routing-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓××original-routing-04-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓✓××✓✓✓×✓✓✓×××××××✓✓✓×original-routing-04-1choice✓✓✓✓×✓×✓×✓×✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓✓✓××✓×××✓✓×✓×××××××××✓✓✓×original-routing-05-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓×✓×✓✓××✓××××××××✓✓××original-routing-05-1choice✓✓✓✓✓✓×✓××✓✓✓✓✓×✓✓✓✓✓××✓✓✓✓✓✓×✓✓××✓×××××××××××××✓✓××original-routing-06-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓×××××××✓✓!✓×original-routing-06-1choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓×××××××✓✓✓××Hard · 111 of 220 decisions publicHard · 111 of 220 public81/11168/11175/11181/11167/11168/11154/11174/11163/11154/11171/11182/11157/111107/11181/11155/11184/11180/11148/11173/11185/11165/11155/11169/11182/11185/11141/111107/11150/11166/11154/11143/11139/11146/11138/11141/11142/11133/11141/11144/11141/11142/11137/11135/11140/11142/11135/11137/11178/11147/510/017/44hard-opus-a-long_policy-01choice✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓×✓××✓✓××✓××××××!××××××✓✓·×hard-opus-a-long_policy-04choice×××××××××××××✓××××××!××××××✓×××××××××××××××××××××✓·✓hard-opus-a-long_policy-08choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓✓✓×✓✓✓×✓×✓✓××××××××✓××××××××✓✓·×hard-opus-a-long_policy-09choice✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓×✓✓××✓✓✓×✓×✓×××××××××××××✓×✓××✓✓·×hard-opus-a-long_policy-11noul✓✓×✓✓✓✓✓×✓×✓✓✓✓✓✓✓×✓✓✓×✓✓×✓✓✓×✓×✓××✓✓××××✓×××✓×✓✓✓·×hard-opus-a-long_policy-13noul××××✓××✓×××××✓×✓×✓✓×✓××✓×××✓×××✓✓××✓✓✓××✓✓✓✓×✓×✓×!·✓hard-opus-a-long_policy-17choice✓×✓✓××××××✓✓✓✓×✓✓✓✓✓!✓×✓✓✓×✓✓✓××××××××××××✓×××××✓✓·×hard-opus-a-long_policy-19noul×✓×✓××✓×✓××××✓✓××✓××✓×××✓✓×✓×××××✓✓✓✓××××✓××✓✓✓××✓·×hard-opus-a-probability-03noul✓×××××✓✓×✓×✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓✓✓✓✓✓××××✓✓·✓hard-opus-a-probability-04choice×××××××××××××✓××××××✓✓××✓××✓××××××××××××✓×××✓××××✓·×hard-opus-a-probability-07choice××✓×××××××✓✓✓✓×✓✓✓×✓✓×××✓✓×✓✓✓××✓✓✓××××××××××××✓×✓·×hard-opus-a-probability-08noul✓✓✓×✓✓××✓×✓✓✓✓✓×✓×✓✓✓×✓×✓✓×✓✓✓✓✓×✓✓×××××××✓×✓×✓×✓✓·×hard-opus-a-temporal_numeric-03noul×✓×✓××××××××✓××✓×✓××✓××××✓×✓××××✓××✓✓××××✓×✓✓××××✓·✓hard-opus-a-temporal_numeric-06noul✓×✓✓✓✓×✓✓×✓✓×✓✓×✓✓×✓✓✓××✓✓×✓×✓×××××✓✓××✓✓×✓×✓✓✓×✓✓·×hard-opus-a-temporal_numeric-07choice××××✓×××✓×✓××✓✓×××✓×✓×✓××××!×××✓✓×✓××××✓×××××××✓×✓·✓hard-opus-a-temporal_numeric-09noul×✓×✓✓✓×✓✓××✓×✓✓×✓✓××✓××✓✓✓×✓××××××✓×✓××✓✓×✓×✓✓✓××✓·×hard-opus-a-temporal_numeric-12score××✓××××××××××✓×××✓✓×!××××××✓××××××××××××✓××✓××✓××✓·✓hard-opus-b-ambiguous-02choice✓✓✓✓×××✓✓×✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×✓✓×××✓×××××××✓✓✓✓✓·×hard-opus-b-ambiguous-03choice×××✓×✓××✓××✓×✓××✓✓××!××××✓×✓✓×✓××××✓✓✓✓✓✓✓✓✓×✓✓××✓·✓hard-opus-b-ambiguous-07choice✓××××✓×✓×××✓×✓✓×✓×××!××✓✓✓×✓×××××××✓✓××××✓×××✓✓×✓✓·✓hard-opus-b-ambiguous-09choice✓✓✓✓✓××✓✓×✓✓×✓✓✓✓××✓✓✓××✓✓×✓×✓××××××××××✓××✓××××✓✓·✓hard-opus-b-ambiguous-10choice✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓××✓✓×✓××××××××××✓×✓✓✓×××××✓✓·×hard-opus-b-ambiguous-11choice✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓×××✓××✓××✓✓××××××✓✓·×hard-opus-b-ambiguous-13choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓!✓✓✓✓✓✓✓×✓×××✓××××××××✓×××××✓✓·×hard-opus-b-multi_hop-03choice××✓✓✓×××××✓✓×✓✓×✓✓×✓!✓×✓×××✓××××✓××××××✓××✓××××✓×✓·×hard-opus-b-multi_hop-04choice××✓×××××××✓✓×✓××××××✓××✓×××✓×××××××××××××××××××××✓·×hard-opus-b-multi_hop-05score×××××××✓×××××✓×✓×××✓✓×✓××××✓××××✓××××××××××××××××✓·×hard-opus-b-multi_hop-07choice✓×✓×✓××✓××✓✓×✓✓××✓×✓!××✓×✓×✓✓✓×××××××××××✓××××××✓✓·×hard-opus-b-multi_hop-08noul✓✓✓✓✓✓×✓✓×✓✓×✓✓×✓✓×✓✓✓×✓✓✓×✓×✓×✓××✓×✓✓✓×××✓×✓×✓✓✓✓·×hard-opus-b-probability-01choice✓✓✓✓✓×✓✓××✓✓×✓××✓✓✓×✓×✓✓✓×✓✓×✓××✓✓✓×✓✓×✓×××✓××××✓✓·×hard-opus-b-probability-02choice×××✓×××✓××××××✓×××××✓×××✓✓×✓×××××××××××××××××××××✓·×hard-opus-b-probability-03noul✓×✓✓×✓✓✓×✓×✓✓✓××✓✓✓✓✓✓×✓✓××✓×✓✓✓✓×✓✓✓✓✓✓×✓✓✓×✓×✓✓✓·✓hard-opus-b-probability-04choice✓✓✓×✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓××✓✓✓✓××××✓×××✓✓×✓✓·✓hard-opus-b-probability-06choice✓××✓✓××✓×××✓×✓×✓✓✓××✓××✓✓✓×!××××✓××✓✓×✓×✓×✓✓×✓××✓✓·✓hard-opus-b-tradeoff-01choice×✓×××✓××××××✓✓××××××✓××××✓×✓×××✓✓××✓××××××××✓××××✓·×hard-opus-b-tradeoff-03choice✓✓✓✓✓××✓✓×✓✓×✓✓×✓✓×✓✓✓✓✓✓✓×✓××××✓✓✓×✓×✓×✓××✓××××✓✓·×hard-opus-b-tradeoff-06noul✓××✓✓✓✓✓×××✓×✓✓✓✓✓✓✓✓✓×××✓×✓✓✓×✓✓××✓×✓××✓××✓✓××✓✓✓·✓hard-opus-b-tradeoff-07choice✓××✓✓××××××✓×✓××✓✓✓×✓✓×✓✓✓×✓××××✓××××✓×✓✓✓×✓✓✓✓×✓✓·×hard-opus-b-tradeoff-08noul✓✓✓×✓××✓××××✓✓×✓××✓×✓××××✓×✓××✓×✓××××✓×××✓×✓✓×××✓✓·✓hard-opus-b-tradeoff-12choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓✓✓✓!✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓××✓×✓✓×✓×××✓✓·×hard-opus-c-long_policy-02choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×✓!✓✓✓✓✓×✓✓✓✓×××✓✓✓×✓×✓×××✓×××✓✓·×hard-opus-c-long_policy-03choice×××✓×××✓×××✓✓✓×✓××✓×✓✓××××✓✓✓✓✓×××××✓✓✓✓××✓×✓××××✓·✓hard-opus-c-long_policy-04noul✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓××✓✓✓✓✓××✓✓✓××××✓✓✓·✓hard-opus-c-long_policy-05score××✓✓××××××✓××✓××××××✓××××××✓×××××××××××××××××✓×✓×!·✓hard-opus-c-long_policy-08score×××××××××××××✓××××✓✓✓××××××✓×××××××××✓×××××××××✓×✓··hard-opus-c-long_policy-10choice✓×✓✓✓✓✓✓××✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓✓✓××✓✓✓✓××✓✓✓✓✓×✓✓×✓✓✓··hard-opus-c-long_policy-11noul××✓✓××××××✓××✓××✓×××××××✓××✓×××××✓✓××××××××××✓✓✓×!··hard-opus-c-probability-03choice✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓×✓✓✓✓✓✓✓✓×××××××✓×××××✓×✓✓✓··hard-opus-c-temporal_numeric-02noul××××××✓✓×✓××✓✓×✓××✓×✓××××××✓××✓✓✓××✓✓✓×××✓✓✓✓××✓×✓··hard-opus-c-temporal_numeric-03choice×××××××××✓×××✓✓✓×××✓!××✓×××✓××✓×××✓×××××××✓×××××××··hard-opus-c-temporal_numeric-04choice×××××××××××××✓××××××!××××××✓×××✓×××✓✓×✓✓×××××××××✓··hard-opus-c-temporal_numeric-06choice×××××××××××××✓××××××✓××××××✓××✓×××✓×××✓××××××××××···hard-opus-c-temporal_numeric-08choice✓×✓✓✓✓×✓××××✓✓✓×✓✓✓×✓××××××✓××××✓×✓××××××××××✓×✓×···hard-opus-c-temporal_numeric-12choice✓✓×✓✓✓✓××✓××✓✓✓✓✓××✓✓×✓✓×××✓×✓✓××××✓✓×✓✓××××✓××××···hard-sol-a-adversarial-01choice✓✓✓✓×✓✓✓✓×✓✓×✓✓×✓✓×✓!✓×✓✓✓×✓×✓×✓×✓✓✓✓×✓✓✓×✓×××××✓···hard-sol-a-adversarial-06noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓···hard-sol-a-adversarial-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓××✓×✓××✓✓××✓×✓···hard-sol-a-adversarial-08noul✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓×✓✓✓✓×✓✓✓×✓×××✓✓···hard-sol-a-adversarial-09score✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓×✓✓✓✓✓×××✓···hard-sol-a-adversarial-11choice✓×✓✓✓✓×✓××✓✓✓✓✓×✓✓✓✓✓✓××✓✓✓✓×✓×✓✓××✓✓✓✓××✓××✓×✓×✓···hard-sol-a-multi_hop-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓×✓✓✓×✓××✓✓××✓···hard-sol-a-multi_hop-02choice✓✓✓✓✓×××✓×✓××✓✓××✓×✓!✓×✓✓✓×✓×✓✓✓×✓××××××××✓✓××××✓···hard-sol-a-multi_hop-05choice✓✓✓×✓✓××✓×✓✓×✓✓✓✓✓×✓✓✓✓✓✓✓×✓×✓✓×××××××××××××✓✓✓×✓···hard-sol-a-multi_hop-07choice✓✓✓✓×✓×✓×✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓××××✓××××××××✓×××✓××✓···hard-sol-a-multi_hop-08choice✓✓×✓××××××✓××××✓××××✓×✓××✓✓✓✓××✓×✓×××××✓×××××××××···hard-sol-a-multi_hop-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓×××✓××××××✓✓✓✓×✓××✓···hard-sol-a-multi_hop-10choice✓×✓××××✓××✓×✓✓✓✓✓×✓×!××✓✓✓✓✓×××✓××××××✓××✓××✓×✓×✓···hard-sol-a-multi_hop-12choice✓××✓×××✓×✓×✓✓✓✓×✓✓××✓××✓✓✓×✓×✓✓✓×✓××✓×××✓✓××✓×××✓···hard-sol-a-trap-02noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××××✓××××✓✓✓✓✓···hard-sol-a-trap-04noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓××✓✓✓✓✓···hard-sol-a-trap-05score✓✓×✓✓✓×✓✓✓×✓×✓✓×✓✓✓✓✓×✓×✓✓✓✓✓✓✓××✓×××✓✓××××××××✓✓···hard-sol-a-trap-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓✓✓✓××✓×××✓✓✓×✓×××××✓✓···hard-sol-a-trap-08choice✓✓✓✓✓✓×✓✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓×××✓×××××✓×××××××✓✓···hard-sol-a-trap-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓✓✓✓×××××××✓···hard-sol-a-trap-13noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×××✓××✓××✓✓✓×✓···hard-sol-a-trap-15noul✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓×✓×××××✓✓✓×✓···hard-sol-b-judge_hard-01noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓××✓✓···hard-sol-b-judge_hard-02noul××××××××××××××××××××!××××××××✓×××××××××✓×××××✓✓✓×···hard-sol-b-judge_hard-05noul✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓×✓✓✓✓✓×××✓···hard-sol-b-judge_hard-08noul×××××××××✓×✓×✓✓×××××!××××✓×✓××××✓××××××✓×××××✓✓✓×···hard-sol-b-judge_hard-10noul×✓×××✓××✓✓×✓×✓✓×✓✓××✓✓××✓✓×✓×××××✓×××✓×✓×××××✓✓××···hard-sol-b-judge_hard-14noul✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓✓×✓✓✓✓××✓×××××✓×××××✓✓×✓···hard-sol-b-judge_hard-15noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓××✓···hard-sol-b-judge_hard-18noul×✓✓✓✓✓×✓✓✓✓✓×✓××✓✓××✓✓×✓✓✓×✓×✓×××××××××✓×××××✓✓✓✓···hard-sol-b-long_policy-01choice✓×✓✓×✓✓×✓✓✓✓✓✓✓×✓✓×✓!✓✓✓✓✓×✓✓×✓×××××××✓×✓×✓✓✓✓×✓✓···hard-sol-b-long_policy-02choice✓✓✓✓✓×✓×××✓✓×✓✓×✓✓×✓✓×✓✓✓✓×✓✓✓××××××××××××××✓×✓✓✓···hard-sol-b-long_policy-05choice✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓!✓×✓✓×✓✓✓×✓×✓×××✓✓✓✓××××××✓×✓···hard-sol-b-long_policy-06choice✓✓×××××✓×✓×✓×✓××✓✓×✓✓×✓×✓××!×✓×××✓×××××✓×✓✓×××××✓···hard-sol-b-routing_hard-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓✓×✓××××✓···hard-sol-b-routing_hard-02choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓×✓×✓×××✓×××✓···hard-sol-b-routing_hard-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓××✓××××××✓××✓××××✓···hard-sol-b-routing_hard-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓×✓×××××××✓···hard-sol-b-routing_hard-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓×××✓×××××××✓···hard-sol-b-temporal_numeric-01choice×××××××××××××✓××××××✓××××✓×✓×××××××××××××××××××××···hard-sol-b-temporal_numeric-02choice××××××✓×✓✓×✓✓✓✓✓✓×✓×✓×✓✓×✓✓✓××✓✓×✓××××××✓×✓××✓×××···hard-sol-b-temporal_numeric-03choice✓××××××××✓××✓✓××✓×✓×!✓✓××✓×✓×××✓×××✓××✓×✓××✓×××✓✓···hard-sol-b-temporal_numeric-04choice××✓×××✓×✓××××✓✓✓××××✓×✓××✓×✓×××✓✓×✓✓✓××××✓✓××✓✓✓×···hard-sol-c-judge_hard-03noul✓✓✓✓×✓××✓✓✓✓×✓✓×✓✓×✓✓×××✓✓×✓×××✓×××××××✓✓×✓××✓✓✓✓···hard-sol-c-judge_hard-05noul✓×✓✓✓✓✓✓×✓✓×✓✓✓✓××✓✓✓×✓✓✓✓✓✓✓✓✓××××✓✓✓✓✓×✓×✓✓×××✓···hard-sol-c-judge_hard-07noul✓✓✓✓✓✓×✓✓×✓××✓✓×✓✓×××✓××✓××✓×××✓✓××××××✓×××××✓✓✓✓···hard-sol-c-judge_hard-08noul✓✓✓✓×✓××✓✓✓✓×✓✓×✓✓×✓✓✓××✓✓×✓×✓×××××××✓××××××✓✓✓×✓···hard-sol-c-judge_hard-09noul✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓×××✓···hard-sol-c-judge_hard-10noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓···hard-sol-c-judge_hard-11noul✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓✓×✓✓✓✓××✓×××××✓✓××××✓✓×✓···hard-sol-c-judge_hard-13noul✓✓✓✓×✓××××✓✓×✓✓×✓✓××✓✓××✓✓×✓×✓×✓✓××××××✓✓×✓××✓✓✓×···hard-sol-c-judge_hard-15noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓×✓××××✓···hard-sol-c-multi_hop-07choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓×✓×✓✓✓✓×××✓✓✓××✓✓×××✓···hard-sol-c-multi_hop-09choice✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓!×✓✓✓✓✓✓×✓✓××✓××××✓××✓×✓×✓✓✓✓···hard-sol-c-multi_hop-11choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓!✓✓✓✓✓×✓✓✓✓××✓✓×××××××××✓×××✓···hard-sol-c-multi_hop-12choice✓✓✓✓×✓×✓×✓×✓×✓✓×✓✓×✓✓✓×✓✓✓×✓×✓×××××✓×××××✓×××××✓✓···hard-sol-c-multi_hop-13choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓×✓×✓✓××××✓···✓ correct · × wrong · ! failed (scored wrong) · · not attempted · — no public outcome in the pinned artifact. A group row counts the public tasks it lists; the tier columns of the table above use all decisions of the tier. Every task id carries its topic (hard-opus-a-long_policy-01 is a long_policy task), and the tag after the id is the published question type: choice — pick one of a defined set of options · noul — whether a stated condition holds · score — a degree along a described dimension. Task ids are shown without their tier prefix; the tier is the group row. Task descriptions are intentionally not included; the task id, tier, topic and type are the published public metadata.How the JevBench Score worksJevBench Score = (Intelligence × Calibration × Speed × Cost)1/4, each axis on 0–100 — the geometric mean. A weak axis pulls the score down hard: a strong axis cannot buy it back. Below 50 Intelligence the score is also multiplied by (Intelligence ÷ 50)², so a system barely better than guessing cannot rank on speed and price.Intelligence — accuracy above chance: per tier, how much of the gap between guessing and all-correct a system closes (0 = guessing, 100 = all correct), weighted: hard 30 %, easy 14 %, standard 28 %, judge 28 % (220 / 72 / 96 / 146 decisions).Calibration — on the hard tier: does “80 % sure” come true 80 % of the time, and does the returned distribution match the exact gold distribution on the probability items.Speed — median and 95th-percentile latency, one request at a time: 0.1 s scores 100, each 10× slower costs 20 points (1 s = 80, 10 s = 60). ⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.Cost — dollars per 1,000 decisions, never per 1,000 tokens: $0.001 scores 100, each 10× more expensive costs 30 points ($0.01 = 70, $0.10 = 40, $1 = 10). Models without a tariff are priced at hosted-provider prices, marked “est.” (how). 💲 $ per 1,000 decisions, not $ per 1,000 tokens. One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.Full scoring rulesJevBench Score. Geometric mean of Intelligence, Calibration, Speed and Cost, 25 % each. If chance-corrected Intelligence is below 50, multiply by (Intelligence / 50)^2; at or above 50 there is no penalty.Intelligence. Per tier: 100 x (accuracy - chance) / (1 - chance), clipped at 0. Chance is 1 / options for each item (1 / levels for score items), then averaged within the tier. Tier weights: hard 30 %, easy 14 %, standard 28 %, judge 28 %. Failed, timed-out or unparseable answers count as wrong.Calibration. Hard tier only, systems that return a probability distribution: mean of (a) 100 x (1 - ECE/0.5), ECE = top-label expected calibration error in 10 bins, and (b) probability fidelity = 100 x (1 - mean total-variation distance) between the returned distribution and the exact gold distribution on the 20 probability items. Label-only systems have none; it counts as 0 in the JevBench Score.Speed. Mean of score(p50) and score(p95) of the serial 242-decision standard+judge run; score(s) = 100 - 20 log10(s / 0.1 s), clipped to 0..100 (0.1 s = 100, 1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo. Production APIs (Jev, djev, classifier.dev, OpenAI, Google, DeepSeek, Chutes) are not adjusted.Cost. US dollars per 1,000 DECISIONS — not per 1,000 tokens. One decision is one whole question: its state, its rubric and its options, which is hundreds to thousands of input tokens. Pooled over all 534 v1.2 decisions; score = 100 - 30 log10(usd / 0.001), clipped to 0..100 ($0.001 = 100, $0.01 = 70, $0.10 = 40, $1 = 10). Measured = public tariff x measured tokens. est. = hosted-provider list price of the same weights or size class x tokens (for a flat-rate service, its published plan price at full use). announced = the provider's published price, not yet charged (free preview), x measured tokens.Ranked. Ranked: a system's own model, with every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank. A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number.Presets. Other views reweight the same four axes and combine them the same way (geometric mean). They are not the JevBench Score.How the ranking moves with other weightsRank and score under the JevBench Score and the earlier views, all combined as a geometric mean (Intelligence : Calibration : Speed : Cost). Highlighted = a different rank than the JevBench Score. Ranked systems only.SystemJevBench Score25:25:25:25 · officialBalanced, no calibration33:0:33:33 · not the defaultEmphasis on Accuracy60:0:20:20 · not the defaultEmphasis on Speed20:0:60:20 · not the defaultEmphasis on Cost20:0:20:60 · not the defaultJev 1.13.0#1 74.4#3 71.8 (rank differs from the JevBench Score)#2 77.1 (rank differs from the JevBench Score)#4 76.2 (rank differs from the JevBench Score)#9 63.1 (rank differs from the JevBench Score)SemIf#2 73.1#2 73.3#3 75.5 (rank differs from the JevBench Score)#2 77.3#4 67.4 (rank differs from the JevBench Score)djev#3 73.0#1 75.8 (rank differs from the JevBench Score)#1 78.5 (rank differs from the JevBench Score)#1 81.7 (rank differs from the JevBench Score)#3 67.9Winnow-12B Q8#4 71.2#4 71.0#4 75.2#5 75.3 (rank differs from the JevBench Score)#10 63.1 (rank differs from the JevBench Score)reflex 4B#5 70.3#6 68.8 (rank differs from the JevBench Score)#5 73.1#17 68.4 (rank differs from the JevBench Score)#5 65.0jqv#6 68.6#14 65.5 (rank differs from the JevBench Score)#9 70.7 (rank differs from the JevBench Score)#14 69.0 (rank differs from the JevBench Score)#14 57.6 (rank differs from the JevBench Score)decision-machine-1#7 68.3#9 67.6 (rank differs from the JevBench Score)#22 65.4 (rank differs from the JevBench Score)#3 76.8 (rank differs from the JevBench Score)#11 61.7 (rank differs from the JevBench Score)decider-35b-a3b#8 67.6#13 66.3 (rank differs from the JevBench Score)#8 71.3#10 71.7 (rank differs from the JevBench Score)#18 56.9 (rank differs from the JevBench Score)open-alternative-jev#9 67.0#7 68.3 (rank differs from the JevBench Score)#17 66.6 (rank differs from the JevBench Score)#6 74.0 (rank differs from the JevBench Score)#8 64.7 (rank differs from the JevBench Score)system-one-open#10 66.6#5 70.3 (rank differs from the JevBench Score)#11 70.0 (rank differs from the JevBench Score)#9 72.9 (rank differs from the JevBench Score)#2 68.0 (rank differs from the JevBench Score)OpenJev#11 66.4#11 66.9#7 71.6 (rank differs from the JevBench Score)#8 73.0 (rank differs from the JevBench Score)#15 57.3 (rank differs from the JevBench Score)SimpleJev Qwen3.8-27B#12 66.3#18 62.0 (rank differs from the JevBench Score)#10 70.2 (rank differs from the JevBench Score)#24 65.5 (rank differs from the JevBench Score)#22 51.8 (rank differs from the JevBench Score)ZeroEntropy zerank-2#13 66.0#15 62.8 (rank differs from the JevBench Score)#29 62.9 (rank differs from the JevBench Score)#15 68.8 (rank differs from the JevBench Score)#16 57.2 (rank differs from the JevBench Score)GPT-5.6 Luna#14 65.9#23 59.5 (rank differs from the JevBench Score)#6 71.8 (rank differs from the JevBench Score)#23 66.1 (rank differs from the JevBench Score)#30 44.3 (rank differs from the JevBench Score)openjev-sglang#15 65.3#19 61.7 (rank differs from the JevBench Score)#12 69.6 (rank differs from the JevBench Score)#18 67.4 (rank differs from the JevBench Score)#24 50.0 (rank differs from the JevBench Score)Qwen3-Reranker-4B#16 63.8#16 62.8#27 63.3 (rank differs from the JevBench Score)#16 68.7#17 56.9 (rank differs from the JevBench Score)reflex-27b#17 63.3#26 57.2 (rank differs from the JevBench Score)#16 67.2 (rank differs from the JevBench Score)#28 61.1 (rank differs from the JevBench Score)#27 45.5 (rank differs from the JevBench Score)LitJev#18 62.7#28 57.0 (rank differs from the JevBench Score)#19 66.0 (rank differs from the JevBench Score)#29 60.7 (rank differs from the JevBench Score)#26 46.1 (rank differs from the JevBench Score)kev 0.6B#19 62.5#12 66.8 (rank differs from the JevBench Score)#30 60.4 (rank differs from the JevBench Score)#13 70.2 (rank differs from the JevBench Score)#1 70.4 (rank differs from the JevBench Score)SimpleJev Qwen3.6-35B-A3B#20 62.5#21 61.0 (rank differs from the JevBench Score)#14 67.8 (rank differs from the JevBench Score)#21 66.3 (rank differs from the JevBench Score)#23 50.6 (rank differs from the JevBench Score)djev#21 62.4#30 54.6 (rank differs from the JevBench Score)#25 63.9 (rank differs from the JevBench Score)#27 62.1 (rank differs from the JevBench Score)#34 41.1 (rank differs from the JevBench Score)jev-local#22 61.8#22 59.6#26 63.9 (rank differs from the JevBench Score)#26 63.3 (rank differs from the JevBench Score)#21 52.5 (rank differs from the JevBench Score)decider-2b#23 61.7#8 67.7 (rank differs from the JevBench Score)#23 65.1#7 73.5 (rank differs from the JevBench Score)#7 64.9 (rank differs from the JevBench Score)Bespoke Nimble 9B#24 60.5#24 58.9#20 65.9 (rank differs from the JevBench Score)#22 66.2 (rank differs from the JevBench Score)#25 47.0 (rank differs from the JevBench Score)Gemini 3.1 Flash-Lite#25 60.1#25 57.6#15 67.5 (rank differs from the JevBench Score)#20 66.3 (rank differs from the JevBench Score)#32 42.8 (rank differs from the JevBench Score)OpenJev#26 60.0#27 57.1 (rank differs from the JevBench Score)#13 67.9 (rank differs from the JevBench Score)#25 64.0 (rank differs from the JevBench Score)#31 42.8 (rank differs from the JevBench Score)kev 4B#27 59.7#10 67.2 (rank differs from the JevBench Score)#18 66.2 (rank differs from the JevBench Score)#12 70.5 (rank differs from the JevBench Score)#6 65.0 (rank differs from the JevBench Score)DeepSeek V4.1 Flash#28 57.5#34 48.4 (rank differs from the JevBench Score)#28 63.2#33 56.6 (rank differs from the JevBench Score)#40 31.7 (rank differs from the JevBench Score)kev 8B#29 56.4#20 61.2 (rank differs from the JevBench Score)#24 64.3 (rank differs from the JevBench Score)#19 66.3 (rank differs from the JevBench Score)#19 53.6 (rank differs from the JevBench Score)Open-Jev 9B#30 55.0#32 52.4 (rank differs from the JevBench Score)#31 59.3 (rank differs from the JevBench Score)#30 59.5#35 40.9 (rank differs from the JevBench Score)system-one#31 54.8#17 62.6 (rank differs from the JevBench Score)#21 65.6 (rank differs from the JevBench Score)#11 70.6 (rank differs from the JevBench Score)#20 53.1 (rank differs from the JevBench Score)jeff#32 54.4#31 53.6 (rank differs from the JevBench Score)#33 48.2 (rank differs from the JevBench Score)#34 54.5 (rank differs from the JevBench Score)#13 58.7 (rank differs from the JevBench Score)Laya#33 54.4#29 55.0 (rank differs from the JevBench Score)#34 47.7 (rank differs from the JevBench Score)#32 56.8 (rank differs from the JevBench Score)#12 61.4 (rank differs from the JevBench Score)Open-Jev 2B#34 51.3#33 50.1 (rank differs from the JevBench Score)#32 54.2 (rank differs from the JevBench Score)#31 58.4 (rank differs from the JevBench Score)#37 39.8 (rank differs from the JevBench Score)OpenDecision#35 40.6#35 41.7#35 35.2#35 46.0#28 44.9 (rank differs from the JevBench Score)openJev Verdict 1.4#36 38.9#37 37.4 (rank differs from the JevBench Score)#38 30.7 (rank differs from the JevBench Score)#37 40.8 (rank differs from the JevBench Score)#33 41.6 (rank differs from the JevBench Score)openJev Verdict#37 38.1#36 40.1 (rank differs from the JevBench Score)#36 33.3 (rank differs from the JevBench Score)#36 43.3 (rank differs from the JevBench Score)#29 44.7 (rank differs from the JevBench Score)kev 0.5B#38 33.2#39 35.4 (rank differs from the JevBench Score)#39 29.4 (rank differs from the JevBench Score)#38 38.9#38 38.7GLiNER2 large#39 29.6#38 36.5 (rank differs from the JevBench Score)#37 31.8 (rank differs from the JevBench Score)#39 37.8#36 40.5 (rank differs from the JevBench Score)smalljev semantic-v9#40 27.4#41 26.9 (rank differs from the JevBench Score)#41 22.6 (rank differs from the JevBench Score)#41 31.4 (rank differs from the JevBench Score)#41 27.6 (rank differs from the JevBench Score)GLiNER2#41 24.0#40 30.3 (rank differs from the JevBench Score)#40 24.6 (rank differs from the JevBench Score)#40 32.6 (rank differs from the JevBench Score)#39 34.6 (rank differs from the JevBench Score)open-jev-deberta-v3-large#42 23.1#42 21.9#42 17.8#42 23.7#42 24.9GLiNER2.5 multi#43 16.6#43 16.4#43 12.6#43 18.0#43 19.5GLiNER2.5 small#44 13.8#44 14.4#44 10.6#44 16.5#44 16.9Mixedbread mxbai-rerank-base-v2#45 0.8#45 0.6#45 0.3#45 0.9#45 0.8BAAI bge-reranker-v2-m3#46 0.7#46 0.5#46 0.3#46 0.8#46 0.7Alibaba GTE Reranker ModernBERT-base#47 0.3#47 0.3#47 0.1#47 0.4#47 0.4Certo v1#48 0.0#48 0.0#48 0.0#48 0.0#48 0.0Compare two systemsPick any two. The first radar shows the four axes of the JevBench Score (0–100, the values in the table above); the second shows accuracy by subject topic, over all tiers. Further out is better on every spoke.System ASystem BA: Jev 1.13.0 — Jev · JevBench Score 74.4 (#1)B: SemIf — Jev rebuild · JevBench Score 73.1 (#2)The four score axesRadar: the four JevBench Score axes, two systemsJev 1.13.0 vs SemIf. Intelligence: 85.7 vs 79.0; Calibration: 82.7 vs 72.6; Speed: 83.3 vs 83.7; Cost: 52.0 vs 59.5.50100Intelligence85.7 · 79.0Calibration82.7 · 72.6Speed83.3 · 83.7Cost52.0 · 59.5Speed includes the latency adjustment for self-hosted and demo endpoints — an assumption, see Limits. A label-only system has no calibration (counted as 0).Values as a tableAxisA: Jev 1.13.0B: SemIfIntelligence85.779.0Calibration82.772.6Speed83.383.7Cost52.059.5JevBench Score74.473.1Accuracy by subject topic — not part of the scoreRadar: accuracy by subject topic, two systemsAccuracy by subject topic, Jev 1.13.0 vs SemIf. Math & numbers (129 items): 87.6% vs 79.1%; Coding & software (56 items): 83.9% vs 96.4%; Rules, policy & law (67 items): 83.6% vs 64.2%; Finance & commerce (64 items): 73.4% vs 60.9%; Support & operations (119 items): 89.1% vs 87.4%; Everyday language (79 items): 100.0% vs 100.0%; Safety & security (20 items): 100.0% vs 75.0%.50100Math87.6% · 79.1%Coding83.9% · 96.4%Rules & law83.6% · 64.2%Finance73.4% · 60.9%Support & ops89.1% · 87.4%Everydaylanguage100.0% · 100.0%Safety &security100.0% · 75.0%Share of each topic's decisions answered correctly, all tiers together — compare the two systems within a topic, not topics with each other.Values and notesTopics mix tiers differently — Everyday language is mostly easy items, Rules & law and Finance mostly hard ones — which is why topics are not compared with each other.Topics: one per item, drafted by a model and checked by hand — method. Held-out items count in the totals; their texts stay private.Topic (items)A: Jev 1.13.0B: SemIfMath & numbers (129)a calculation decides the answer: arithmetic, word problems, probability, dates, units87.6% 113 of 12979.1% 102 of 129Coding & software (56)code, SQL, repositories, developer tools and IT systems83.9% 47 of 5696.4% 54 of 56Rules, policy & law (67)applying written rules: company policies, contracts, regulations, eligibility83.6% 56 of 6764.2% 43 of 67Finance & commerce (64)money: payments, refunds, invoices, orders, expenses, insurance payouts73.4% 47 of 6460.9% 39 of 64Support & operations (119)support tickets, incidents, logistics, scheduling desks and routing work to a team89.1% 106 of 11987.4% 104 of 119Everyday language (79)short everyday messages: intents, assistant requests, reading a detail out of a text100.0% 79 of 79100.0% 79 of 79Safety & security (20)untrusted or injected instructions, fraud, moderation, access and security triage100.0% 20 of 2075.0% 15 of 20Jev alternatives, open source and self-hostingThe table above compares the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.What are open-source alternatives to Jev?The highest-ranked open entrants in this run are SemIf (#2, 73.1), djev (#3, 73.0), Winnow-12B Q8 (#4, 71.2), reflex 4B (#5, 70.3). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.How is JevBench scored?The official score is the geometric mean of chance-corrected Intelligence, Calibration, Speed and Cost, weighted 25% each. Version v1.3.0 uses 534 decisions (220 hard) and penalises systems below 50 Intelligence. Open Method and tiers for the exact rules, or inspect the MIT-licensed harness and public tasks.How do I submit my model?Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.Held-out hard-tier detailWith about 110 items on each side, ordinary noise is roughly ±9 percentage points. Read a system's public-minus-held-out gap against the field mean (-0.7 points across 49 complete systems): only an outlier against that field is meaningful. “Not public” does not mean “not seen”, because held-out items were sent to hosted APIs.SystemHard publicHard held-outPublic − held-out gap (95% interval)Field mean gapSemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)61.3% (68/111)57.8% (63/109)+3.5 points [-9.5, +16.4]-0.7 pointsOpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)64.0% (71/111)67.0% (73/109)-3.0 points [-15.6, +9.6]-0.7 pointssystem-one (Qwen3-8B, Sean Goedecke)48.6% (54/111)51.4% (56/109)-2.7 points [-15.9, +10.5]-0.7 pointsJev 1.13.0 (TypeSafe AI)73.0% (81/111)75.2% (82/109)-2.3 points [-13.8, +9.3]-0.7 pointssystem-one-open (Gemma 4 E2B LoRA on an L4)48.6% (54/111)49.5% (54/109)-0.9 points [-14.1, +12.3]-0.7 pointsBespoke Nimble 9B (Bespoke Labs)62.2% (69/111)68.8% (75/109)-6.6 points [-19.2, +5.9]-0.7 pointsopenjev-sglang (Qwen3.6-35B-A3B on SGLang)73.0% (81/111)69.7% (76/109)+3.2 points [-8.7, +15.2]-0.7 pointsGPT-5.6 Luna (low reasoning effort)96.4% (107/111)92.7% (101/109)+3.7 points [-2.3, +9.7]-0.7 pointsGemini 3.1 Flash-Lite73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 pointsopen-jev-deberta-v3-large (local CPU)37.8% (42/111)34.9% (38/109)+3.0 points [-9.7, +15.7]-0.7 pointsDeepSeek V4.1 Flash (thinking default)96.4% (107/111)93.6% (102/109)+2.8 points [-2.9, +8.6]-0.7 pointsNeedle 3 (Cactus, 2-bit, local CPU) (partial)38.6% (17/44)— (0/0)— points [—, —]-0.7 pointsQwen3.8 27B (Chutes TEE) (partial)92.2% (47/51)— (0/0)— points [—, —]-0.7 pointsNeedle 3, options as tools (post-hoc adapter mode) (partial)— (0/0)— (0/0)— points [—, —]-0.7 pointsopen-alternative-jev (Qwen3.5-4B, IkerMoel)56.8% (63/111)56.9% (62/109)-0.1 points [-13.2, +13.0]-0.7 pointsBAAI bge-reranker-v2-m337.8% (42/111)35.8% (39/109)+2.1 points [-10.7, +14.8]-0.7 pointsCerto v1 (AltSlate Labs)33.3% (37/111)30.3% (33/109)+3.1 points [-9.2, +15.4]-0.7 pointsclassifier.dev (fast tier)70.3% (78/111)70.6% (77/109)-0.4 points [-12.4, +11.7]-0.7 pointsdecider-2b (Mapika)49.5% (55/111)45.0% (49/109)+4.6 points [-8.6, +17.8]-0.7 pointsdecider-35b-a3b (Mapika)66.7% (74/111)64.2% (70/109)+2.4 points [-10.1, +15.0]-0.7 pointsdecision-machine-1 (milliseconds.ai)48.6% (54/111)45.0% (49/109)+3.7 points [-9.5, +16.9]-0.7 pointsdjev (thinking)76.6% (85/111)78.9% (86/109)-2.3 points [-13.3, +8.7]-0.7 pointsdjev (Maisa, diffusion-gemma)67.6% (75/111)71.6% (78/109)-4.0 points [-16.1, +8.2]-0.7 pointsGLiNER2 large (Fastino)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 pointsGLiNER2.5 multi (Fastino, 287M)33.3% (37/111)42.2% (46/109)-8.9 points [-21.6, +3.9]-0.7 pointsGLiNER2.5 small (Fastino, 74M)31.5% (35/111)34.9% (38/109)-3.3 points [-15.8, +9.1]-0.7 pointsGLiNER2 (Fastino, gliner2.5-base)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 pointsAlibaba GTE Reranker ModernBERT-base31.5% (35/111)35.8% (39/109)-4.2 points [-16.7, +8.2]-0.7 pointsjeff (Logan Markewich, GLiFormer 400M)38.7% (43/111)36.7% (40/109)+2.0 points [-10.8, +14.8]-0.7 pointsjev-local (Qwen3.5-9B)58.6% (65/111)59.6% (65/109)-1.1 points [-14.1, +11.9]-0.7 pointsjqv (Qwen3-32B zero-shot)61.3% (68/111)67.9% (74/109)-6.6 points [-19.2, +6.0]-0.7 pointskev 0.5B29.7% (33/111)32.1% (35/109)-2.4 points [-14.6, +9.8]-0.7 pointskev 0.6B (research preview)43.2% (48/111)36.7% (40/109)+6.5 points [-6.4, +19.5]-0.7 pointskev 4B (research preview)36.9% (41/111)47.7% (52/109)-10.8 points [-23.8, +2.2]-0.7 pointskev 8B (research preview)45.0% (50/111)49.5% (54/109)-4.5 points [-17.7, +8.7]-0.7 pointsLaya (Convai Innovations, ModernBERT-large 421M)35.1% (39/111)33.0% (36/109)+2.1 points [-10.4, +14.6]-0.7 pointsLitJev (Qwen3.8-27B)72.1% (80/111)74.3% (81/109)-2.2 points [-13.9, +9.5]-0.7 pointsMixedbread mxbai-rerank-base-v236.0% (40/111)44.0% (48/109)-8.0 points [-20.9, +4.9]-0.7 pointsOpen-Jev 2B (Zefan Cai)41.4% (46/111)44.0% (48/109)-2.6 points [-15.7, +10.5]-0.7 pointsOpen-Jev 9B (Zefan Cai)59.5% (66/111)62.4% (68/109)-2.9 points [-15.8, +10.0]-0.7 pointsOpenDecision (ModernBERT-large zero-shot)34.2% (38/111)32.1% (35/109)+2.1 points [-10.3, +14.6]-0.7 pointsOpenJev (thinking, BF16)76.6% (85/111)79.8% (87/109)-3.2 points [-14.1, +7.7]-0.7 pointsopenJev Verdict 1.436.9% (41/111)38.5% (42/109)-1.6 points [-14.4, +11.2]-0.7 pointsopenJev Verdict (heman10x, ModernBERT-base 151M)37.8% (42/111)38.5% (42/109)-0.7 points [-13.5, +12.1]-0.7 pointsQwen3-Reranker-4B49.5% (55/111)50.5% (55/109)-0.9 points [-14.1, +12.3]-0.7 pointsreflex-27b (Qwen3.8-27B)75.7% (84/111)76.1% (83/109)-0.5 points [-11.8, +10.8]-0.7 pointsreflex 4B (kshetrajna12)60.4% (67/111)66.1% (72/109)-5.7 points [-18.4, +7.0]-0.7 pointsSimpleJev Qwen3.6-35B-A3B65.8% (73/111)67.0% (73/109)-1.2 points [-13.7, +11.3]-0.7 pointsSimpleJev Qwen3.8-27B73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 pointssmalljev semantic-v939.6% (44/111)36.7% (40/109)+2.9 points [-9.9, +15.8]-0.7 pointsWinnow-12B Q873.0% (81/111)68.8% (75/109)+4.2 points [-7.8, +16.2]-0.7 pointsZeroEntropy zerank-251.4% (57/111)43.1% (47/109)+8.2 points [-4.9, +21.4]-0.7 pointsAccuracy is correct / attempted; invalid responses count as incorrect. The interval is the unpooled two-sample normal 95% interval for a difference in proportions. Partial systems are shown but excluded from the field mean.Public-split policy. Training on JevBench's public split is allowed and should be declared with each submission. Rankings continue to use all benchmark items. We report held-out results separately so that specialisation on public tasks is visible. Held-out means not publicly released, not guaranteed unseen: hosted systems receive these tasks during evaluation. We periodically issue fresh tasks to reduce the value of prior exposure.Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.How costs are estimatedOne decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.SemIf — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decisionWinnow-12B Q8 — ~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))reflex 4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))jqv — ~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))decider-35b-a3b — ~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))open-alternative-jev — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decisionsystem-one-open — ~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decisionOpenJev — ~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decisionSimpleJev Qwen3.8-27B — ~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))openjev-sglang — ~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decisionreflex-27b — ~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))LitJev — ~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))kev 0.6B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))SimpleJev Qwen3.6-35B-A3B — ~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))djev — ~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures includedjev-local — ~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)decider-2b — ~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))Bespoke Nimble 9B — ~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))OpenJev — ~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decisionkev 4B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))kev 8B — ~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))Open-Jev 9B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))system-one — ~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decisionjeff — ~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))Laya — ~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))Open-Jev 2B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))OpenDecision — ~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))openJev Verdict 1.4 — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)openJev Verdict — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]kev 0.5B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))GLiNER2 large — ~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)smalljev semantic-v9 — ~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))GLiNER2 — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]open-jev-deberta-v3-large — ~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decisionGLiNER2.5 multi — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)GLiNER2.5 small — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)Certo v1 — ~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))classifier.dev — ~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.Qwen3.8 27B — ~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decisionNeedle 3, options as tools — ~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]Needle 3 — ~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decisionReference prices by size class ($ per million input / output tokens)dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted onceusd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.No tariff, measurement, item, answer or rank changed. The prices before and after:Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)SemIf — $0.0230 → $0.0224 (-2.31 %)system-one-open — $0.0157 → $0.0149 (-5.13 %)OpenJev — $0.0672 → $0.0656 (-2.36 %)GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)openjev-sglang — $0.1346 → $0.1313 (-2.48 %)Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)system-one — $0.0915 → $0.0894 (-2.29 %)openJev Verdict — $0.0039 → $0.0037 (-5.21 %)GLiNER2 — $0.0039 → $0.0037 (-5.21 %)open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)Needle 3 — $0.0249 → $0.0238 (-4.36 %)Who could not be measured, and whyAn exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.open-jev (Dasein Labs) — MLX on Apple Silicon only. Its own README says Linux containers cannot reach the Apple GPU, so a RunPod NVIDIA GPU cannot run it.open-jev (JoshuaSP) — A DiffusionGemma 26B-A4B serving wrapper rather than new trained weights. It was demonstrated on an H100; no public endpoint exists and no suitable 80 GB RunPod host was available in this round.mini-jev (Mikhail Rakutko (r-ms)) — Public weights exist and fit a normal GPU, but the implementation covers Choice/Noul and explicitly does not measure Score. A faithful full-suite adapter would require new interface work rather than a mechanical endpoint adapter.system-one-gemma (Akash Kamat) — The adapter is public, but its Gemma base is gated behind Google’s licence terms. We do not accept binding terms on Florian’s behalf.jevlike (Vincent Wang-Maścianica) — Only Doom and chess vision checkpoints are released; there is no general text-decision checkpoint for this suite.AlexWortega/openjev (Alex Wortega) — Its released NLI and task-specific heads do not define a distribution over an arbitrary supplied label set. Inventing that mapping would measure our assumption.Needle 3 (Cactus Compute) — Its native response is a chosen label plus one accept/refuse confidence, not a categorical distribution over the supplied labels. The options-as-tools adaptation remains published as a partial run.Succinct Router 14M (Pedro Marques) — A router over three fixed GPT settings, not a general typed-decision model.jev-model-router, Director, Loki (various) — Applications built on decision models, not decision models themselves.ProgramAsWeights (ProgramAsWeights) — The compiler still requires GitHub authentication and the available path would expose held-out rubrics to a third party. No public weights or anonymous endpoint are available.EigenJev (EigenJev) — The endpoint requires authentication and no public weights or runnable implementation are published.NanoJev (NanoJev) — Public weights exist, but the server exposes a different schema (including boolean rather than Noul) and lacks the full structured/null contract. It needs substantive compatibility work before a fair full-suite run.Werr (pCwOrM) — Its documented server imports a module (scratch.jevbench_eval.optimize_werr_jevbench) that is not in the public repository, so the submitted configuration cannot be started; its engine also sends telemetry about each request to an outside server by default.DIY Jev (VakeDomen) — The repository named in the request (github.com/VakeDomen/DIY-Jev) answers 404, so there is nothing to run.SimpleJev RWKV variants (SimpleJev) — The public demo exposes RWKV IDs, but it does not identify their exact checkpoints or licences. Without reproducible model provenance, we do not publish benchmark rows for them.Method and tiersGeometric mean of Intelligence, Calibration, Speed and Cost, 25 % each. If chance-corrected Intelligence is below 50, multiply by (Intelligence / 50)^2; at or above 50 there is no penalty.Revision v1.3.0. v1.3.0 scoring-only release: Intelligence is chance-corrected per tier and scores below 50 receive the growing near-chance penalty. Calibration, Speed, Cost, ranking eligibility, tasks and measurements are unchanged.easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchangedjudge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchangedhard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.v1.2 numbers are not comparable with v1.1 or v1.0 (different tiers and scoring). The v1.0 page keeps its own numbers, calibration plots and per-family tables.Limits534 decisions is a pilot, not a census, and it is English-only.The weights are a choice. The JevBench Score weights the four axes equally and multiplies rather than adds them; if a wrong decision costs you more than a slow or expensive one, pick “Emphasis on Accuracy” above — the table of views shows what other weightings would do.The latency adjustment (×2, +0.15 s on our own servers) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. The official Jev API is presumably under high load, given the public interest. Serving under load trades per-user speed for throughput: in the NVIDIA chart shown by SemiAnalysis, moving to the throughput-maximising setting cuts per-user tokens per second by far more than 2×. That chart is a 1.8T mixture-of-experts model on GPU clusters, not a 4B model on one GPU, so it supports the direction and size of the effect, not our exact factor. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo, and a measurement under load is planned.Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.Latency is one origin at one time of day; hosted endpoints, public demos and a local CPU are different kinds of latency. Public demo endpoints are shared with everyone else using them.Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.CreditHarness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.Alibaba GTE Reranker ModernBERT-base — Alibaba-NLP, Apache-2.0 — huggingface.co/Alibaba-NLP/gte-reranker-modernbert-baseBAAI bge-reranker-v2-m3 — BAAI, Apache-2.0 — huggingface.co/BAAI/bge-reranker-v2-m3Bespoke Nimble 9B (Bespoke Labs) — Bespoke Labs, Apache-2.0 (weights); repository without a licence file as of 19 Sep — github.com/bespokelabsai/nimbleCerto v1 (AltSlate Labs) — AltSlate Labs, MIT — huggingface.co/altslate/certo-decision-modelclassifier.dev (fast tier) — mrmps (@michael_chomsky), MIT (code); hosted service — classifier.devdecider-2b (Mapika) — Mapika, Apache-2.0 — huggingface.co/Mapika/decider-2bdecider-35b-a3b (Mapika) — Mapika, Apache-2.0 — huggingface.co/Mapika/decider-35b-a3bdecision-machine-1 (milliseconds.ai) — milliseconds.ai (Baptiste Laget), proprietary API, closed weights — www.milliseconds.aiDeepSeek V4.1 Flash (thinking default) — DeepSeek, open weights, proprietary API route — api-docs.deepseek.comdjev (Maisa, diffusion-gemma) — Maisa (David Villalón), Apache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weights — github.com/Davipar/djev-devdjev (thinking) — David Villalon / Maisa, Apache-2.0 — github.com/Davipar/djev-devGemini 3.1 Flash-Lite — Google, proprietary API — ai.google.devGLiNER2 (Fastino, gliner2.5-base) — Fastino, Apache-2.0 — github.com/fastino-ai/GLiNER2GLiNER2 large (Fastino) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2-large-v1GLiNER2.5 multi (Fastino, 287M) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2.5-multi-v1GLiNER2.5 small (Fastino, 74M) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2.5-small-v1GPT-5.6 Luna (low reasoning effort) — OpenAI, proprietary API — platform.openai.comjeff (Logan Markewich, GLiFormer 400M) — Logan Markewich, MIT (code); GLiFormer weights per their model card — github.com/logan-markewich/jeffJev 1.13.0 (TypeSafe AI) — TypeSafe AI, proprietary API — docs.typesafe.aijev-local (Qwen3.5-9B) — us (GitHub), no licence stated in the repository (public code); Apache-2.0 base weights — github.com/us/jev-localjqv (Qwen3-32B zero-shot) — hjmurmur (Octalab), Apache-2.0 (Qwen3-32B weights); serving code public — github.com/Octalab-Inc/jqvkev 0.5B — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kevkev 0.6B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kevkev 4B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kevkev 8B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kevLaya (Convai Innovations, ModernBERT-large 421M) — Convai Innovations, Apache-2.0 — huggingface.co/convaiinnovations/layaLitJev (Qwen3.8-27B) — Zhengxu Yu, Apache-2.0 (code); Apache-2.0 base weights — github.com/zhengxuyu/litjevMixedbread mxbai-rerank-base-v2 — Mixedbread, Apache-2.0 — huggingface.co/mixedbread-ai/mxbai-rerank-base-v2Needle 3 (Cactus, 2-bit, local CPU) — Cactus Compute, Apache-2.0 (model and package) — github.com/cactus-compute/needleopen-alternative-jev (Qwen3.5-4B, IkerMoel) — IkerMoel, Apache-2.0 (code and weights) — github.com/ikermoel/open-alternative-jevOpen-Jev 2B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-JevOpen-Jev 9B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-Jevopen-jev-deberta-v3-large (local CPU) — Kotoba Labs, Apache-2.0 (model card); DeBERTa-v3 keeps its own terms — github.com/kotoba-lang/typed-decisionsOpenDecision (ModernBERT-large zero-shot) — Deepan Wadhwa, Apache-2.0 — github.com/deepanwadhwa/OpenDecisionOpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) — razorback16 / Codiv, Apache-2.0 (repo and weights) — github.com/razorback16/openjevOpenJev (thinking, BF16) — razorback16, Apache-2.0 — github.com/razorback16/openjevopenJev Verdict (heman10x, ModernBERT-base 151M) — Hemant (heman10x), Apache-2.0 — github.com/Heman10x-NGU/openJev-verdict-2.0openJev Verdict 1.4 — Hemant (heman10x), Apache-2.0 — huggingface.co/heman10x/rlcd-modernbert-151mopenjev-sglang (Qwen3.6-35B-A3B on SGLang) — ekzhang, no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms — github.com/ekzhang/openjev-sglangQwen3-Reranker-4B — Qwen, Apache-2.0 — huggingface.co/Qwen/Qwen3-Reranker-4BQwen3.8 27B (Chutes TEE) — Qwen / Chutes, open weights — chutes.aireflex 4B (kshetrajna12) — kshetrajna12, MIT (code, adapter); Apache-2.0 (base) — github.com/kshetrajna12/reflexreflex-27b (Qwen3.8-27B) — kshetrajna12, MIT code; Apache-2.0 Qwen weights — github.com/kshetrajna12/reflexSemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) — Theodore Lee (TheoLeeCJ), MIT (code); Qwen3.5 weights Apache-2.0 — github.com/TheoLeeCJ/openjevSimpleJev Qwen3.6-35B-A3B — Featherless AI, Apache-2.0 (Qwen weights); repository licence not stated — github.com/featherless-ai/simple-jevSimpleJev Qwen3.8-27B — Featherless AI, Apache-2.0 (Qwen weights); repository licence not stated — github.com/featherless-ai/simple-jevsmalljev semantic-v9 — Aditya (isHeSatoshi), Apache-2.0 — github.com/isHeSatoshi/smalljevsystem-one (Qwen3-8B, Sean Goedecke) — Sean Goedecke, no licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0 — github.com/sgoedecke/system-onesystem-one-open (Gemma 4 E2B LoRA on an L4) — mithalouni, MIT (repository LICENSE; Gemma weights keep Google’s terms) — github.com/mithalouni/system-one-openWinnow-12B Q8 — Eldan Ring, Apache-2.0, including the applicable Gemma 4 base/derivative licence terms — huggingface.co/EldanRing/Winnow-12BZeroEntropy zerank-2 — ZeroEntropy, Apache-2.0 — huggingface.co/zeroentropy/zerank-2-rerankerAuthors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.