JEV-27B decision model keeps the Qwen backbone frozen
AutoTrust's JEV-27B decision model is an Apache-2.0 add-on to a frozen Qwen, not a new base model. The company blog and a September 28 newsroom piece agree on the shape and disagree, slightly, on the calendar.

in this block
The JEV-27B decision model is AutoTrust AI's open attempt to answer the small, typed questions agents ask all day, without fine-tuning away the model that still has to write. The Hugging Face post is dated September 27, 2026. RuntimeWire, published September 28 at 2:25 p.m. CT, says the release was September 28. Both describe the same object: Apache-2.0 weights, a frozen Qwen3.8-27B backbone, and a detachable decision block.
TL;DR - AutoTrust says the trained decision block is 108.9 million parameters, about 0.4% of the full model, and that training took about 9.2 hours on one NVIDIA B200. The Hugging Face post also lists a 24-slot decision head of 123 thousand parameters inside that budget. - On six benchmark groups the company ran itself, the mean is 84.07 for JEV-27B and 83.85 for hosted TypeSafe Jev 1.13, a gap of 0.22 percentage points. JEV-27B is ahead on four groups and behind on two. - An independent set, gazelle93/decision-models-under-pressure, is the check that is not the company's own labels. At 16 options, AutoTrust says its model reaches 96% of Jev's published accuracy. RuntimeWire stresses that this still trails the hosted model.
What actually happened
AutoTrust's post frames the job as System 1 versus System 2. System 1 is a fast typed decision: yes or no (noul), a choice among 2 to 16 options, or a score from 0 to 5. System 2 is ordinary generation and step-by-step reasoning, left on the original language-model head. The backbone is named as Qwen/Qwen3.8-27B, and the company says that tower is bit-identical to the release and was not trained. The decision path is a LoRA of rank 16 plus a small head. In serving, a request either calls the LoRA module jev-decision or leaves it off.
The reason they refuse to merge the adapter is a measurement from the earlier 9B model, not a slogan. Merging, they write, dropped HumanEval from 70.7% to 61.6% even though prose perplexity barely moved, from 3.15 to 3.30. On this 27B pair they say HumanEval is 78.0% (128 of 164) with the decision block on or off, and that all 164 completions are byte-identical to the base model. RuntimeWire repeats the byte-identical claim and adds the right caveat: that test covers the coding outputs they checked. It does not prove every prompt behaves the same in every deployment.
The teacher is closed. AutoTrust says TypeSafe Jev 1.13 is a hosted model from TypeSafe AI, that AutoTrust is not affiliated with it, and that the student was trained on published output distributions in the Apache-2.0 corpus SargeDev/jev-distill-corpus-v3. RuntimeWire names co-founder and CEO Daniel Tang and a 2025 PhD from the University of Luxembourg. Those lines are not in the Hugging Face post opened here.
Speed claims are the company's, and RuntimeWire flags the comparison. The Hugging Face post says a median 137 milliseconds per decision on one B200, about 130 decisions per second where the hosted API did 23, and 4.2 milliseconds per decision at a batch of 128. Third-party API figures in that post are a 238 millisecond mean and a 291 to 301 millisecond median, and the post says those include the network. A local timer is not a bake-off. The same post says full bf16 fits one H100 80 GB at about 100 decisions per second, against about 220 on a B300 in their deployment. The download line says about 54 GB. RuntimeWire says roughly 54.7 GB. Two RTX 5090s are described as untested estimates. None of this is a price.
What the score does and does not mean
The head-to-head table is dated as a run on September 26, 2026. Means: JEV-27B 84.07, hosted Jev 1.13 83.85. The company is ahead on JevBench (88.70 versus 87.18), OpenJev text (73.89 versus 72.96), Nimble (92.91 versus 91.84), and MASSIVE-en (87.71 versus 87.14). Jev is ahead on Kev by 1.77 points (85.52 versus 83.75) and on VitaminC by 1.00 point (78.46 versus 77.46). The open baselines in that table, including NeoHorse-Jev-4B at a mean of 77.70, are values the company says it did not re-run. A lead of 0.22 points, measured by the student, is a press-release-sized number. RuntimeWire says so in plainer words: the edge is slim and company-measured, and a deployment test matters more than the headline average.
Distribution match uses the teacher's own outputs as labels. On 25,376 held-out questions, AutoTrust reports mean KL of about 0.017, an expected calibration error of 0.0009, and temperatures near 1.00. The company says those rows sit in the same 53 domains as the training data, so the match shows fidelity, not truth. A published poker spot makes the same point: a solver always checks, Jev shoves with 0.62, and JEV-27B shoves with 0.63. Option order alone flips about 7.0% of Jev's 16-way answers and 7.4% of the student's.
The independent benchmark is the paragraph to keep. AutoTrust's table, using published Jev numbers and its own run, shows accuracy at 2, 4, 8, and 16 options of 0.890, 0.801, 0.782, and 0.769 for Jev, against 0.876, 0.784, 0.767, and 0.740 for JEV-27B. That is the 96% figure at 16 options. RuntimeWire calls it the strongest external check in the release and notes the student remains behind. The company says none of that benchmark's data is in the training set, and that closer open models the benchmark's author ran, Laya and zero-shot DeBERTa-v3-large, landed at 90% of Jev. Community leaderboards, the company says, are still to come.
What a frozen backbone changes for a buyer
The JEV-27B decision model does not remove the cost of a 27B model. It moves that cost onto your machines. RuntimeWire's point is the one a procurement note should steal: a single B200 headline is not "runs on an ordinary server." Local inference can be the right answer when a hosted API is a privacy problem. It is a worse answer if you needed the API's 23 decisions per second and you do not have the card. The smaller sibling, JEV-9B, is described as 18 GB in bf16, 2.6 times faster, with mean KL about 0.019 to Jev and 90% of Jev's accuracy at 16 options. The company says the 27B is the better choice when the option list is long or System 2 still matters.
A pattern the post suggests, and says it has not benchmarked, is to let the decision head answer when it is confident and to hand the rest to generation on the same engine. Thresholds are yours. The post is explicit that known weaknesses of the teacher, including multi-hop reasoning, arithmetic, counting, and adversarial inputs, are inherited. Confidence gating is not a safety proof.
For another small model aimed at decisions rather than essays, see the Strands Decider note. For a different decision index with its own name collision risk, see the Fastino GLiDE note. AutoTrust's name warning belongs in the same drawer: JEV-27B is not TypeSafe Jev, and the post says they share no weights and no code.
What to do as a reader (not a trade)
This is not investment advice, and there is no token in either opened page. The JEV-27B decision model is a weights release, which is a different object from a ticker. If you are deciding whether to try the weights, separate three claims about the JEV-27B decision model. The architecture claim, a frozen backbone plus a small head, is described consistently by the company post and by RuntimeWire. The 0.22 point mean lead is AutoTrust's run of both sides. The 96% of Jev at 16 options is the number that comes from a benchmark the company says it did not write. Ask for the leaderboard submission they say is still pending before you repeat "on par" as a fact about your tasks. Until that submission exists, the JEV-27B decision model is a documented student, not a crowned one.
Write down the calendar instead of merging it. Blog post September 27. RuntimeWire's release date September 28. Benchmark run September 26. Those are three stamps, not one. The primary write-up is AutoTrust's Hugging Face post. The second page is RuntimeWire's September 28 article, which is useful because it refuses to treat the company's average as an independent verdict.
Not financial advice. DYOR, ser.