Back to Blog
Tools & Resources6 min readSeptember 29, 2026

Jeff Is an Open, Self-Hosted Jev Alternative at 0.8B Parameters

Jeff is a free, Apache-licensed 0.8B model that copies Jev's decision API and, per its README, answers in about 28 ms on a Mac. When a SaaS team should self-host classification.

Emma Watson

Emma Watson

Growth at NeedBase

A free, open-weight copy of TypeSafe's Jev API reached the Hacker News front page today. Jeff is a set of fine-tuned Qwen3.5 and Gemma 4 models built for one job: you describe a situation and list some options, and it returns a probability for each option. It doesn't generate text, so there's nothing to parse. According to the README, the smallest model has 0.8 billion parameters and makes a decision in 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max. The weights are Apache 2.0 and the code is MIT.

If your product calls GPT or Claude just to answer "which queue does this ticket go to?", this matters to you.

What Jev is, and what "Jev-compatible" means

Jev is a hosted model from TypeSafe AI, released this month in early access behind a waitlist. TypeSafe calls it a "System One model". It takes a state, which can be text or JSON, plus a set of typed questions, and returns calibrated probabilities instead of prose. There are three question types: choice (pick one option), score (a point on a scale) and noul (yes or no). MarkTechPost reports the price as $42 per billion input tokens, with output free, and says TypeSafe admits it cannot prove that price is unsubsidised. TypeSafe has not disclosed the model's size, and you cannot self-host it.

"Jev-compatible" means Jeff accepts the same request shape at the same /v1/systemone endpoint, so code written against Jev should point at a Jeff server with little change. The README says the project is not affiliated with TypeSafe. It started as a fork of AutoJev, an earlier open recipe that fine-tuned a 27B Qwen model for the same purpose.

How it was built

The headline claim is that everything ran at home. According to the README, the 0.8B model trained in about two hours on a single RTX PRO 6000 workstation GPU. The synthetic training data was written by an open model running on two Nvidia DGX Sparks, and testing was done on a MacBook. The author says no closed-model output went into the training data. The README gives no money cost.

That is serious workstation hardware, and commenters on Hacker News pointed out that renting would be saner for most people. You only need it to reproduce the training, not to use the models.

What the benchmarks actually say

On the author's five-benchmark panel, Jeff-0.8B scores 79.1 overall and the 2B scores 83.1. Jev's published figure is 83.0. The breakdown matters more than the headline.

Classification and grounding are strong. The 0.8B scores 96.4 on Financial PhraseBank (sentiment) and 86.1 on RAGTruth (hallucination detection). Jev's published figures there are 77.0 and 77.3.

Reasoning is weak. The 0.8B scores 64.0 on BBH against Jev's 94.3, and 47.6 on JevBench's hard tier against 73.3. The README says so plainly: "Small models don't reason."

There are two caveats. First, the README notes that the Jev figures were measured on a different sample of the same benchmarks, so this isn't a strict head-to-head. Second, benchmarks are not your data. Commenters on Hacker News reported mixed results. One said they got roughly 70% with Jeff against 94% with Jev on their own classification task. Another said the 0.8B was "completely useless" for tagging job ads, and the 2B was better but still missed fields.

The author's answer is fine-tuning. The README describes a voice-navigation fine-tune on about 11,000 examples that took roughly half an hour on one GPU and lifted held-out accuracy from 31.7% to 95.8%.

What this means for a SaaS team

Most SaaS products have a handful of small decision points that currently go to a large LLM: routing support tickets, tagging intent, flagging abuse, deciding whether a message needs a human, or picking which expensive prompt to run. A model like Jeff turns each one into a local function call measured in tens of milliseconds, with no charge per request.

It is worth it when you make a high volume of short classification calls, when latency sits in the user's path (autocomplete, voice commands, live routing), when your data can't leave your infrastructure, or when you're tired of parsing JSON out of chat completions.

It isn't worth it when the decision needs multi-step reasoning, or when you have more than 26 options: the README says options past position 26 are effectively never chosen, so the server refuses those questions. It also isn't worth it if you'd have to run a GPU just for this. On a 32-thread CPU the README reports 463 ms for the 0.8B, which is slower than many hosted APIs.

Keep the older options in mind too. Several commenters on Hacker News argued that for a fixed set of labels you'll often do as well or better with a sentence-embedding model and a logistic regression or SVM, or with a fine-tuned ModernBERT. One reported 93.0% on Banking77 with all-MiniLM-L6-v2 and a simple classifier. What Jeff and Jev offer instead is zero-shot use: you change the options by editing a sentence, not by retraining.

What to actually do

1. Find your decision calls. Search your codebase for LLM calls whose output is really a label, a yes/no or a score. Log real inputs from each one along with the answers you currently accept.

2. Run Jeff locally. Build a test set of a few hundred labelled examples first. The quick start uses uv, downloads Jeff-Qwen3.5-0.8B from Hugging Face and serves it on port 8765. There's an MLX backend for Apple silicon Macs. Send it your test set and measure accuracy and latency on your own hardware.

3. Compare against a baseline. Try the same test set with your current LLM and with an embeddings-plus-classifier pipeline. If you can get off the waitlist, try Jev as well. Pick the cheapest option that clears your accuracy bar.

4. Write options as consequences. The README's advice is to "reason in code, decide with Jeff". Describe each option clearly and consistently, use short keys, and do any calculation in your own code before you ask the question.

5. Keep a fallback. Use Jeff's probabilities to send low-confidence cases to a larger model or to a person.

The bottom line

Jeff shows that the decision-model idea behind Jev doesn't depend on one vendor. A 0.8B open model can take the same requests, run on a laptop in about 30 ms, and cost nothing per call. It is a classifier, not a thinker, and early reports from real users are mixed. Treat it as a candidate to test against your own data, not a drop-in replacement. If you run enough small labelling calls to notice them on your API bill, spend an afternoon benchmarking it.

Found this useful?

Share it with a founder who needs it.

Ready to launch your product?

Join thousands of makers who launched on NeedBase.

Submit Your Product โ†’