Jeff: Home-Trained Jev-Compatible AI Models Hit 28ms Decisions
AI News

Jeff: Home-Trained Jev-Compatible AI Models Hit 28ms Decisions

5 min
9/29/2026
Jeff AIJev-compatibledecision modelsopen-source AI

Jeff: The 0.8B Decision Model That Trained at Home and Matches Jev's Accuracy

In the rapidly evolving landscape of AI decision models, a new open-source project is turning heads. Jeff, created by developer Firelex, is a family of small, Jev-compatible decision models that deliver remarkably fast, accurate classification without requiring cloud GPUs or massive infrastructure. The project, which hit the Hacker News front page on September 28, 2026, demonstrates that high-performance AI decision-making can now happen on a single workstation GPU.

Jeff is not just another clone of TypeSafe's proprietary Jev system. It's an independent implementation that speaks the same request format, making it a drop-in replacement for developers who want local, private, and fast decision-making. With the 0.8B model achieving 79.1% accuracy on a five-benchmark panel and the 2B variant hitting 83.1%—essentially matching Jev's published 83.0%—Jeff proves that small models can compete with much larger systems on classification tasks.

What Makes Jeff Different: Speed and Local Training

The headline number is speed. Jeff's 0.8B model makes a single decision in about 22 milliseconds on an NVIDIA RTX PRO 6000 and 28ms on an Apple M4 Max using MLX. That's a fraction of the 114–212ms per call that Jev's API reportedly takes, even accounting for network latency. The key architectural choice is that Jeff returns calibrated probabilities in a single forward pass—no text generation, no parsing, just pure classification.

Equally impressive is how Jeff was trained. Everything ran on local hardware: one RTX PRO 6000 for training (about 2 hours for the 0.8B, 3.5 for the 2B), two DGX Sparks running Qwen3.8-Flash-Next to generate synthetic training data, and a MacBook for testing. No cloud GPUs were used, and no closed-model outputs appear in the training data. A closed model was only used to spot-check data quality. This is a fully reproducible, home-grown AI project.

Benchmark Performance: Where Jeff Shines and Where It Struggles

Jeff's performance is nuanced. On classification-heavy benchmarks, it doesn't just match Jev—it beats it. On Financial PhraseBank, Jeff-Qwen3.5-0.8B scores 96.4% versus Jev's 77.0%. On RAGTruth, the 2B model hits 88.9%, tied with the much larger AutoJev-27B. These are tasks where the model must choose between options, which is exactly what Jeff is designed for.

However, on reasoning-heavy benchmarks like BBH, JudgeBench, and JevBench's hard tier, Jeff falls well short. The 0.8B scores 64.0% on BBH versus Jev's 94.3%. This is expected—Jeff is a System 1 model, built for fast, calibrated decisions, not multi-step reasoning. The project's README is refreshingly honest about this limitation, stating: "Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning."

continue reading below...

The Games Test: Zero-Shot Decision-Making in Action

To test Jeff's zero-shot capabilities beyond benchmarks, the team had it play three classic games: Doom, Frogger, and Pac-Man. Each turn, the model receives a description of the situation and the legal moves in words, and it picks one. The results are illuminating. Jeff-0.8B matches a hand-coded rule bot in Doom (6.55 kills) and nearly matches it in Frogger (10.3 crossings vs 10.25). In Pac-Man, it collects 57 of 98 pellets, a massive improvement over the untrained model's 25.8.

These games serve as a stress test for zero-shot generalization. The tasks are unlike anything in the training data, yet Jeff still makes sensible decisions. Notably, the 0.8B model outperforms the 2B on games, despite scoring lower on benchmarks—a cautionary tale about relying solely on benchmark scores.

Developer Experience: API, Fine-Tuning, and Integration

Jeff is designed to be slotted directly into code. The API is Jev-compatible, so developers can switch with minimal changes. A simple curl command can send a situation and a list of options, receiving back calibrated probabilities for each. Three question types are supported: choice (pick one of up to 255 options), noul (yes/no probability), and score (a point on a described scale). Multiple independent questions can be batched in one request.

For those who need better accuracy, Jeff supports fine-tuning. The team reports that a voice-navigation fine-tune on ~11k app-specific examples took about 30 minutes on one GPU and improved held-out accuracy from 31.7% to 95.8%. The training pipeline is fully open-source, starting from the AutoJev recipe, and includes a leak filter to prevent data contamination.

Licensing and Availability

Jeff is released under permissive licenses: the code is MIT (including AutoJev's original copyright notice), and the model weights are Apache 2.0. Three models are available on Hugging Face: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B. The training data itself is not released, as some sources are share-alike (CC BY-SA), but all sources are documented.

This open approach stands in contrast to Jev, which is API-only with undisclosed model size. For developers who value privacy, local execution, and cost efficiency, Jeff presents a compelling alternative.

Why Jeff Matters for the AI Ecosystem

Jeff represents a growing trend: the democratization of AI decision-making. It shows that you don't need a 27B model or a cloud API to get good zero-shot classification. A 0.8B model trained on a single GPU can match or beat much larger systems on the right tasks, at a fraction of the latency and cost.

The project also validates the AutoJev training recipe, proving it can be applied to smaller student models with excellent results. As Firelex puts it: "Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe." This independence is crucial for the open-source ecosystem, offering a viable path forward for developers who want Jev-like functionality without vendor lock-in.

While Jeff won't replace large reasoning models for complex tasks, it fills a vital niche: fast, reliable, local decision-making for routing, moderation, intent classification, and more. With sub-30ms latency and Apache 2.0 weights, it's ready for production use today.