Jeff is a family of very small decision models from the independent developer firelex, released on 28 September 2026. The premise is narrow on purpose: you describe a situation, list the options in plain words, and Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing, no schema to coax out of a chat model. Version 1.1 fine-tunes Qwen3.5 and Gemma 4 into 0.8B and 2B checkpoints that choose among up to 254 options, up from 26 in v1.0, and the README reports about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max through MLX. The whole project was trained on local hardware: one workstation GPU for training, two DGX Sparks for the synthetic data, and a MacBook for testing.
Key Features
- Probabilities instead of prose: Jeff returns a calibrated score per option in one forward pass, so downstream code branches on a number rather than parsing free text.
- Zero-shot options: The categories do not need to be in the training data. You describe them at call time, which suits support queues, intent routing, moderation labels, and voice commands.
- Two sizes for different budgets: The 0.8B checkpoint runs anywhere; the 2B trades some latency for accuracy.
- Jev-compatible request format: If you have written code against TypeSafe's Jev API, the same request shape applies, which makes it a plausible local fallback.
- Long lists actually work: The v1.1 release raised the option ceiling from 26 to 254, and the 0.8B model moved from 40 percent to 95 percent on the project's long-list test.
- Honest documentation: The README states plainly that these models approach and sometimes beat Jev on benchmarks while not matching Jev's reasoning, which is a rare thing to read in a model card.
Use Cases
Who Should Use This Tool?
- Developers adding a routing layer: When an agent or pipeline needs "which of these 40 intents is this", a small local model is cheaper and more predictable than a frontier call.
- Teams with data that cannot leave the building: Training and inference both run on local hardware, so nothing is sent to a hosted API.
- Anyone doing classification at volume: At tens of milliseconds per decision, the economics differ sharply from per-token API pricing.
Problems It Solves
- Wasting a frontier model on a small judgement: Deciding between 12 support categories does not need a reasoning model and its bill.
- Unreliable output parsing: Returning probabilities removes the brittle step where you scrape a label out of generated text.
- Data residency: Local training and inference keep sensitive classification data on your own machines.
Pricing
Jeff is free and open source under the MIT license, with weights published on Hugging Face and version 1.0 kept available as a revision. Your cost is hardware. The README notes that the 0.8B trains in roughly two hours and the 2B in about three and a half hours on one RTX PRO 6000, so a short fine-tune on your own examples is measured in an afternoon, not a quarter.
Advantages & Unique Selling Points
Compared to Competitors:
- Speed at the edge: Single-digit-to-tens-of-milliseconds decisions make it viable inside a request path, which is where most hosted classifiers fall down.
- Reproducible provenance: All training data was synthetic, written by an open model, with a closed model used only to spot-check sample quality.
- Fine-tuning is documented as the escape hatch: The README reports a voice-navigation fine-tune that moved held-out accuracy from 31.7 percent to 95.8 percent in under half an hour.
What Makes It Stand Out:
- Same request format as Jev, so migration is a configuration change.
- The 254-option ceiling turns it from a toy into a usable router.
- Calibration is the product, not an afterthought.
User Reviews
The project arrived on Hacker News in late September and drew the reaction small-model releases usually get: immediate interest from people who want a cheap routing layer, and immediate scepticism from people who assume a 0.8B model must be a toy. The benchmark numbers are self-reported, and the honest caveat in the README about reasoning capability is doing a lot of work here: it sets expectations at "fast, well-calibrated judgement" rather than "replaces your model". Treat the accuracy figures as the author's, and re-measure on your own data before you depend on them.
Getting Started
Quick Start Guide
- Pick a size: Start with the 0.8B checkpoint and confirm latency on your own hardware.
- Describe your options: Write the categories in plain language at call time; there is no label file to fit.
- Call it like Jev: Send the situation and the option list in the documented request format.
- Measure before trusting: Run your own held-out set and compare calibration, not just top-1 accuracy.
- Fine-tune if it is close: If accuracy is nearly good enough, a short fine-tune on real examples is the documented next step.
Integration
- Local inference stacks including MLX on Apple silicon and CUDA on RTX hardware.
- Existing Jev client code, since the request format is compatible.
- Agent routers and pipelines that need a fast classifier step ahead of a larger model.
- Hugging Face for weights, including the pinned v1.0 revision.
Frequently Asked Questions
Is Jeff a replacement for Jev?
No. The README positions it as a much smaller model that approaches Jev on classification benchmarks while not matching its reasoning, and suggests a fine-tune when zero-shot accuracy is not enough.
How many options can it choose between?
Version 1.1 supports up to 254 options, up from 26 in version 1.0.
How fast is it?
The project reports about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max via MLX. Measure on your own hardware before you design around those numbers.
Can I fine-tune it on my own data?
Yes, that is the documented path when zero-shot accuracy is insufficient, and the README describes the training hardware and time involved.
Does it send data anywhere?
Training and inference both run locally. Nothing about the workflow requires a hosted API.
Alternatives
If Jeff is not the right fit, consider these alternatives:
- Jev: TypeSafe's hosted System One model when you want the original API and larger reasoning capacity.
- LangGraph: Better when the missing piece is orchestration around several decision calls rather than the classifier itself.
- MiroFish: A different local-first approach, aimed at multi-agent simulation rather than classification.
- Buzz: When the missing piece is a workspace for agents, not a decision model.
Tips & Best Practices
- Calibrate, do not just rank: A probability you can threshold is more useful than a label, but only if you check the calibration on your own data.
- Keep option phrasing stable: Zero-shot means the wording of your options is part of the interface, so version that text like you version code.
- Try the smaller checkpoint first: Latency budgets are easier to reason about when you start from the floor.
- Fine-tune with real examples: Synthetic data gets you started; production examples get you the last 60 points.
Conclusion
Jeff is a well-scoped tool: it does one thing, returns a number instead of a paragraph, and runs on hardware you already own. If your bottleneck is a routing or classification decision sitting in front of a larger model, the 0.8B checkpoint is cheap enough to test this week, and the fine-tuning path documented in the README is what turns it from a clever demo into a component you can ship.
Comments
No comments yet. Be the first to comment!