# Jeeves Lets a 0.3-Second Classifier Think. Its p90 Reaches 17.1 Seconds.

**Summary:** Jeeves lifts development-set accuracy from 0.775 to 0.825, but full reasoning reaches a 17.1-second p90. I would use its confidence gate to reserve that delay for uncertain decisions.

- Canonical: https://markhuang.ai/news/jeeves-reasoning-17-second-tail
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-29
- Section: News
- Tags: Jeeves, Decision Models, AI Inference, Model Evaluation, Open Source AI
- Source: [PostHog Jeeves on GitHub](https://github.com/PostHog/jeeves)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![Fast decision tokens take a direct route while one uncertain token travels through a deeper reasoning path](https://cdn.markhuang.ai/news/jeeves-reasoning-17-second-tail/hero.webp)

*Jeeves makes the routing choice visible: keep the easy decisions fast, and spend reasoning time where uncertainty can justify it.*

[PostHog's Jeeves](https://github.com/PostHog/jeeves) adds a reasoning stage to a 9B typed-decision model. On 325 development questions, full reasoning raised accuracy from 0.775 to 0.825. It also moved latency from about 0.3 seconds without reasoning to a 3.3-second median and a 17.1-second p90.

My read is that the extra accuracy earns Jeeves a place in a decision pipeline, though I would hesitate to put every request through it. A classifier is attractive because it turns a piece of state into a bounded answer and calibrated probabilities without generating an essay. Make every easy decision wait for a long reasoning chain, and part of that advantage disappears.

The useful product control is already in the repository: Jeeves can skip reasoning when its initial confidence clears a threshold. I would build the deployment around that switch and pay the extra latency only when the first answer is uncertain.

## Reasoning improves the answer, then changes the service

Jeeves starts with Qwen3.5-9B, adds LoRA weights and a pointer head, and trains the model to answer yes/no, multiple-choice, and rating questions through a Jev-compatible API. It can produce a reasoning chain before the pointer head assigns probabilities to the allowed options. The code is MIT licensed. The released weights inherit Qwen3.5's Apache 2.0 license, and each source dataset keeps its own license.

The gains are substantial on the author's test setup. With reasoning, Jeeves scored 0.889 on an out-of-domain and held-out test aggregate, ahead of the published comparison figures of 0.822 for Kev-9B and 0.857 for Jev. On the same checkpoint, reasoning improved the project's 2,962-item test split from 0.804 to 0.840.

The gain matters, but requests do not benefit equally. Jeeves still trails Jev on knowledge-heavy tests: 0.793 versus 0.900 on MMLU, and 0.739 versus 0.840 on MMLU-Pro. More inference time did not erase the limits of the 9B base model.

## The middle setting looks more like a product

Jeeves exposes three useful operating points. Full reasoning reached 0.825 accuracy on the 325-question development set, with 1,138 mean reasoning tokens and a 17.1-second p90. Turning reasoning off produced 0.775 at about 0.3 seconds. A hybrid setting capped reasoning at 768 tokens and skipped it above a 0.9 confidence threshold; that reached 0.806 accuracy, with a 2.0-second median and 5.6-second p90.

The hybrid is where I would begin, followed by recalibration on the decisions that matter to my application. There is no universal threshold. [Kev-9B's model card](https://github.com/jaredpalmer/kev/blob/main/docs/model-cards/kev-9b.md) shows why: its shipped temperature fit some external workloads but needed refitting on another, where out-of-fold calibration lowered expected calibration error from 0.131 to 0.037 without changing accuracy.

A confidence gate works only when its confidence means something on local data. I would measure how much traffic stays within an acceptable error rate, then send the uncertain or costly cases to reasoning while obvious routing decisions remain fast.

> **Info:**
>
> My deployment rule: reasoning should spend a declared latency budget on uncertainty. If a team cannot name the errors that deserve the slow path, it is too early to make thinking the default.

## The 0.935 score has a public boundary

Jeeves also reports 0.935 accuracy across 231 public JevBench items, compared with 0.866 for Jev, and 0.865 on the 111 public hard items. Those numbers got my attention. The source is also clear that the sealed judge tier is excluded.

An [independent audit of the same public subset](https://github.com/Zefan-Cai/Open-Jev/blob/main/docs/jevbench-public.md) describes its composition as 72 original, 48 easy, and 111 hard tasks. It also notes that the full benchmark contains another 303 private or judge tasks. The current [JevBench project](https://github.com/fstandhartinger/jevbench) now uses public and sealed decisions in its official score precisely because public items can be trained on or selected against.

Nothing I inspected shows that Jeeves trained on JevBench. The repository labels the result as public and makes no claim about the sealed tier. I read 0.935 as a strong public-set result, while the production error rate remains unknown. I made the same distinction when looking at [OpenJev's smaller comparison](https://markhuang.ai/news/openjev-3-8-point-102-row-asterisk). A reproducible interface and an encouraging score are enough to justify the next test, but they cannot replace it.

## The H100 belongs in the decision

The latency figures come from one H100, and the FP8 inference kernel requires NVIDIA Hopper hardware. Reproducing the full training pipeline calls for eight GPUs. Jeeves does include the weights, training code, data-building scripts, an SDK, and a diffusion drafter that the project says raises single-question chain generation from 109 to 176 tokens per second at its default block size.

Open code makes the system inspectable, but teams still have to pay for hardware and remeasure latency on their own setup. A team considering Jeeves needs to measure memory use, throughput under concurrency, queueing at the p90, and the cost of keeping suitable GPUs available. The repository does not publish a dollar cost per decision.

I would also keep returned reasoning out of user-facing logs until I had a reason to expose it. The project warns that it did not train for language consistency, so the chains are not reliably interpretable. The typed answer and probability distribution are the contract. The chain is an internal computation, not an audit explanation.

## My pilot would test the router

I would start with a labelled slice of the actual workload and run all three modes. I want to see how accuracy, calibrated coverage, p90 latency, and GPU demand move as I change the confidence threshold. Then I would inspect two failure groups closely: easy decisions that waste time reasoning, and confident mistakes that never reach the slow path.

Jeeves makes a credible case that a small typed model can benefit from deliberate reasoning. Its public benchmark result deserves a closer look, and the repository is unusually direct about the slower tail and weaker knowledge scores. The 17.1-second p90 does not put me off the project. It tells me how I would use it: as a fast router with a reasoning lane for the decisions that give it trouble.
