Skip to content
Thrivbe
← All experiments
In progressStarted September 20, 2026Updated September 24, 2026

Can a fast "System 1" model replace an LLM call for routine yes/no and ranking judgments, and what does it take to trust it?

Jev as a Judgment Primitive

Partly. Jev answers in under a second for a fraction of a cent and is stable, but not automatically right: a coin flip on bug-versus-noise triage, a higher score than the old scorer on job listings (on labels that scorer had already filtered), and far ahead of a local model at screening 82 applicants. Trust came from a baseline, data-derived thresholds, hard rules in code, and shadow mode first.

  • Python
  • TypeSafe Jev (System One)
  • Claude Sonnet (comparison baseline)
  • laya (open-source local model, comparison)
  • Claude Code hooks

Relevant services: Thrivbe AI

Hypothesis

Many decisions in our systems are small semantic judgments: is this relevant, is this urgent, which of these options fits. An LLM call can make them, but at several seconds and real cost per call we only make them where it is unavoidable, and everywhere else we fall back to keyword rules that miss meaning. Jev, from TypeSafe, is pitched as a "smart if statement": a fast System 1 model that scores answers in parallel instead of writing tokens. If it is fast and cheap enough, semantic decisions can live in places an LLM never could. The open question is whether its answers are right, and what evidence we need before we let it decide anything.

What we built

Jev takes a piece of state plus named questions of three kinds: noul (yes or no, returns a probability), choice (one label from a list, with a probability map and confidence) and score (a grade on a described scale). Our pattern is the same in every test: Jev judges, then a deterministic threshold in our own code decides. Guidance we settled on: if a regex or rule can do it, use code; if it needs generation or a summary, use a normal LLM; if it is an intuitive structured judgment, try Jev.

Tests, in the order we ran them:

  • Benchmark against Sonnet (20 September). Two small tests with weak or no ground truth. On relevance filtering of 30 articles, Jev gave the same answer on 30 of 30 repeated runs and took about 0.9 s per call against 5.3 s for Sonnet. The two agreed on 24 of 30; nobody adjudicated the six disagreements, so accuracy there is unknown. On 30 error-feed issues labelled from past triage comments (title-only input, rough labels), both got 14 of 30 right; AUC was 0.49 for Jev (a coin flip) and 0.62 for Sonnet.
  • A tradeability gate with a keyword baseline. In a research pipeline we used Jev to keep off-topic methods papers out of a hypothesis pool, and ran the fair-baseline step: a keyword rule agreed with Jev on 95 of 100 hypotheses. Of the 5 disagreements, Jev was right on 4. The keyword rule failed on ordinary word collisions and on one legitimate trade it wrongly rejected.
  • Contact scoring. One run over 1,946 contacts took 143 s, cost $0.083 and had 0 failures. We have no accuracy number for it.
  • A log-only news-label job. Jev labels five crypto news feeds next to a keyword baseline every 15 minutes since 21 September. Nothing reads the output; a kill-or-promote review is scheduled for 19 October.
  • Job-search back-test (20 September). An offline test on a personal job-search pipeline with 128 human approve or reject labels (26 approved, 102 rejected). Jev's probability reached AUC 0.78 (95% range 0.67 to 0.88) against 0.63 for the pipeline's current score. A 4-level grade scored 0.79. Combining Jev with the old score was worse (0.74) than Jev alone. As a pre-filter, dropping everything below p=0.1 lost 0 of 26 approved jobs and would skip about 73% of random jobs, but the existing low-score cut also lost none, so the gain is saved scoring calls, not better filtering. 278 calls, 0 errors, 0.73 s median. Jev also ignored an explicit location rule (a wrong-location job scored 0.76), so location stays in code. It then went live in shadow mode on 21 September: answers stored beside the old score, nothing changed.
  • Jev versus a local open model (21 and 22 September). On a real screening task with 82 freelance applicants, scored against a careful three-agent Sonnet read of every application, Jev detected templated applications at AUC 0.84 and its fit score correlated with the reference at r=0.66. It flagged 3 of 78 as off-platform contact and caught the one real case at 0.96. The open-source laya base checkpoint got AUC 0.51 (r=-0.02) and 33 of 78 false off-platform flags; its "typed decisions" checkpoint got AUC 0.58, cut false flags to 6 of 78 and then missed the one real case. Latency was 740 ms per call for Jev against 170 ms for laya (free, local).
  • Inbox triage (20 September). A read-only run that ranked 200 client inbox threads into reply-now, reply-soon, pitch, review and archive buckets took about 7 s of Jev time and $0.004. Nothing sends. A second pass asked whether a client knowledge base (422 markdown files) already answered each of the 12 messages needing a reply: 0 of 12 were answered (5 had no matching content, 7 were not lookups at all, such as thank-yous). A planted question whose answer was written down scored 0.41 and came back "partial", against 0.03 for the others, so the check does detect an answer when one exists.
  • Chat-claim calibration (24 September). A first-pass claim checker for community chat got 4 of 6 claims with known answers right.
  • Two Claude Code hooks, both shipped switched off. A model router sizes each typed message as tiny, everyday, large or hardest and suggests which model tier to delegate to. A skill picker chooses one of roughly 327 skills for a request (groups of 40 plus a runoff, acting only above 60% confidence). Both fail open on any error or timeout. We have no accuracy number for either; we tested them on a handful of requests and have not measured them.
  • A "needs attention now" scorer (24 September). Runs inside the attention router every 15 minutes and orders approvals and notifications. It records would-page and would-suppress verdicts but does not change paging. Only seven hard-coded system sources are sent, redacted, and approvals send only an action type and age.
  • A maze demo. A 17x13 grid where Jev picks up, down, left or right each tick. It is a toy, and it taught the composition lesson (see Learnings).

Learnings

  • Fast and stable is not the same as right. Jev gave the same answer on 30 of 30 repeated runs and still scored 0.49 AUC on the error-feed task. Do not use it for triage of real bugs or any hard-to-undo action based on that test alone.
  • Compare against a real baseline, not against nothing. The keyword-baseline step on the tradeability gate is the only test where we can say Jev beat the thing it would replace, and the margin was 4 of 5 disagreements.
  • Threshold matters more than the model. A guessed floor of 1.5 was discarding a legitimate study at 0.91. Looking at the real score distribution put the floor at 0.75. Derive cutoffs from the distribution, never from a handful of examples.
  • It is only as good as the labels, and ours were uneven. The job-search back-test labels exist only for jobs the old scorer had already ranked high, which makes the AUC comparison unfair to the old scorer. Twenty-six approvals is a small sample; zero lost of 26 still permits a true loss rate near 11%. One prompt, no cross-validation. A reason to run shadow mode, not a proof.
  • Nuance is where the small local model fell over. The task needed reading between the lines of cover letters. A smaller open model produced noise on it, while Jev tracked a careful human-style read. Local models may still suit simple surface-level routing; we have not tested that here.
  • Retrieval was the hard part, not the judgment. In the inbox work Jev's scores separated cleanly (0.41 against 0.03) at every stage, and three retrieval fixes made the difference: strip URLs from keywords, weight by rarity and divide by document length, and send whole pages instead of short windows around the first keyword hit. The knowledge base also simply did not hold answers to outsiders' mail.
  • It matches topic, not claim. On the chat-claim check it got two of six wrong, including one claim it called supported that was false. Treat it as a first-pass sorter and read the rows it surfaces yourself.
  • Keep hard rules in code. Jev ignored a location rule in the job test, and in the maze demo it oscillated in dead ends until code pruned illegal and recently visited moves before offering options. With that composition it solved 4 of 6 maps (shortest paths 13 to 24 steps, solved in 14 to 39) and failed the two that need 15 or more consecutive moves away from the goal. Jev made every decision; code pruned the options.
  • Meta questions fail, content questions work. Asked whether loading a skill would help, Jev gave 0.18 to a request that clearly named a platform task. Reworded to ask about the message's content, the same request scored 0.95. It cannot reason about our own tooling. A related gap in the router: a short follow-up like "No, the other one." was rated tiny at 0.84, so confidence alone does not catch conversation-dependent replies.
  • Notification noise is a source problem. On the scorer, Jev could not separate routine auto-heal give-ups from real failures (all about 0.60). Of 140 scored notifications it surfaced 1. The value is in ordering approvals, from a weak signal of action type and age, and the better fix is muting the noisy source.
  • Default-deny data. An adversarial review found redaction alone lets some names and amounts through, so an allowlist of sources is the real guard.
  • Expect overload. The API intermittently returned HTTP 529. Every caller needs retry with backoff; a single try fails silently.
  • Vendor numbers are unverified. The stated 70 to 500 ms latency and pricing are claims; we measured roughly 0.7 to 0.9 s per call from a laptop.

Open: nothing has yet run unattended long enough to prove it. Shadow reports for the job-search pipeline and the scorer, and the news-label review, are still ahead.

Log

  • 2026-09-24: "Needs attention now" scorer went live in shadow mode inside the attention router: scores written, approvals sorted by it, paging untouched. Chat-claim calibration run: 4 of 6 right.
  • 2026-09-22: Model router and skill picker hooks built, both off by default. Jev versus local model write-up finished on the 82-applicant task.
  • 2026-09-21: Job-search pipeline shadow mode went live. Contact scoring, the tradeability gate and the log-only news-label job shipped, and the keyword-baseline check ran. Maze demo reached its v4 state and was parked.
  • 2026-09-20: Started. Benchmark against Sonnet on two small tasks, job search back-test (128 labels), and read-only inbox triage run.