Back to BlogI Tested a Decision Model 486 Times, Then Built a Product on It

I Tested a Decision Model 486 Times, Then Built a Product on It

Jev Lab runs 24 everyday AI decisions three ways: Jev, an LLM prompt, and plain code. Jev got 95.7% of them right and lost 6 of the 24 projects. MKT·INTEL is the market research agent I built on it afterwards, and it still runs without Jev.

AIPythonNext.js

The question a demo can't answer

Most of what an AI product does all day isn't writing. It's deciding. Is this ticket urgent? Is this tool call safe? Is this chunk relevant? Should this answer ship? People usually hand those to the same big model that writes the prose, with a prompt that ends in "respond with one word".

Jev (typesafe/jev-1.13) takes a different approach. You give it a state and a question, and it returns a typed decision: a choice, a score or a yes/no, each with a confidence. Your code acts on that value. The pitch is that a small model built for bounded decisions beats a general model that's been asked to pretend.

A pitch can't settle that, and neither can a demo. The only way to know is to count the mistakes. So I built two things:

  • Jev Lab puts Jev, a prompted LLM and plain code on the same labelled rows and counts every error.
  • MKT·INTEL is a real product where Jev makes the calls that matter, with a fallback for when it isn't available.

Jev Lab: Source · Live

MKT·INTEL: Source · Live


Jev Lab: 24 decisions, three ways each

Each project is one decision an AI product actually makes: classify a ticket, gate a risky tool call, filter RAG chunks, verify a citation, catch a prompt injection, route to an agent, qualify a lead. Every project gets decided by:

  1. Jev, called through OpenRouter's Decisions API.
  2. An LLM baseline, the usual prompt, often a structured-output variant as well.
  3. No model at all: regex, a permission list, a points system, or "store everything".

All three variants see the same rows, and every call is a real API call on my key. The whole suite costs about $0.51 per full run.

Accuracy vs cost across all 24 experiments

Every project is a folder with an experiment.py that exposes run_experiment(input) → Result, plus a dataset, a diagram of how it's wired with and without Jev, and offline tests. core/projects.py globs for projects/NN_*/experiment.py, so adding project 25 means adding a folder and nothing else. The Next.js UI and the original Streamlit page both read that same contract.

What the numbers say

These come from the latest saved run of every project, scored as an exact label match:

Variant Correct Accuracy
Jev 465 / 486 95.7%
Prompted LLM (plain) 90 / 98 91.8%
Prompted LLM (structured output) 198 / 213 93.0%

The headline number matters less than the shape underneath it.

On small, bounded decisions, Jev was cheaper and faster. In 04 action routing it scored 19/20 against 16/20 for both LLM variants. In 14 coding tool gate it scored 23/24 against 19/24 for structured output. On ticket classification, Jev cost about $0.000017 per decision against $0.000053 for the structured LLM, and came back in 0.87 s instead of 1.4 s.

Plain code lost badly, every time a decision needed judgement. Regex caught 12/24 PII cases. Trusting citations as-is got 9/20. A static linter got 9/20. These baselines are free and instant, and they're wrong half the time. That's the real case for putting any model in the loop.

Jev lost 6 of the 24. This is the part I trust most:

Project Jev Winner
05 model router 15/20 LLM router, 18/20
17 agent escalation 18/20 LLM self-confidence, 20/20
20 fraud prescreen 17/21 escalation ladder, 19/21
24 agent harness 11/14 plain agent, 13/14
06 tool selector 19/20 native tool calling, 20/20
18 support agent 19/20 single agent, 20/20

The losses cluster. Once a decision sits inside a longer agent loop, or when "which model should handle this" depends on how hard the task is, a general model that sees the whole context did better. The agent harness is also where Jev cost the most, about $0.0023 per decision, so there it lost on accuracy and wasn't cheap either.

One caveat. The rows are mine and the datasets are small, 14 to 24 per project. Read these as directional, not as a leaderboard.

Keeping a public demo from draining my key

The deployed site calls real models with my key, so the API caps usage: 20 single runs per visitor per hour, 300 a day in total, and even fewer full dataset runs. The counters live in memory. With those limits, the worst case stays under $2 a day, and there's a credit limit on the OpenRouter key behind all of it.


MKT·INTEL: using it for real

Benchmarks tell you where a tool is good. MKT·INTEL is where I found out whether that holds up inside a product.

You ask something like "pricing teardown of Shopify vs BigCommerce" and a team of agents researches it on the live web. The result is a brief where every claim has a source.

MKT·INTEL research view

graph LR
    Q["Question"] --> P["Plan<br/>companies + domains<br/>3-6 search tasks"]
    P --> R["Research<br/>search + crawl to Markdown"]
    R --> A["Analyze<br/>cited claims, changes,<br/>contradictions, gaps"]
    A --> D["Decide<br/>Jev scores each change"]
    D -->|"needs more research?<br/>max 2 loops"| R
    D --> S["Synthesize<br/>brief"]
    S --> DB["Persist<br/>InsForge"]

It's a LangGraph state graph behind FastAPI, and every step streams to the UI over SSE. You watch the plan, each search, each crawl, each provider switch and each decision as it happens, in a terminal-style interface.

Where Jev sits

OpenAI plans, extracts and writes. Jev only decides. For every change the research turns up, it answers four typed questions:

  • Is it real? A probability.
  • Which type is it? A probability for each type.
  • How big is the impact? A score from 0 to 100.
  • How strong is the evidence?

The app maps those onto alert / investigate / monitor / ignore. A fifth question, needs more research?, can send the graph back to the research step, at most twice. The brief labels unverified changes as unverified. They never appear as fact.

That split came straight out of Jev Lab. Jev did well on bounded questions with a fixed set of answers, so that's the only job it has here.

It still works without Jev

If you don't bring an OpenRouter key, the decisions run on an LLM decision agent that uses your own key and returns the same typed shapes. The rest of the graph can't tell which one answered. I didn't want a product that stops working when one provider does.

Web data gets the same treatment. The app tries Context.dev first, then Tavily, then Firecrawl, and a credit breaker skips any provider that has run out, so a run doesn't die halfway because one account hit zero.

The rest of it

  • Eight playbooks: change brief, competitor profile, pricing teardown, battlecard, market landscape, customer pain, market sizing (the TAM/SAM/SOM maths is checked server-side) and go/no-go opportunity.
  • Every report exports as PDF, an editable PowerPoint, Excel, Markdown or JSON, and comes with a methodology section that lists per-topic completeness, gaps, contradictions and source quality.
  • Bring your own keys. OpenAI, Anthropic, Gemini or OpenRouter, with a fast model for planning and a strong one for the brief, both picked from the provider's live model list. Keys are Fernet-encrypted at rest and tried before the platform keys.
  • Keyboard first. Ctrl K opens the command palette, 1-7 navigates, j/k moves through lists.

Next.js 16 on Vercel, FastAPI in Docker on Render, InsForge for auth and Postgres with row-level security.


What I'd change

The live run streams and the usage meters in MKT·INTEL live in process memory, so the backend has to stay at one instance. Moving the run registry to Redis pub/sub fixes that, and it's first on the list. After that come scheduled change tracking with Context.dev monitors, and Slack and email alerts when a high-impact change lands.

On Jev Lab, the datasets are what I'd grow. Twenty rows is enough to catch a 50% baseline. It isn't enough to tell 95% from 90% with any confidence.

The lesson carries over from the Docker auditor. Nothing here counts as "better" until it has a column of numbers next to it. Jev Lab gave me that column before I built the product, and it told me where not to use Jev.


Jev Lab is an independent set of tests. I'm not affiliated with TypeSafe AI or Jev, and I paid for every run with my own OpenRouter key.

Related Posts

"Whatever you lose, you'll find it again. But what you throw away you'll never get back."

- Himura Kenshin, Rurouni Kenshin