Most AI models write. Jev doesn’t. TypeSafe launched it on September 15 with $40M behind it, and it only answers questions:

  • Yes or no, as a probability. “Are these the same product?” 0.94.
  • A pick from a list you give it.
  • A score on a scale you define.

Answers come back in under half a second. Input costs $0.042 per million tokens. Output is free, because there is almost none.

“choice” maps to “match” statement, “score” maps to sorting, “noul” short for bernoulli maps to if-statements

— Diogo Almeida, TypeSafe CEO, on Hacker News

In other words: an if-statement that understands English.

The Launch Claims Are Soft

  • “444x cheaper, 193x faster.” TypeSafe measured how often Jev agreed with GPT-6 Astra and Fable 5.1, not how often it was right.
  • “Zero hallucination.” It means the answer always comes back in the right shape. One Hacker News comment put it best: “it can’t emit an invalid type, but it can still emit a completely wrong valid value.”
  • “Calibrated.” This is the claim that matters. When Jev says 90% sure, is it right 90% of the time? TypeSafe published nothing to show it.

So I tested it on my own data.

The Queue Nobody Was Going to Clear

Pricogni, my competitor pricing product, matches a retailer’s products against competitors’ listings. Is “Brand X Body Wash 1L Lavender” the same product as a competitor’s “Brand X Sensitive Body Wash 1L”? Code handles the easy cases: same barcode, or same brand, size and name. The unsure ones go to a human.

There were 9,081 unsure ones. Nobody was ever going to review them. The plan from June was an AI judge, and it was shelved because a frontier model cost too much per check.

Jev answers two questions per pair: are these the same product, and if not, what’s different? The build took an evening.

same: {
  type: 'noul',
  instructions:
    'Are `catalogue` and `competitor` listings for the same retail product: ' +
    'same brand line, same variant (shade, flavour, scent, strength, form) ' +
    'and same pack size?',
},

32 Cents

Pairs judged9,081
Cost$0.32
Time13 minutes
Rejected49%
Confirmed21%
Still unsure30%

Code turns the probability into a decision: 0.8 or higher confirms, 0.2 or lower rejects, anything in between stays with a human.

Most rejects were flankers: same brand, same size, different product. Shampoo against the conditioner from the same range. A knee support against the ankle support. My matcher saw “same brand, same bottle” and guessed. Jev read the names.

I checked 50 decisions by hand. 48 held up.

The one time it overruled a human

One pair had a human decision, and Jev disagreed with it. A reviewer had matched a roll to the same brand’s roll at double the width. Jev said 6% likely to be the same, and called it a different product. The reviewer was wrong.

Update, September 30: It Got the Job

Eight days later, Jev stopped advising and started deciding. Its verdict now sets which matches Pricogni publishes.

Fifty hand-checked rows weren’t enough to justify that. Barcodes were. I hid the barcodes from 1,450 pairs, let Jev judge them, then used the barcodes as the answer key. Jev’s share of that test cost 5 cents.

  • More accurate. Published matches went from 78% right to 90% right.
  • More matches found. It caught 70% of the real matches, up from 62%.
  • Wrong with confidence. Moving the cutoff from 0.7 to 0.9 barely changed accuracy. When Jev is wrong, it is usually sure of itself. That answers the calibration question, and not in Jev’s favour.
  • Blind spot. On pairs where brand, pack and size already agreed, Jev rejected 9 of 250, and all 9 were real matches. Since September 30, a second model double-checks those rejections before they stand.

The unsure pairs no longer wait for a human either. They go to Gemini 3.8 Flash with web search, which settled 208 of 216 in testing, at about $0.002 each. That’s roughly 55 times Jev’s price, but it only sees the pairs Jev can’t decide.

In production, one Amazon run cut the unsure pile from 2,661 to 82. A weekly check hides the barcodes on 400 products and re-matches them. The first run scored 90.3%. An alert fires below 85%.

What Jev Can’t Do

  • Maths, dates, images. It reads names. Size and pack contradictions stay in code.
  • Ignore instructions in the data. Competitor descriptions are scraped text, and TypeSafe’s own limitations page says text written to steer it can move the answer. Fine for product matching. Not for anything touching money or access.
  • Prove it’s new. Sean Goedecke argues you can get most of this from a regular model forced to answer in one token. He may be right. For me it doesn’t matter. The price is what changed.

The Takeaway

  • The price changed the design. “Check every uncertain match with an AI” was already a published pattern. It was too expensive to run until it cost 32 cents.
  • Free output is the real pricing story. A judge returns a few numbers. Paying writing rates for a few numbers made every AI-judge idea a budget line.
  • Measure before you hand over the keys. The barcodes made measuring fast. Once the numbers were in, handing over the decision was easy.