The Thirty-Cent Judge: TypeSafe's Jev on a Real Product-Matching Queue
TypeSafe launched Jev on September 15 claiming 193x faster and 444x cheaper than frontier LLMs, zero hallucination, and calibrated probabilities. The launch evals measure agreement with GPT-6 and Fable 5.1, not correctness, and no calibration curve has been published. I had a better test in Pricogni: 9,081 low-confidence product matches a human review queue was never going to clear. One Noul, one Choice, 150 lines, 32 cents, 13 minutes. Half the queue was flankers, a fifth was publishable, and the one time it disagreed with a human reviewer the model was right.
READ




