Architecture Tag

Architecture

Posts related to architecture

44 posts

← Back to all posts

The Thirty-Cent Judge: TypeSafe's Jev on a Real Product-Matching Queue

TypeSafe launched Jev on September 15 claiming 193x faster and 444x cheaper than frontier LLMs, zero hallucination, and calibrated probabilities. The launch evals measure agreement with GPT-6 and Fable 5.1, not correctness, and no calibration curve has been published. I had a better test in Pricogni: 9,081 low-confidence product matches a human review queue was never going to clear. One Noul, one Choice, 150 lines, 32 cents, 13 minutes. Half the queue was flankers, a fifth was publishable, and the one time it disagreed with a human reviewer the model was right.

Read more →

Nvidia Buys the Hub: Hugging Face, $12.9 Billion, and Whether MLX Is in Trouble

Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion, about 86 times its revenue. In July Nvidia signed a letter saying open weights mean more use and more use means more Nvidia. Now it wants to own the shelf the weights sit on. Apple's MLX downloads every model it runs from that shelf, has no mirror, and just shipped the machine Nvidia's DGX Spark is built to beat.

Read more →

Tokens per Megawatt: OpenAI's Jalapeno Chip and Why Power Is Now the Price of Inference

OpenAI published the first measured results for Jalapeno, its Broadcom-built inference chip, at Hot Chips on August 25. The headline is 1.5 to 1.9 times more work per watt than Nvidia Blackwell. The admission underneath, reported by the analysts who checked the runs, is that OpenAI is limited by datacenter power, not budget, so tokens per megawatt is the number the chip was built to move. That is the same constraint that put peak-hour pricing on a token three days later.

Read more →

The Editor Is Now a Host

Cognition killed Windsurf overnight via an over-the-air update, rebranded it Devin Desktop, made the default UI an agent command center instead of a code editor, and shipped an open Agent Client Protocol so Codex, Claude, and OpenCode can all run inside it. The bet underneath: the IDE wins by being the place agents report for work, not by having the best autocomplete. The editor was always the wrong center of gravity.

Read more →

Cheap Is a Hardware Strategy

Google led I/O 2026 with a cheap, fast Gemini Flash instead of a frontier behemoth, and everyone read it as conceding the top of the market. Wrong read. Cheap isn't a model strategy, it's a silicon strategy. Google owns every layer from the TPU to the search box, which is why it can give intelligence away while its rivals rent the compute to compete with it, some of them for $40 billion.

Read more →