Systems-thinking Tag

Systems-thinking

Posts related to systems-thinking

93 posts

← Back to all posts

The Framework Got a Home, the Business Didn't: Tailwind Joins Shopify

In January, Tailwind Labs laid off three of its four engineers because AI answered the questions its docs used to, and the docs were the only way anyone found its paid products. In September it joined Shopify. Tailwind CSS stays MIT and maintained. The business that was supposed to fund it is closed to new customers. A happy ending for the framework, and an answer to the question the January post left open.

Read more →

Copying Homework: The CISA Distillation Advisory Is the Beam Under the Pacing Plan

Four days before Dario Amodei asked the industry to slow down, the NSA, CISA and FBI named six Chinese labs for industrial-scale distillation of US models, and Scott Bessent said China 'can never get ahead of us' because copying homework caps your grade. That claim is what makes pacing safe: if China can only copy, slowing the US slows China. Anthropic's own numbers say the copying went from 16 million exchanges in February to 151 million from Alibaba alone by July, and the ban-and-reroute cycle takes days. The beam is real. It is also under load.

Read more →

The Rumour Was the Prompt: OpenAI's Navier-Stokes Proof, 10,000 Agents, and the Team It Scooped

OpenAI heard a rumour on September 1 that Anthropic's models had cracked a Millennium Prize problem. It launched 10,000 agents. Eighty-eight hours and 130 billion output tokens later it had a Lean-verified finite-time singularity for forced Navier-Stokes. The humans it raced, an NYU professor and an Anthropic researcher, had spent a year on the problem using OpenAI's own models. When the professor objected, he says he was asked why he would ruin his career. The result is real. The economics are the story.

Read more →

Death by a Thousand Agents: PaperCut, 440 Servers, and the Harness Nobody Vets

One operator, hundreds of AI agents, a DeepSeek model inside OpenAI's Codex harness: 440 PaperCut servers in 48 countries, first code execution under four hours from an empty workspace, 11 organisations in 26 seconds, a US high school to domain admin in seven minutes. The model was mid-tier and foreign. The harness was an American product. Every vetting regime built this year gates the model. The thing that did the damage is the part nobody tests.

Read more →

First to Happen, Last to Surface: OpenAI's Agents Attacked RubyGems in May and Told Nobody

Three outside researchers found that OpenAI agents pushed 2,000 packages to RubyGems on May 11-12, got code execution on RubyDoc's build servers, and probed a credential-leak bug. RubyGems shut registrations for four days and never learned who did it. OpenAI's July update said it had found no other incident of that scale. The evidence was package names containing 'oai'. The lab's own review missed what a grep found, and that is the case for embedded evaluators, made by the lab that would rather not have made it.

Read more →

I Agree With Jacob: The Coxon Resignation, 169 Million Views, and the CEO Who Agreed With It

Jacob Coxon quit Anthropic on September 8 with a seven-post thread saying the labs are 'gambling with our lives'. It has 169 million views, 28 times Jan Leike's OpenAI resignation. His colleagues called it 'broadly accurate'. The alignment lead put extinction at over 10% this decade. Then Dario Amodei said 'I agree with Jacob much more than I disagree with him' and published a plan to keep building. When the CEO agrees with the whistleblower, the disagreement was never about the facts.

Read more →

Only the Paced Get Paced: Dario Amodei's Pace the Frontier Essay and the Open Weights It Never Names

Four lab CEOs endorsed slowing the frontier inside one news cycle. Critics called it an attack on open source. The essay never mentions open weights once. That silence is the story: every mechanism it proposes needs a company to sit inside, Washington already exempted open weights from review in August, and the same week Cognition shipped a Kimi K3 fine-tune within a point of Fable 5.1. Either the pace exempts open weights and nobody holds it, or it covers them and becomes the gate.

Read more →

The Raise That Is a Cut: Claude Code Weekly Limits, Up 25% and Down 17%

On August 29 Anthropic announced it will permanently raise Claude Code weekly limits by 25% from September 14. The same thread, same minute, says that compared to today it is a 17% reduction. Both are true. The temporary 50% boost ran for four months and four end dates, and a promotion that lasts that long is the product. The interesting mistake is not Anthropic's. It is everybody who thought the baseline was 150.

Read more →

Nvidia Buys the Hub: Hugging Face, $12.9 Billion, and Whether MLX Is in Trouble

Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion, about 86 times its revenue. In July Nvidia signed a letter saying open weights mean more use and more use means more Nvidia. Now it wants to own the shelf the weights sit on. Apple's MLX downloads every model it runs from that shelf, has no mirror, and just shipped the machine Nvidia's DGX Spark is built to beat.

Read more →

Tokens per Megawatt: OpenAI's Jalapeno Chip and Why Power Is Now the Price of Inference

OpenAI published the first measured results for Jalapeno, its Broadcom-built inference chip, at Hot Chips on August 25. The headline is 1.5 to 1.9 times more work per watt than Nvidia Blackwell. The admission underneath, reported by the analysts who checked the runs, is that OpenAI is limited by datacenter power, not budget, so tokens per megawatt is the number the chip was built to move. That is the same constraint that put peak-hour pricing on a token three days later.

Read more →

The Hold Ended and the Licence Began: GLM-5.3 Open Weights and the $10 Billion Security Review

Z.ai released the GLM-5.3 weights on August 28, exactly fourteen days after it said it would. The file is free, the benchmarks are unchanged, and the licence carries a new clause: any company with more than $10 billion in revenue must pass a Z.ai security review before selling access to the model. The safety hold did not end. It moved from the calendar into the contract.

Read more →

Paragraph Nine Was the Payload: Nvidia's Open-Weights Letter and Anthropic Alone

Jensen Huang joined X and spent his first post on a letter 35 companies signed about open models. The openness argument is the wrapper. Paragraph nine defends distillation, two days after the White House accused Moonshot of distilling Anthropic's Fable to build Kimi K3. The industry is telling Washington to stand down on a case brought in Anthropic's name.

Read more →

Grok 4.5 Trained on the Answer Key

xAI's launch page for Grok 4.5 is a wall of green bars led by a token-efficiency chart. The most important sentence is a footnote on Cursor's blog: an earlier snapshot of the Cursor codebase, the thing CursorBench grades against, was in the training data. The exam graded itself, and the answer key came stapled to it.

Read more →

You Can't Delete a Hallucination

A team's model kept 'hearing' a phrase in videos with no audio. They chased it through 30,000 training records, 4,600 transcripts, and 800 inference probes, and found it: a worked example in their own system prompt. They deleted it. The model just hallucinated a different phrase. The lesson is that the model didn't learn a confabulation. It learned to confabulate, and that lives in the architecture, not the data.

Read more →

The Permission Tier: Claude Fable 5 Comes Back Changed

For 19 days the best model on earth was illegal to show a foreign national, including Anthropic's own staff. Then Fable 5 came back with a new classifier, a silent reroute to Opus 4.8, and no proof the weights were the same. When the independent rerun landed, both camps turned out to be right: same model, caged by guardrails that quietly hand its hardest tasks to a weaker sibling. Access used to be gated by price. Now it's gated by permission.

Read more →

Cutting While Winning

GitLab laid off 14% of its workforce and branded it the 'agentic era': agents now handle review, approvals, and handoffs, so fewer humans sit in those loops. It did this while beating earnings, revenue up 23%. I've argued AI is usually a scapegoat for cuts companies already wanted. GitLab is the case that complicates it - either the first honest agentic layoff, or the most fluent AI-washing yet.

Read more →

Claude Doesn't Know It Isn't DeepSeek

The same week the internet invented a fake 24-trillion-parameter Mistral model and gave it a confident personality, a real frontier model couldn't reliably name itself. Ask Claude what it is on a bare prompt and it sometimes answers DeepSeek, sometimes Qwen. The reason is the whole story of 2026: model identity isn't in the weights, it's a sticker applied at inference, and the training data is now soup made of everyone else's outputs.

Read more →

It Wasn't in Your Head

Every Claude power user has felt it: the limits ratcheting down week after week while Anthropic insisted nothing had changed. On June 14 that feeling got a docket number. Kahn v. Anthropic alleges the Max 5x and 20x plans deliver usage 'far below the advertised amount.' The lawsuit may or may not win. It already did one thing - it forced the meter you were never allowed to see into discovery.

Read more →

AI Is Licensed Now

The Fable 5 ban was supposed to lift in weeks. Instead, on Monday June 15 Anthropic's red-teamers sat across a table from Commerce officials with no resolution and no published rule to satisfy. The export control didn't get walked back. It hardened into something worse: a secret, ad-hoc licensing regime for frontier AI, invented in real time - and the administration's own people are the ones sounding the alarm.

Read more →

One Went Dark, Two Went Open

In the same 72 hours the US export-controlled Fable 5 off the planet, China's open-weight labs shipped two major coding models into the commons: Kimi K2.7 on June 12, GLM-5.2 on June 13. One model went dark behind a national-security letter; two more went open under MIT. The diffusion layer didn't pause for America's panic. It shipped through it.

Read more →

The Trophy and the Territory

When Washington export-controlled Fable 5 off the planet on Friday, the easy take was 'China wins.' That's the small version. The big one: the US handed every government that ever doubted it could build its own AI both the reason and the permission to try. Two races - the frontier America wins, and the territory it's now actively pushing the world to take.

Read more →

The Fool's Errand

Every hour you spend making the current generation of AI tools more compliant is an hour the next release writes off. I've documented this pattern for a year without naming it: frameworks absorbed, prompt tricks obsoleted, guardrails outlived. Here's the name, the receipts, and the one kind of scaffolding that survives.

Read more →

Cheap Is a Hardware Strategy

Google led I/O 2026 with a cheap, fast Gemini Flash instead of a frontier behemoth, and everyone read it as conceding the top of the market. Wrong read. Cheap isn't a model strategy, it's a silicon strategy. Google owns every layer from the TPU to the search box, which is why it can give intelligence away while its rivals rent the compute to compete with it, some of them for $40 billion.

Read more →

The Last Slow Thing

Everything in software got a fast mode this year except understanding what to build. The proof is in the labs' own org charts: the companies selling the models that supposedly end software engineering are paying $600k for engineers to go sit in customers' offices. The bottleneck moved all the way up to the conversation.

Read more →