Agentic AI's 2FPS Moment
Audience: Engineers watching open models converge with hardware they can own.
Reading time: ~15 minutes.
In 1998 I was a young software engineer at Microsoft, writing DirectX code on a monstrously expensive 3D accelerator that rendered simple scenes at 2 frames per second. It felt like bullshit, and I was the one writing it. Then it shipped. Consumer cards ran the same code at 30 FPS for $125, and the accidental Moore’s Law that gaming volume financed kept compounding from there: fixed-function pipelines became programmable shaders, researchers hijacked the shaders for general math, Nvidia formalized the hijack as CUDA, and CUDA caught crypto, then AlexNet, then the training runs behind the LLM I now type wishes into. The entire AI era sits on silicon that exists because tens of millions of people wanted 3D games on Windows to look better.
Everyone running a 2-bit quant of a 753-billion-parameter model at 3 to 9 tokens per second is having the 1998 feeling. 3 to 9 tok/s on a 2-bit quant is this era’s 2 FPS. The feeling of absurdity is what building for hardware that does not exist yet feels like from the inside.
A while back, in LLM Inferencing Costs are Going to $0, I argued that the marginal price of everyday intelligence was headed to zero and that productized local-model solutions would show up soon. That post was about price. This one is about ownership, because sometime this summer we crossed a border that price doesn’t describe: frontier-class AI stopped being only a thing you rent and became a thing you can own. Not comfortably, not cheaply, and at 2 FPS. But the border is behind us, and everything interesting follows from that.
The Border Is Behind Us
My capability threshold is frozen on purpose: “Opus 4.6-class,” meaning the model Anthropic shipped in February 2026. I picked it because 4.6 is the oldest model that still feels sufficient to me for serious agentic work, and I keep meeting engineers who say the same. The frontier has moved well past it since (Fable 5 arrived in June) and I use the new stuff daily. That’s a fact about the frontier. It doesn’t move the threshold. When Fable 5 shipped I ran a controlled model swap on an unchanged harness to check whether it should: First Fable. A step up, not a step function.
In June, Zhipu’s GLM-5.2 reached that threshold in the open. The facts, from the model card: 753 billion parameters, about 40 billion active per token, plain MIT license, a million tokens of context. On Zhipu’s own benchmark table it clears Opus 4.6 on the hard coding evals and closes in on 4.7. Vendor numbers, so discount accordingly. The independent datapoint is better: the UK’s AI Security Institute measured GLM-5.2 performing about level with Opus 4.6 on their cyber evaluations, and open-weight models as a class trailing the closed frontier by four to seven months, down from six to ten a year ago. A government security lab now measures the gap between the frontier and the commons in months.
Two things about GLM-5.2 matter more than the benchmarks. It ran inside Claude Code, Cline, Roo, and Goose on day one, because the ecosystem standardized on frontier-shaped APIs; the models are becoming drop-in parts. And the 4-bit build, roughly 400GB, fits a 512GB Mac Studio, where a community benchmark measured 17.7 tokens per second. The 2-bit build squeezes into a 256GB machine at 3 to 9 tok/s with visible quality loss. That is the 2 FPS experience, and it’s real, replicated, and MIT-licensed.
Then the cadence got silly. On July 16, Moonshot shipped Kimi K3: 2.8 trillion parameters, 16 of 896 experts active per token, trained quantization-aware at 4-bit so its native form is already 1.4TB. Moonshot claims the thing open models have always failed at, long-horizon agentic work: on SWE-Marathon, a multi-day agentic benchmark, its table shows K3 edging out Opus 4.8 and comfortably clear of Fable 5. Hold that one loosely. The claim comes from Moonshot’s own harness while the rivals ran on others’, and the same frontier model Moonshot credits with a near-tie scores far lower on the benchmark’s official leaderboard. The mismatch tells you as much about the state of benchmarking as about the model. Moonshot’s own launch copy concedes K3 “still trails” Fable 5 and GPT-5.6 Sol. The weights are promised for July 27, five days from this writing, expected under a Modified-MIT license (K2’s version added an attribution clause that only binds very large commercial deployments). On time, at full precision, with nothing worse than an attribution clause, and the cadence holds.
Three days after K3, Alibaba previewed Qwen3.8-Max at 2.4 trillion parameters, claiming second place behind only Fable 5, open weights promised. DeepSeek’s V4 line shipped in April under MIT at 1.6 trillion. Tencent’s latest is Apache-licensed. Three Chinese models over two trillion parameters, all open or promised open, inside five weeks.
Whose Brain Is It
It’s worth being precise about who “they” are, because the shorthand I keep hearing, that the Chinese government gives away frontier models, is wrong in an interesting way.
The labs are private companies having extremely capitalist years. Zhipu listed in Hong Kong in January and the stock is up roughly 2,400 percent since. Moonshot raised $2 billion at a $20 billion valuation in May and is reportedly negotiating its next round at up to $50 billion ahead of an IPO, on annualized revenue that tripled to $300 million between March and June. These are venture-backed firms making the classic second-place move: open the weights, commoditize the leader’s margin, recruit the commons as free distribution and free QA. Nobody open-weights a decisive lead.
What the state supplies is everything around the labs. The State Council’s “AI+” plan makes open-source AI an explicit national directive, with adoption targets written down like grain quotas: 70 percent of key sectors by 2027, 90 by 2030. Seventeen-plus city governments hand out compute vouchers worth up to $280,000. Beijing subsidizes API access. Xi endorsed open-source diffusion by name at the World AI Conference this month. The government doesn’t write the models. It pays for the gym, the coaches, and the plane tickets, and it has made giving the results away a matter of national strategy. Meanwhile there is no US-origin frontier-parity open-weight model. None. The export-control irony writes itself: Washington built a dam and got a floodgate.
The uncomfortable part of depending on this pipeline is that it has an off switch, and the off switch is in Beijing. Open-model progress looks like physics, like heat finding its level. It is a strategy, and strategies get discontinued: genuine parity would end the cadence from one side, competitive collapse from the other, and an IPO-minded CFO could end it from inside. This month added a sharper wrinkle. China’s commerce ministry is reportedly consulting on export controls for model weights themselves, up to and including a ban on publicly releasing the most capable ones. Read that again. The most concrete motion toward weight export controls anywhere in the world right now is Beijing considering whether to stop giving the brains away. July 27 is the live test.
The Machines Went Backward
Here’s where the story stops rhyming with 1998 and starts running in reverse.
In March 2025 Apple shipped the 512GB M3 Ultra Mac Studio, the machine that made all of this locally hostable in the first place. In March 2026 the 512GB option quietly disappeared from the store; no announcement, just a missing configurator row. On May 5 the 256GB option followed. The biggest Mac you can order new today has 96GB of unified memory. The window in which you could buy a new Mac that fits GLM-5.2 lasted one year, and it closed three months before GLM-5.2 shipped.
Look at the orange line. It drops.
The used market did what markets do. Secondhand 512GB M3 Ultras now list around $24,000 to $28,000 against a $9,499 original price. Asking prices, not sold prices, but the direction is unambiguous. This is a consumer computer appreciating like a crypto-era GPU, because it turned out to be the last of its kind.
The reason is the memory supercycle, and the numbers are violent. Conventional DRAM contract prices rose 18 to 23 percent in Q4 2025, then a record 90 to 95 percent in Q1 2026; PC DRAM alone crossed 100 percent in a single quarter. LPDDR5X, the class Apple builds unified memory from, rose roughly another 80 percent in Q2 (the 12GB-module print went from $77 to $146). The Q3 forecast, 13 to 18 percent, is what passes for cooling. Gartner’s full-year view: memory and SSD prices up 130 percent through 2026, PC prices up 17, PC shipments down 10.4 percent, the steepest contraction in a decade.
The mechanism is allocation, not appetite. A gigabyte of HBM consumes three to four gigabytes of commodity DRAM wafer capacity, AI will eat about a fifth of global DRAM wafer output this year, and the fabs, led by Samsung’s HBM4 line conversions, chose the datacenter. The consequences have names. Micron shut down Crucial, its consumer brand, to concentrate on AI customers. Nvidia raised the DGX Spark from $3,999 to $4,699 mid-cycle, citing memory costs; that works out to about $37 per gigabyte. Even the ~$97K DGX Station ships with seven of its eight HBM stacks enabled, salvaged silicon at six figures, because right now not even Nvidia can get enough memory for Nvidia.
Two numbers decide whether a model runs on your desk: whether the weights fit, and how fast you can read them. Capacity and bandwidth. Generation speed is roughly bandwidth divided by the bytes you touch per token, times an efficiency factor of one-half to three-quarters. Mixture-of-experts is what makes desk-scale frontier models possible at all: GLM-5.2 reads only its ~40B active parameters each token, about 20GB at 4-bit, so the M3 Ultra’s 819GB/s of memory bandwidth yields high-teens tok/s instead of low single digits. The whole hardware problem in one sentence: local AI is a memory product, and memory is what the datacenter is currently strip-mining.
Which is the inversion. In 1999, consumer volume financed the silicon, and the hardware rose to meet the software. In 2026, the datacenter outbids you for the wafer, so the software has to shrink to meet the machines, and remarkably the models are volunteering: K3 was trained quantization-aware precisely so that its ship-quality form is the small one. As for when the squeeze ends, pick your pessimist. Morgan Stanley sees prices peaking around Q4 2026 and declining from late 2027. Intel’s CEO relays what suppliers told him: “no relief until 2028.” SK Hynix’s CEO, on the day his company listed on Nasdaq, said 2027 will be the worst supply year in the industry’s history and demand will outrun supply beyond 2030. He would say that. He also might be right.
The Wave Comes Anyway
Against that backdrop the roadmaps are almost comically aggressive. Most of what follows is announced or reported rather than shipped, and I’ve flagged the rumor-grade parts.
Apple, per Gurman’s July reporting (single-source, rumor-grade): a Mac Studio refresh around October with an M5 Ultra tested for up to 768GB, supply permitting; the M6 generation skips its Pro, Max, and Ultra tiers entirely; a base M7 in the first half of 2027, fabbed by Intel on 18A-P of all things; and in 2028 an M7 Ultra designed for up to 1.5 terabytes of unified memory with Blackwell-class ambitions, explicitly conditional on the shortage easing. Note the shape of that bet. K3’s native form is 1.4TB. The open frontier of this July already fills the machine Apple may ship two years from now.
Nvidia put its plan on a Computex slide. The RTX Spark superchip, co-announced with Microsoft as the platform for an “agentic AI” Windows: 128GB of unified memory and about a petaflop, in fall-2026 laptops and desktops from every major OEM, at an estimated $1,800 to $2,900. (Nvidia claims it runs 120-billion-parameter models on device; vendor claim, hardware not shipped yet.) Behind it, on a public roadmap: Vera Rubin Spark in 2027-28 on LPDDR6, which roughly doubles memory bandwidth, then Rosa Feynman Spark in 2029-30. The entry price of a 100B-class local box roughly halved in one year, in the middle of a memory supercycle.
AMD sells a $2,000 box with 128GB that generates 55 tok/s on gpt-oss-120b, and now a $3,999 first-party developer machine marketed, in AMD’s own words, for the “agent computing” era, which is the closest any silicon vendor has come to naming “home LLM server” as a product category. Strix Halo’s rumored 2027 successor moves to LPDDR6 at 460 to 690GB/s. And the cluster people are having fun: macOS quietly shipped RDMA over Thunderbolt 5, EXO pools four Mac Studios into 1.5TB of unified memory for just under $40K, and four clustered $2K Framework boards run DeepSeek R1 671B at 24 tok/s. A dual-EPYC build with 768GB lands between $6K and $14K, if you can stomach 6 to 8 tok/s and some assembly.
What still doesn’t exist is the obvious machine: one box, $8K to $15K, 512GB to a terabyte. I went looking; nobody ships it. That empty middle is where this whole thesis lives, and whoever fills it first, Apple with a 768GB M5 Ultra, Nvidia with a fat Rubin Spark, AMD with a Medusa Halo, gets to sell the $125 card of this cycle.
There’s a weirder wave behind that one: silicon that gives up generality entirely. Reports surfaced this week (The Information, via Tom’s Hardware; Google won’t confirm) of “Frozen v2,” a Google server chip that etches part of Gemini’s architecture directly into the silicon. Engineers on the project reportedly expect six to ten times the tokens per watt of the newest TPUs, targeting 2028: weights updatable, architecture frozen in transistors. Etched came out of stealth in June with a transformer-only ASIC and a billion dollars in contracts; Nvidia paid a reported $20 billion to license Groq’s inference technology. This is the opening chain running in reverse: graphics silicon spent twenty years going from fixed function to programmable, and inference silicon is hardening back into fixed function, because when one workload eats the world you stop paying the generality tax. None of this is a consumer product yet. All of it is where 120 FPS comes from.
So here’s the frame-rate map I’d actually defend. 2 FPS is now: 3 to 9 tok/s on quantized giants, on a discontinued $25K machine or an $8K cluster with sharp edges. 30 FPS is 2027-28: a buyable half-terabyte box, ship-quality 4-bit models, 20 to 50 tok/s, overnight agent fleets as the normal way software gets built. Gated on the DRAM cycle, not on model quality. 120 FPS is roughly 2029-30, and I label it extrapolation: dedicated inference silicon at consumer prices, hundreds of tokens per second, resident models in everything. The dates wobble. The direction hasn’t wobbled once since llama.cpp showed up.
Nobody Buys the Box to Save Money
An honest aside before the fun part: the naive pitch for local AI, that it saves you money on tokens, is dead on arrival. The subscription frontier is heavily subsidized: running a $200/month Claude Max plan flat out would cost up to ~$3,650 a month at API rates by third-party estimates, and open-model coding plans start around $12.60 a month. Against $200 a month, a $25K used Mac Studio pays back in about a decade. Nobody buys the box to arbitrage tokens.
The box sells the things subscriptions meter. Parallelism: the weights are paid for once, and the fifth concurrent agent costs KV-cache and shared bandwidth. Privacy: the repo, or the conversation, never leaves the building, which in regulated shops is a precondition rather than a preference. And throughput over latency: 17 tok/s is painful at 2 p.m. and irrelevant at 2 a.m., and every increment of agent autonomy converts more work from the first kind into the second. The constraint that survives the conversion is whether you can trust what the fleet did while you slept, which is why I keep saying verification is the ceiling. The cloud is a latency machine. The box is a throughput machine. The router that decides, per task, which side of that seam each job runs on is the most interesting piece of software in this picture; I’ve written about where that goes in Agent-Hypervisors. (Full disclosure: you are reading agent-hypervisor.ai. The thesis found its bag.)
The durable economics are componentization. DirectX’s legacy wasn’t cheaper rendering; it made 3D a free part in every consumer box, and industries condensed out of that vapor. Nobody prices a game per triangle. When an Opus-class brain is a free part, you get the products per-token pricing forbids: firmware that maintains itself, appliances with a staff engineer inside, software that ships with its own developer. The floor is a new place products come from.
The Min-Spec Clock
How will we know the flip is happening? Not by a substitution event. Cheap bandwidth didn’t kill the telcos on a date; one day WhatsApp was simply the assumption, and the industry repriced around it. Hardware transform-and-lighting did the same to software rendering: no crossover ceremony, just a quarter after which new games assumed the GPU. The tell is system requirements, and the min-spec era has already started at the small end. Copilot+ PCs require a 40+ TOPS NPU and 16GB. Apple Intelligence has a hard 8GB memory floor; the iPhone 15 was excluded over DRAM, not over its neural engine. Apple’s Foundation Models framework now hands every app a free on-device model, no API key, no network. The remaining milestone is the first mainstream agent product whose requirements read “64GB unified memory minimum.” It hasn’t happened yet. When you see it, the flip already happened.
The arithmetic under the timing is two numbers. Call L the lag, in years, from a frontier release to the same capability running at home. The parity half of L has nearly vanished; a government lab now clocks it at four to seven months, so hardware fit dominates, and L sits near a year and falling. Call A the age of the oldest model that still feels sufficient. Today that’s Opus 4.6, so A grows a year per year for as long as the testimony holds. The crossover is L ≤ A: the day the oldest model that still feels magical is old enough to run at home. The band lands in 2027-28 without heroic assumptions. Two asterisks, both honest. Frozen weights age: a 4.6-class model in 2030 is a brilliant engineer four years into a coma, which quietly makes the MIT license load-bearing, because refresh and fine-tune rights are what keep the coma reversible. And expectations inflate: the same engineer who calls 4.6 sufficient forever is running five parallel Fable windows today. If “sufficient” inflates faster than the lag falls, the crossover recedes like a horizon. My bet is that it won’t, and I flag it as the bet it is.
Regulation Gets a Vote, Not a Veto
Now the section everyone asks about. The facts first, because they surprised me.
The only US export control that ever touched model weights (it covered closed weights only) was rescinded in May 2025; no replacement has appeared. The White House’s March 2026 AI framework spends its four pages on preempting state laws and limiting developer liability, and says nothing about open weights. The live bills in Congress are procurement bans on Chinese-origin models for federal agencies; none has passed. The EU AI Act’s general-purpose-AI obligations become enforceable on August 2, and they bind providers, the companies placing models on the market. Article 2(10) of the same law states that it does not apply to natural persons using AI in a purely personal, non-professional activity. The strictest AI law on Earth stops, deliberately and in writing, at the household door.
We’ve run this movie before. In the nineties the US government classified strong cryptography as a munition, and it lost, not by repeal but by obsolescence: the Ninth Circuit held in Bernstein that source code is protected speech (the opinion was later withdrawn for rehearing, but the writing was on the wall), Zimmermann’s printed PGP source book made the paper-versus-electrons distinction untenable, and in January 2000 the controls effectively ended. Weights that have touched a torrent are just as permanent as source code in a bookstore. What regulation can actually grip is atoms and use: chips at borders, liability after harm, procurement lists, insurance. So it shears rather than stops: audit and attestation regimes push enterprise work toward the metered cloud, where an API call is a logged, insurable event and an overnight local swarm is a compliance void, while sovereignty and residency rules push the same work home. Regulation moves the boundary between the local floor and the frontier spike. It doesn’t un-exist the floor.
And this summer supplied a nineteen-day case study. Per reporting around the launch, a Commerce Department export-control order took Fable 5, the most capable model in the world, offline on June 12; it came back July 1. I have no view to offer here on the order’s merits. I just want you to notice what the incident demonstrated. The frontier has an off switch, the off switch sits in a government office, and it works. Every downloaded copy of GLM-5.2 kept running.
The Quake Question
The 1998 story has a hole in it, and the hole is the point. Consumer 3D had Quake: the thing millions of people wanted badly enough to finance an accidental Moore’s Law. Nobody bought a Voodoo card for the DirectX dev kit. So what’s the Quake of local AI? It isn’t coding agents. We’re the DirectX engineers in this story, grinding away at 2 FPS and muttering that it feels like bullshit.
Look at what actually needs the box, and notice that every item needs it because some form of control can’t follow it there.
The robot can’t send its eyes over Wi-Fi. Figure’s Helix runs entirely onboard, a 7B vision-language model and an 80M control policy on two embedded GPUs; 1X’s home humanoid runs its model on the robot for reliability and privacy. Today’s robot brains are small, which means robots currently pull edge silicon rather than half-terabyte boxes. But cognition at reflex distance can’t ride a WAN, and that never changes. Latency: cut.
The always-on model can’t ride a meter. Apple already gives every app a free resident model, and the ambient endpoint of that trend, a model that watches and listens to everything, always, is unpriceable per token. Nobody pays a subscription denominated in their own attention. That demand only clears on owned silicon. The meter: cut.
Then the things people will never say to a server. I was going to be coy here, but the data is blunter than I would have dared to be. OpenRouter, which routes model traffic for thousands of apps, published an empirical study of a hundred trillion tokens: roleplay is the single largest category of usage, about a third of everything they carry, and roughly half of all tokens routed to open-weight models. One router’s traffic, not the world’s, but it’s the largest usage dataset anyone has published. The companion, porn included, isn’t a lurid footnote to the local-AI story. It’s the measured majority of open-model demand. Anyone surprised hasn’t read the history of media: porn and gambling were the first paying customers of the VCR, the early internet, and streaming video. Vice funds the buildout, because its customers pay first and complain least, and then everyone else moves into the infrastructure vice financed. Observation: cut.
The fourth category is everything the lab would have said no to, and here’s the paragraph I owe you. The property that makes a local model censorship-proof for a dissident is the same property that makes it safeguard-proof for an abuser. Refusals are a deployment feature; deployments are now yours; stripping the safety training out of an open model is a fine-tune, not a research program. The UK’s AI Security Institute calls open-weight release “a persistent and irreversible risk of misuse,” and they’re not wrong. The sentence was written as a warning and reads as a spec sheet. There is no version of this technology where you get the dissident’s freedom without the abuser’s, and anyone selling you one is selling something else. Permission: cut.
Four categories, four cut leashes. That’s the answer to the Quake question, and it’s also the answer to the bigger question this essay has been circling, because they’re the same answer. The demand and the uncontrollability are the same fact. What finances the ride from 2 FPS to 120 is the desire for compute that nobody else can meter, watch, or refuse.
If that still sounds abstract, make the list of every way anyone controls AI today: the refusal trained into the model, the usage policy, the API key, the rate limit, the audit log, the deprecation notice, the export order that turned off the frontier for nineteen days in June. Every item on that list is a property of the deployment, not of the weights. All of them assume the model runs on a computer you rent. Weights on your own machine dissolve the whole list at once. Not weakened. Gone.
I watched the last version of this from inside Microsoft. The PC spent years being worse than the mainframe. It won by being yours, and the spreadsheet gets the credit, but the disruption was IBM no longer being in the room. GLM-5.2 grinding on a secondhand Mac Studio at seventeen tokens per second is worse than the frontier, and it’s mine. There’s nobody else in the room.
You are at 2 FPS. The people who remember 1998 know how that felt, and they know how it ended.
References:
- Zhipu / Z.ai (2026). GLM-5.2 model card. Parameters, MIT license, context, and (self-reported) benchmarks. Local quant sizes via Unsloth; the 512GB M3 Ultra throughput figure via oMLX.
- UK AI Security Institute (2026). How far behind the frontier are leading open-weight models on cyber? The 4-7 month gap, GLM-5.2 ≈ Opus 4.6, and the “persistent and irreversible” line. July 17.
- The Decoder (2026). Kimi K3 launch coverage. July 16. K3 benchmark numbers are Moonshot self-reported on its own harness; see also Abundant AI’s SWE-Marathon for the official leaderboard those numbers don’t map onto.
- OpenRouter (2026). State of AI: a 100-trillion-token usage study. Roleplay at ~33% of traffic and ~52% of open-model tokens. January.
- MacRumors (2026). 512GB option pulled (March 5) and 256GB follows (May 5).
- Gurman, M. / Bloomberg, via Tom’s Hardware (2026). The M6/M7 roadmap and the conditional 1.5TB M7 Ultra. July 13. Rumor-grade.
- Nvidia (2026). RTX Spark announcement with Microsoft (May 31) and the three-generation Spark roadmap.
- TrendForce (2026). Q1 DRAM price forecast, +90-95% (February 2) and Q3 deceleration to +13-18% (July 3).
- Gartner (2026). Memory costs to cut PC and smartphone shipments. February 26.
- Tom’s Hardware (2026). HBM is eating your RAM. The wafer-trade mechanism; also the SK Hynix CEO’s 2027 outlook and the Frozen v2 report (July 21, The Information sourcing, unconfirmed by Google).
- Bloomberg (2026). Moonshot in talks at a $50 billion valuation. July 21.
- ChinaTalk (2025). China’s new AI plan. The AI+ initiative and its open-source directives.
- Akin (2025). BIS rescinds the AI Diffusion Rule. The end of the only weights export control. See also EU AI Act Article 2 for the personal-use exemption.
- EFF. Bernstein v. US Department of Justice; Zimmermann, P. (1995). PGP source code book preface. The Crypto Wars precedent.
- Geerling, J. (2025). 1.5TB of VRAM on Mac Studio and clustering four Framework mainboards. What the cluster path actually costs and yields.
- Figure (2025). Helix. Onboard robot cognition, and why.
- Zatloukal, P. (2025-26). LLM Inferencing Costs are Going to $0, The Bitter Lesson of Agentic Coding, Agent-Hypervisors, and First Fable. The price argument, the verification ceiling, the router at the seam, and the frozen-threshold test this essay leans on.