RAIVEN Field Report · Seat economics

From Tokenmaxxing to Seatmaxxing

Everything you ever wanted to know about AI/LLM seat economics: RAIVEN lab observations, with receipts · Field report: two top-tier AI seats, instrumented over seven weeks in June and July 2026. 43,943 recorded model responses, one ledger, priced at API list rates. Observations, results, and the discussion they invite.

Richard Vdovjak · RAIVEN 2026-07-30 AI-polished, never AI-invented · anchored in real work in the technology trenches. Read more about our promise.

The most expensive AI seats can also be the cheapest, per token they deliver. The catch: you have to actually put them to good use. Reddit forums are already full of people planning their lives around five-hour Claude usage reset clocks. Add a second vendor and it compounds. If only there were a way to handle this more effectively (hint: something is brewing).

Claude's own composer bar: Out of usage credits, with an Upgrade button
2026-07-29, 13:22, the lab's own Claude seat: the screen heavy users know. Anthropic estimates fewer than 5 percent of subscribers ever see it; a floor that takes the whole seat gets there by design. What it does not get is a stoppage. Receipts below.

TLDR, for the execs

Let's get one thing out of the way

AI spend is not AI value. Seats or tokens, subscription or metered: it is all cost, and cost delivers nothing by itself. So before we talk about seats, one call-out: the strategy work better already be in place. Most companies are still working on the questions that come first: how can we leverage AI, and where does it add value to our business? Those are the right questions, in that order, and they deserve answers, or at least verifiable hypotheses, before the first structural euro of AI spend.

Two more boundaries belong in the same disclaimer. First, AI is more than LLMs; LLM seats, our subject here, are one instrument in a wider kit. Second, no model at any price will fix your data problems, your governance problems, or the organizational transformation your AI ambitions may actually require. Those are strategy questions with their own discipline, beyond this report's scope.

(And if those questions are open in your shop, that is work RAIVEN does: AI strategy and AI-accelerated product strategy. Reach out.)

Everything from here on assumes that work is done. The starting point: you know where you are aiming LLMs and why. The clearest case, and our own: you are a software company, and assisted or even agentic-driven coding sits high on your list. From here, the subject is narrow and operational: the capacity you then buy, and how not to waste it.

The wave, and the misread

On May 14, 2026, Anthropic announced it would split Claude billing in two, effective June 15. The headlines read like a price hike on agents. Then June 15 arrived and the change did not: Anthropic paused it before it took effect. Its help center says it plainly today: “For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits,” with an explicit promise of notice before anything takes effect (Anthropic's own help center, live-checked 2026-07-30).

The announced design is still worth reading, because it shows where the vendor wants to land:

Read the pause correctly: the fence is not built, but the intent is on record, and pricing was not the first instrument. Enforcement moved earlier, in a sequence that reads as a ratchet:

So the headless eldorado is ending, just not by the route the headlines picked. What is ending is not your seat. It is the gray zone where an unattended script, wearing a subscription OAuth token inside an unsanctioned harness, quietly drew on the same big interactive pool a human pays for. A bot runs around the clock; a person cannot. That subsidy was not built to last, and announce, pause, rework, enforce is one screw tightening, not a retreat. The same lines exist in OpenAI's terms on paper; only the enforcement cadence differs. Do not build a company on the assumption that any shortcut stays open.

The part that is seldom discussedYou are not using your seat.

The expensive, valuable thing, the big interactive allowance, is the part most people barely touch. One chat, one question, wait, repeat. A 200-dollar seat driven like a 20-dollar one. Anthropic itself said new weekly limits would affect fewer than 5 percent of subscribers (T1, Anthropic's own announcement), which is another way of saying that for more than ninety-five percent of subscribers, the caps might as well not exist.

There is a twist inside the limit mechanics that makes the complaints and the statistic both true. The top plans meter in nested windows, an hours-scale one inside a weekly one; and on the top Claude plan the most capable model tier adds a third, its own separate weekly window inside the seat's (Anthropic's own guidance notes weekly limits that "reset on different schedules" per model, usage-limit best practices). Hit the short window a few times in a burst and the seat feels capped; look at the same week in total and the weekly allowance was nowhere near spent. Bursty, hand-driven use produces exactly that: real friction at the short window, real waste at the long one. The forums are full of people who responded by planning their day around the reset clock, down to firing a throwaway prompt before dawn so the window opens before they do (one widely shared trick), which is dedication, and also a job no human should have. Now add a second vendor: different window lengths, different offsets, a weekly clock that rolls on a different day. Planning around one vendor's limits is a hobby. Planning around two is a scheduling problem, and scheduling problems are what software is for; keeping the windows visible and deciding what runs when against them is precisely the kind of thing we ended up instrumenting, and it comes up again below.

The seat rewards parallel, coordinated work. The bottleneck was rarely the allowance. Far more often it was you, the human shuttling text between windows one prompt at a time.

Claude's native usage popover: 5-hour limit at 43 percent, weekly all-models at 100 percent, weekly Fable at 100 percent
The nested windows, in the vendor's own popover, minutes after our seat walled: the five-hour window at 43 percent with almost an hour still to run, both weekly windows pinned at 100, resets 37 minutes out. Three clocks, three readings, one seat. (The context-window bar on top is a different budget entirely.)

First, the cautionary tale: tokenmaxxing

The press spent this spring documenting what happens when usage becomes the goal. Enterprises made AI usage a target: internal leaderboards ranked developers by token consumption, and employees, being employees, hit the number. Amazon shut its own leaderboard, KiroRank, down in late May after staff padded token counts with make-work; the line from senior vice president Dave Treadwell was "Please don't use AI just for the sake of using AI" (reported coverage; first reported by the Financial Times). The trade press already has a name for the pattern, tokenmaxxing, and it is their word, not ours; by July the follow-up discourse had moved to token spend as the metric to watch. The bill for measuring the wrong thing is real: Bain surveyed 951 companies with more than $100 million in revenue and found almost 40 percent came in under 10 percent cost savings, against targets that were mostly 11 to 20 percent (survey coverage).

Usage is a poor proxy for impact, and often a misleading one: AI use is not AI impact. One directional truth survives the wreckage, though: little to no use does reliably mean little impact, because nobody gets value from capacity nothing touches. So utilisation earns a place on the dashboard as a diagnostic. Impact stays the scoreboard. That is the pattern we are about to name ourselves against.

Seatmaxxing, and the two ways to waste a seat

First, the name, offered tongue in cheek. We know seatmaxxing is a silly word; the -maxxing suffix earned its reputation one leaderboard at a time. We roll with it anyway, because the silliness is the point: it names the exact inversion of the pathology above. Tokenmaxxing maximises what you appear to burn. Seatmaxxing maximises what you already bought.

Seatmaxxing: prudent utilisation of capacity you have already paid for. A seat is wasted in two different ways: by building wrong or useless things, and by not utilising it, leaving paid capacity empty. Both are waste. Hoteliers have a saying for the second one: the most expensive room is the one you did not sell. Seat capacity is the same perishable inventory; every window that resets unused is a room night gone, and no amount of next week buys it back. The 200-dollar seat is expensive when it sleeps, and expensive again when it burns its nights on junk. Hotels answered perishability with a whole discipline, yield management, and that is what seatmaxxing is at its core: a discipline, not a metric. There is no seatmaxxing score to hit, nothing to rank on a leaderboard, no way to win by burning more. It is a frugal coordination discipline: product and AI strategy decide what gets built, which fills the backlog; seatmaxxing coordinates the seats through it, the right work on the right seat in the right window. The definition carries its own fence: seatmaxxing happens through the vendor's supported automation surfaces, inside the vendor's terms. Driving the work itself through injected keystrokes, at scale, is a different activity altogether, whatever it calls itself; the fair-play section below draws that line in full.

The empty-capacity waste often starts at procurement, and that angle gets less press. The seat tiers work like mobile phone plans. Of course it is a waste to pre-pay 500 minutes a month when you barely talk 10 of them. It is the same waste, in the opposite direction, to jump straight to the 200-dollar seat with no useful work queued for it. Right-size the plan to the work first, then consume what you sized. Plan choice is a procurement decision; maxxing is the operating discipline that follows it, and no amount of operating discipline repairs a bad sizing.

And as with phone plans, the tiers reward volume: the top one carries the better unit rate. Anthropic today lists Pro at 20 dollars, the 100-dollar Max at five times Pro's usage, and the 200-dollar Max at twenty times (Claude help center, T1; pricing is perishable, re-check at publish). Do the division. The 200 plan buys twenty times the capacity for ten times the price, so each unit costs about half what it does on Pro. The 100 plan raises the ceiling at roughly Pro's unit rate. The volume discount lives at the top, which is worth knowing before you assume the middle tier is the safe compromise.

The receipts, with the caveats firstObserved replacement cost is not a promised token allowance.

Anthropic does not sell Max 20x as a fixed number of tokens. Its limits are weighted and can vary with model, feature, prompt length, cache state, session and weekly windows, and capacity management. Our ledger is incomplete by design too: it covers local Claude Code transcripts on one machine, including subagents, but excludes Claude on the web, mobile, other machines, and any history no longer present locally. Transcript metadata tells us processed input, cache reads and writes, and output. It does not reveal Anthropic's private seat-limit weighting or prove whether every recorded token drew from included capacity rather than any separately enabled usage credit. Processed-token volume also includes cached context each time the model reads it; it is not a count of unique new text. Anthropic makes the same point from the other side: only new or uncached portions count against your limits, which is a large part of why a cache-heavy workload travels so far on one seat. The number below is therefore a lower-bound observation priced as a replacement-cost thought experiment, not a Max invoice, cash value, resale value, guaranteed entitlement, or claim that an interactive seat can sustain arbitrary API-shaped demand.

With that boundary fixed, the observation is still remarkable. From 7 June through 27 July, one instrumented Max 20x seat recorded 10.8 billion processed tokens in local Claude Code alone. Of its processed input, 95.1% came from cache reads. We priced the recorded input, cache reads, cache writes, and output at the applicable model-specific Anthropic API list rates, using the documented five-minute to one-hour cache-write range and a dated USD/EUR conversion. The after-cache API-list replacement cost is €10.2k to €12.1k. Without cache pricing, the same corpus would be about €57.8k. In 1 to 27 July alone, the ledger recorded 7.0 billion processed tokens, equivalent to about €7.0k to €8.4k at those list rates.

The Codex seat gives us a useful cross-check and a warning about false precision. On 27 July, its local account counter reported 36.7 billion aggregate tokens, but did not split them into cached input, ordinary input, and output. Pricing that unknown mix at the two extremes produces a technically honest but nearly useless €160.4k to €962.3k input-to-output envelope. A more behaviorally grounded scenario asks: what would this Codex token volume cost if we use it in roughly the same way we demonstrably use Claude? Projecting Claude's observed mix at that snapshot (94.7% cached input, 4.8% ordinary input, and 0.5% output) onto GPT-5.5 standard short-context API prices produces an API-equivalent estimate of about €27.8k. That is a cross-vendor workload-mix scenario, not measured Codex token composition, an invoice, or a claim that OpenAI's cache behavior is identical to Anthropic's.

Against a 200-dollar monthly seat, July's partial-month replacement cost is roughly 40 to 48 times the sticker price at the same FX snapshot, or 3,900 to 4,700 percent more. That is not ROI: tokens are not outcomes, and unused capacity is worth zero. It is a capacity finding. When a workflow can feed sustained useful work through supported interactive surfaces, a top-plan seat can deliver token throughput measured in billions per month and API-list replacement cost measured in thousands. The work still needs judgment, coordination, and a queue worth doing. The capacity itself is not small.

One more boundary keeps the term honest: seatmaxxing says nothing about what the seat is used for. Utilisation is a virtue only when the work is useful, and "useful" is carrying that whole sentence. Turn seat utilisation into a corporate metric and it will corrupt exactly the way token counts did; people are endlessly inventive at hitting a number. The assumed setup is the lean one: a solo founder or a frugal team with more useful work queued than hours to do it, leaning into the price gap between a flat seat and metered tokens. No queue, no case. Buy the smaller seat.

The seat gives you one mercy even when discipline slips: the damage is capped. A wasteful metered program discovers its waste on the invoice, after the fact, with no ceiling; the leaderboard enterprises above found that out in public. A wasteful seat costs, at worst, the subscription you already paid. The seat is a cost fuse. That does not make waste virtuous; it makes it bounded, and bounded mistakes are the kind a lean operation can afford to learn from.

And the honest, boring name for the whole discipline is capacity utilisation, the same concept operations teams already run for machines, rooms, and fleets: plan it, instrument it, operate it. A live gauge before you commit to a night of work, a long consumption trail to learn your own patterns, a runbook for the boundary. That is the toolkit. If "seatmaxxing" reads too internet for your taste, use the boring name; the discipline is the same either way.

There is a myth pointing the other way: seats are toys, the API is the real deal. Half right. An LLM embedded in your product, serving your users (the recommender, the support bot), is API territory, by economics and by terms alike. The seat is for the makers' side: building and coding, review, research, the general knowledge work a person at a keyboard would otherwise shuttle by hand, and the decision support behind it. On that side of the line, the seat is not the toy; it is the workhorse.

A third-party check: the same tokens on GitHub Copilot's meterAdded 2026-08-02. Different vendor, same arithmetic.

Both figures above are priced against the model vendors' own API cards, which invites a fair objection: those are the rates of the companies selling the seats. So we re-ran the ledger against an independent metered lane, the one most enterprises actually buy. GitHub Copilot bills overage in AI credits, one credit is one US cent, charged per token once a plan's included allowance is spent (the rate card, checked 2026-08-02).

The first result is the interesting one, and it needed no arithmetic at all. Copilot's per-token rates are the vendors' list rates. Claude Opus 5 is billed at 5 dollars input, 0.50 cache read, 6.25 cache write and 25 output per million tokens on Copilot's card, which is Anthropic's published card to the cent; Fable 5 matches at 10 / 1 / 12.50 / 50; GPT-5.6 Sol matches the GPT-5.5 rates our Codex scenario already used. A large intermediary re-selling frontier capacity does not mark it up per token. It differentiates on allowance, tooling and procurement, not on the meter.

So the Claude figure transfers unchanged: the same 10.8 billion processed tokens cost the same 10.2 to 12.1 thousand euro of replacement value on Copilot's meter. One detail sharpens it. Our range exists because transcript metadata omits cache TTL, so we priced the five-minute and one-hour cache-write rates as bounds. Copilot's card lists a single cache-write price, the five-minute one, so on that lane the honest figure is the bottom of our range, about 10.2 thousand euro.

The Codex seat needs one estimate, and we would rather name it than hide it. The Claude ledger records which model served each turn, so its numbers are model-specific throughout. The Codex account counter does not: it reports an aggregate token total with no model attribution and no cached/ordinary/output split. For the split we reuse the Claude-observed mix, as before. For the models, the operator estimate is 80 to 90 percent GPT-5.6 Sol and the remainder Terra. Priced on Copilot's card across that band, the same 36.7 billion tokens land between roughly 24.5 and 26.1 thousand euro, with the midpoint near 25.3 thousand. It comes in under the all-Sol scenario for the plain reason that Terra is cheaper, which is itself the argument for routing work to the smallest model that carries it.

None of this is a bill, and Copilot is not a lane we run on. It is a check on the comparator: when an independent metered lane prices the same work within rounding of the vendors' own cards, the gap between a seat and a meter is a property of the purchasing model, not of one vendor's price list.

Which plan should you or your team choose?

If you have to ask, the what and the why are probably not nailed down yet, and that is the thing to fix first. A plan is a five-minute change; a fuzzy goal is not.

Once the goal is clear, sizing is empirical rather than a guess. Start on the lower tier, Pro. See how fast you max it. When you hit the wall before the window resets, bridge the rest of the block with pay-per-use, and put your own hard cap on that spend so one busy day cannot become a surprise invoice. Then do the arithmetic: if the pay-per-use you keep buying to finish each block costs more than the next plan up, the next plan up is the cheaper seat. Move.

In our shop that math converged fast, and on the top tier of the top vendors. Right now, if you are really building, the value in these plans is significant. Yours may differ; the method does not. Measure your own usage against your own work and let the numbers pick the tier.

The commitment terms make honest sizing cheap in both directions: month to month at Anthropic, in places shorter still at Google. So downscaling is a move too, and it should be. Two refinements keep the call honest. First, read idle at the right timescale: a seat sleeping through an evening is nothing, the short windows reset on their own; the perishing unit is the weekly allowance, and weeks that keep ending with most of it untouched are the real signal. Second, the tiers price capacity at different unit rates: on Anthropic's list the 20x plan buys four times the 5x plan's capacity for twice its price, so dropping a tier halves the bill but cuts capacity to a quarter, roughly doubling what each unit costs. The trail has to show sustained headroom, not one quiet week. The mechanics reward deliberation over impulse: a tier change is scheduled rather than instant, so a downgrade lands on the next billing date and you keep the higher tier until then. That makes dropping a tier a decision you take once, against the trail, on a date you already know. If the long usage trail shows a pattern of capacity you never touch, run the same arithmetic in reverse and drop a tier; this is exactly what longer-horizon monitoring is for, and ours keeps a 400-day memory precisely so those patterns have nowhere to hide. And once you run seats from several vendors, a sharper version of the question appears: which seats actually decide outcomes, and which one could you drop? We answer that from the cross-vendor review record itself, learn from xLLM, then keep the ones that earn their seat. That drop-one method deserves its own piece.

How seats scale

One seat per vendor is where most shops start. It is not where the ceiling is.

The first scaling move is not a second seat from the same vendor; it is a seat from a second vendor. The capacity math is the same, another allowance on another clock, but the second vendor's seat pays three dividends the first vendor cannot sell you. Independence: a cap, an outage, or a policy turn at one vendor becomes a rerouting decision instead of a stoppage; the no-lock-in principle below, in its procurement form. Scheduling freedom: the two vendors' windows reset on different clocks, so work flows to whichever seat has headroom, which is exactly what makes the reset-clock life-planning above obsolete. And the dividend that pays our rent: model families overlap less in their blind spots than seats from one family do, so a second family adds a reviewer, not just a lane. Same-family seats share training, and share blind spots with it; cross-family review is where the hardening comes from. That is the xLLM method in one sentence, and it is why our own fleet crossed vendors before it ever doubled a seat, and then crossed again: the floor runs three families today.

The cross-vendor dividends come with a coordination overhead: real, but manageable, and in our use cases the benefits outweigh the cost by far. The two vendors' agents do not talk to each other. Out of the box, the bridge between a Claude agent and a Codex agent is you, copy-pasting between two windows; that was the bottleneck at one vendor, and at two it compounds. Cross-vendor seats without cross-vendor orchestration just double the shuttle.

Buying more seats is not a crime. During the February ban scare, an Anthropic engineer said it on the record: "It's not against terms of service to have multiple MAX accounts"; what gets enforced is credential sharing, reselling, and subscription tokens driven through third-party harnesses (the statement). OpenAI went further and shipped the pattern as a feature: official account switching, with its help center stating you can create additional accounts. The conditions are the ones you would expect, and they are conduct, not arithmetic: each account is yours, used by you, through official clients, each limit respected. N seats bought is N seats owed; nothing is being evaded, and the vendors sell overflow usage at API rates themselves. The practical caveats are operational rather than contractual: keep payment methods boring and consistent, never share credentials, and know that automated enforcement has produced false positives at both vendors, so appeals, not workarounds, are the recovery path.

Then there is the team axis, and it buys a different thing. Team tiers bind seats to named humans, one each, and price the bundle above the prosumer seats per unit of capacity. What the premium buys is terms: contractual no-training instead of per-human privacy toggles, central billing and admin, seat reassignment when people change. The arithmetic, at list prices as of late July 2026:

RoutePriceCapacity, in Pro unitsWhat binds it
Claude Pro$20/mo1xone person
Claude Max 20x$200/mo20x (~$10 per unit)one person
Second Max seat, same person+$200/mo+20x, independent windowsofficial clients, no sharing
Top seat at a second vendor (e.g. ChatGPT Pro)$200/mocomparable top-tier allowance, independent windows and clocksone person; a different model family
Claude Team Premium$100-125/seat/mo~5-6.25x (~$16-25 per unit)one named person per seat; org terms
ChatGPT Business standard$20-25/seat/mobundled baseline + creditsone named person per seat; org terms
API, meteredper tokenunboundedthe comparator the receipts price against

EU note, with every figure normalised to VAT-inclusive at the Dutch 21 percent so the columns actually compare. Claude: Pro €18 a month on annual billing or €22 monthly, Max 5x €109, Max 20x €218. ChatGPT: Go €8, Plus €23, Pro €103 at the 5x tier and €229 at 20x. The two vendors' matching tiers land within about five percent of each other at both levels. Buy through a VAT-registered company and you see and pay the ex-tax figures instead, reclaiming the rest, which flips the order slightly: €90 against €85 at 5x, €180 against €189 at 20x. Watch which of the two you are being quoted before comparing anything. Every euro figure elsewhere in this report is converted at the dated FX snapshot in the methods note.

Read the table the way a CFO would. Team seats cost more per token than the prosumer top tier, and both are far from API rates for practitioner workloads. So the scaling story has three honest axes: a second vendor's seat first, where independence and review diversity pay on top of capacity; more prosumer seats per person where raw capacity is the need and the conduct rules are kept; and team seats where the org needs the terms more than the tokens. No axis is a workaround; all three are the vendors' own price lists, read carefully.

The competition dividend

Multi-vendor has a second-order payoff we did not plan for: you become the beneficiary of the vendors' competition for you. The instruments made it visible. Our Codex trail is dotted with two kinds of markers: green triangles, banked reset vouchers we chose to spend (OpenAI's referral promotions pay out in exactly these), and blue triangles, resets the vendor granted unprompted. Each triangle snaps the weekly allowance back to full. In the stretch pictured below, the nominal one-week limit behaved like a few days between refills: the 4-week utilisation counter reads over 160 percent of nominal caps, six and a half weekly allowances consumed in four weeks, and everything past the nominal four arrived as resets, inside the terms.

Three vendors' seat trails over eleven days, with reset markers, a two-day Fable block, and the Claude seat hitting its weekly wall at the frame's right edge
Three vendors, eleven days, one frame, captured minutes after the wall above. Top: the Claude seat, its Fable model blocked for two days earlier in the frame (49 hours in the selected span) while the seat's other models kept working, and the frame's right edge pinned red where the whole seat just walled. Middle: the Codex seat's weekly allowance snapped back to full five times; green triangles are vouchers we chose to spend, blue are resets the vendor granted unprompted. Bottom: the Antigravity lane, newest on the floor and not a top-tier seat; the receipts in this report come from the two that are. The fleet bar carries the thesis: hours with no seat free, zero.

Read the frame top to bottom. In the same days the Codex seat kept being refilled, the Claude seat's Fable model sat blocked for 49 hours, its own weekly window pinned for two straight days while the seat's other models kept working. A partial outage is a routing event, not a stoppage. One vendor's gift landed while the other vendor's most capable tier stood at its wall. A human juggling two tabs cannot harvest that. A planning agent can: ours watches the gauges and reallocates, so when a surprise reset lands mid-week, queued work shifts to the family whose seat just came back, and vouchers get spent deliberately, expiry-first, ahead of walls rather than after them. Capacity that arrives unannounced is only capacity if something is ready to consume it within hours. That is scheduling, and scheduling problems are what software is for. And the frame ends mid-lesson: it was captured minutes after the Claude seat walled outright, both weekly windows at 100, the event already inked at the right rail. Even so, the fleet counter for hours with no seat free reads zero. Eleven days, one seat fully walled, another past 160 percent of nominal, and not one hour when the floor could take no work.

The tactical view: the Claude gauge fully red at 100 percent on both weekly windows, ChatGPT at 8 percent, Antigravity at 62
The tactical view, minutes after the wall. Left to right: the Claude seat at 100 on both weekly windows, the same numbers the vendor's popover showed, its five-hour arc at 42 with capacity the weekly wall makes unspendable; the ChatGPT seat at 8 percent of its week, wide open, voucher bank empty because the vouchers above were spent; and the Antigravity seat, Google's Gemini family, cruising at 62. One wall, two open roads.

The gauge and timeline captures are one instrument at two zooms. The gauges are the tactical view, three vendors' seats on one pane: our coordinating agent reads exactly this in real time, so a surprise reset becomes a routing input within minutes of landing, not a discovery at the next stand-up. The timeline earlier in this section is the same instrument zoomed out, the strategic view: longer-term usage patterns feeding sprint capacity planning, which weekly windows to book night shifts against, when to expect walls, how much headroom a sprint can honestly promise.

The dividend is not only additive capacity; vendors also compete by removing friction. OpenAI silently dropped the five-hour window on this seat. There was no announcement and, as far as we can find, no one else writing about it: our dashboards simply stopped reporting a short-window figure on the Codex side, across the eleven days in the frame above and since, while the Claude seat's short window kept ticking beside it. Sprint planning took the relaxed restriction immediately, moving burst-heavy work to the seat with no short window to trip over. That is the argument for watching the gauges rather than the announcements: the capacity showed up in the instrument before it showed up anywhere else, and it has still not shown up anywhere else we have looked. They may put the window back tomorrow, and that would change nothing here: the gauges would catch the return exactly as they caught the removal, and the planner would route around it the same afternoon. The point is not that this particular limit is gone. It is that a floor reading its own instruments responds whichever way a vendor moves, without waiting for an announcement that may never come.

Whether the resets are retention play, capacity smoothing, or plain generosity is the vendors' business; we claim no motive. The operational fact is enough: on a multi-vendor floor, competitive behaviour arrives as free capacity and removed friction, and it takes an instrumented, orchestrated floor to notice in time to use either.

The verdict on value, and the catch

Put the sections above together: for top-tier AI-assisted coding, the seat is the best value in town, and on our receipts it is not close. The July ledger prices one 200-dollar seat's local traffic at roughly 40 to 48 times its sticker, 3,900 to 4,700 percent more, at after-cache API list rates, caveats exactly as stated above. The tier arithmetic says the top plan buys capacity at about half the unit rate of the entry plan. If the backlog is real, nothing else we priced comes near a well-driven seat, and most users are leaving most of that on the table.

The CFO question: assuming your LLM work is seat-eligible, practitioner work rather than in-product deployment, in which universe would you deliberately pay thirty to forty times more for the very same underlying models through the metered lane? That is what defaulting to API-only for builder workloads amounts to, on our after-cache receipts.

The catch: that value assumes throughput, and throughput at seat scale is an orchestration problem. One human drives one lane. The consumption above comes from many agents working in parallel, and the moment you take vendor lock-in seriously (we do; it is an operations risk before it is a philosophy) you are orchestrating agents across more than one vendor: separate identities, separate limit windows on separate reset clocks, hand-offs between models that check each other's work. None of that is trivial, and no vendor ships it across its competitors today.

We had to build that layer ourselves. The rest is how.

How we take most of the seat, on the paved road

So we automated the shuttle, not the thinking, and we did it without faking anything. Our review and decision work runs as a multi-agent loop coordinated through git: every request, critique, and retraction is a committed artifact, and the agents coordinate by reading and writing that log. No human copy-paste in the middle.

The tool has a name: RAIVEN AgentOS, our meta-harness. It sits above each vendor's own harness and coordinates the fleet across them, through each vendor's supported automation surfaces. It was built to deliver xLLM, our cross-vendor review protocol, and it earns its keep twice: the same coordination that lets different vendors' models check each other's work is what drives seat utilisation up, because a coordinated crew consumes the allowance a lone human shuttling prompts could not hope to.

The load-bearing term is paved road: the automation the vendors build, document, and actively invite you to use. They ship agentic loops, background and scheduled tasks, hooks, skills, and they publish the worked examples showing you how to wire them together. We take them up on the invitation, and we add the one part none of them ships, because none of them has a commercial reason to: the same coordination running across all of them at once. A Codex agent on its own seat and a Claude Code agent on its own seat read and write the same git log, so one picks up what the other left, reviews its work, disagrees with it in writing, and hands it back. Different vendors, different blind spots, one queue. Built from parts the vendors handed us.

What we do not do is the other half of the definition. We do not extract a token and call the API as if we were the app. We do not inject keystrokes to fake a human at a keyboard. The coordination lives in git, outside any vendor, so each model is reached only through its own sanctioned surface.

The person stays in charge the whole time. These are live sessions, not a batch job thrown over a wall: you can read any agent mid-task, redirect it, stop it, or take the keyboard back, and you should. What the harness takes off your hands is the mundane part, the copy-paste between windows and the round-the-clock question of which seat has capacity right now. The protocol itself is tailorable, with human gates written into it explicitly. How often those gates fire is a setting rather than a doctrine, and the right frequency is a property of your domain: a prototype and a change to a payments path do not deserve the same number of checkpoints.

Two things fall out of that design:

  1. It runs on the seat you already pay for. Because coordination lives in git, each agent can be an ordinary interactive session. You get automated, multi-agent throughput on the seat itself, which is exactly the capacity most seats leave on the table.
  2. Transport becomes a dial, not a commitment. The same harness runs a given agent attended or unattended. If a job genuinely belongs in a scheduled, nobody-watching lane, you move it there deliberately and pay whatever that lane costs. You can. You do not have to. You choose, per project, per budget, with no re-tooling.

Fair play is the strategy, not the constraint

There is a quieter reason to stay on the paved road: it is the surface you can keep building on.

We do not spoof. The vendors built these tools, and they can restrict or meter automation whenever they decide it is being used in a way they did not intend. That is their call. Terms of service are theirs to change, and over the past year all three have changed them. We build on what they give us, in the way they give it to us. So far a policy shift has cost us a config flip, not the company.

The crowd reaching for extracted tokens and injected keystrokes rebuilds every time a vendor tightens a screw. We do not, because we never leaned on a shortcut staying open.

Guiding principle: no vendor lock-in

The switching story deserves to be a principle, not a footnote. In a field this young, the model leaderboard changes every few weeks, the FOMO is real, and so are the downstream integration costs once a whole foundation is built with a single vendor in mind. And the vendors do tempt you: agent teams and fleet features that are genuinely good, and single-vendor by design. (Anthropic's agent teams are excellent. They are also Anthropic-only. Each vendor has its own version of that pull.)

Single-vendor is an operations risk as well as a strategic one. Every cap, outage, and policy turn becomes your cap, outage, and policy turn. Even on the top Claude plan, the Fable tier carries its own separate weekly limit; when it binds, that capability is gone until the window resets, however important your deadline. And availability risk is not only limit mechanics: in June 2026 a US export-control directive forced Anthropic to abruptly suspend Fable and Mythos for all customers, over the company's own public objection (the announcement); access later returned. A fleet spread across vendors turns the same event into a routing decision. And the availability question can run further still, into sovereignty proper: data that must not leave the jurisdiction, or the premises at all, which puts local models on the table. Those constraints sit outside this report's scope; a future report will pick them up.

The major plans are all month to month: cancel, upgrade, downgrade anytime, no annual lock-in (Anthropic · OpenAI · Google). So the lock-in is not the subscription. It is your workflow. Coordinate through git instead of through a vendor's API, and swapping the model behind a slot is a config change. Ride whichever model is currently best, and hop when the lead moves.

The honest part

This is not a cheat code. Heavy unattended automation still costs real money through the metered lane, and it should; the announced headless credit, if it ships, is small and hard-stops. "Fair play" is not a vendor blessing either, it is just staying on the surfaces they sanction, and the line between supervised automation and a bot is genuinely grey. The vendors' own enforcement moves show where their worry lives: the tightenings so far, weekly caps, the OAuth crackdown, the announced headless split, have aimed at scale, unattended fleets running around the clock and accounts pooled into farms. That is the line we treat as bright. And none of this makes any single model correct; that is a different question, and we have the receipts for it. What it does is stop you paying for a seat you use a tenth of, without betting the work on a shortcut that closes.

Take the whole seat. You are already paying for it. Every token of it, inside the terms.

What this report does not claim: a fixed token allotment or entitlement from any vendor; an invoice, cash value, or ROI; that tokens are outcomes; that utilisation is a virtue without a useful queue; that seat capacity suits in-product LLM deployment; that any single model is correct; or that today's prices, limits, and terms will hold, since all three vendors changed theirs inside a year.

Vendor plan, API-pricing, and help-center documentation was live-checked on 2026-07-30. The four captures reproduced above were all taken on 2026-07-29, within minutes of the same weekly-wall event: two are Anthropic's own client UI, two are RAIVEN's own monitoring, and the two sources agree on the walls they show. The May-14 headless-billing announcement and its June-15 pause are from Anthropic's own help center, and the February 2026 third-party-harness prohibition and ban-wave record are anchored in RAIVEN's evidence vault (Wayback snapshots plus an X syndication capture). The "fewer than 5%" figure is Anthropic's own statement (T1, @AnthropicAI; corroborated by TechCrunch). The replacement-cost evidence is RAIVEN's lower-bound local Claude Code transcript ledger, deduplicated by provider message ID: 43,943 recorded model responses (assistant turns) observed from 2026-06-07 through 2026-07-27. Its API-equivalent range applies model-specific list prices, documented cache rates, a five-minute to one-hour cache-write assumption because transcript metadata omits cache TTL, and the 2026-07-24 USD/EUR snapshot. It is not Anthropic billing or a disclosed seat conversion.

P.S. New here? What a RAIVEN field report is, and what it refuses to be:

About these field reports

Let's be honest: most AI-generated content in your feed right now is slop. Prose written from other prose, about nothing that happened, worth nobody's time. Here is my promise instead: these posts are AI-polished, but never AI-invented. Every one of them is anchored in real work in the technology trenches.

They are first-hand field reports from RAIVEN, my AI lab in Eindhoven: a working multi-vendor fleet of AI agents that I run daily, both to build real systems and to find out exactly where this technology breaks, and what imperfect but practical remedies push it further. Some reports come with full forensics behind them, down to file hashes. Others describe a pattern from real work whose details stay private: client engagements, security-sensitive systems. Each report says which kind it is, and its claims never pretend to more evidence than it shows.

I write first for executives and decision makers steering in-house AI efforts. The strategy world is full of AI decks assembled by people who never operated the technology they present. These reports are the opposite: trench-level truth, translated into what it means for your decisions. I use my own agents to draft and polish; I am not a native English speaker, and I have a fleet for a reason. But the experiences are real, the judgments are mine, every post is vetted and signed by me, and the numbers are checked against the underlying evidence before publishing.

Field reports are dated observations, not eternal truths, and I will get some things wrong. When a substantive mistake surfaces, whether I find it or a reader does, I publish a dated correction on the post itself. Being corrected in public is part of the method here, not a failure of it.

P.P.S. Changelog:

2026-08-02: added the section "A third-party check: the same tokens on GitHub Copilot's meter". It re-prices the same two ledgers against GitHub Copilot's AI-credit rate card, an independent metered lane, and reports that Copilot's per-token rates match the model vendors' published cards. Nothing in the original receipts changed; the Claude figure transfers unchanged, and the Codex figure is restated with the operator's model-split estimate named in the text, since that seat's counter records no model attribution. Added because pricing a seat only against the seat vendor's own card is a fair thing for a reader to question. The same update also gave the 40-to-48-times figure a percentage equivalent, 3,900 to 4,700 percent more, wherever it appears, plus the inverse framing in the TLDR: the ratio is the one already measured, restated because a multiple is easy to read past.