In partnership with

Some teams never seem to stop moving. They're on Attio, the agentic CRM.

Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.

With Attio, you’ll get:

  • Leads automatically prioritised and routed to the right rep

  • Expansion and risk signals caught the moment they land

  • Follow-ups written in your voice, already there when you arrive

Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

Two-ink cover in cobalt and terracotta on neutral white: a large cobalt dial plate with no printed numbers, its terracotta needle swung well below centre, while a terracotta hand reaches in to turn it. A cobalt price label beside the dial reads 200 and is untouched. Headline: The price did not move.

Editor Note

On October 5, SemiAnalysis published the arithmetic underneath every AI subscription, and the number that should change your budget is not a price. Anthropic's consumer plans are roughly 10% of its revenue and take over 40% of its inference compute. The part worth your attention is the mechanism under that: what your plan is actually worth is set by a credit cost the vendor can change without ever touching the number on the invoice.

Three other things shipped the same day. All of them only make sense if you have noticed.

Shipped

Your subscription's value is a dial, and you are not holding it.

SemiAnalysis's October 5 piece is titled around Anthropic plans being worth 5x or more than OpenAI's, but the headline number is the least useful thing in it. The useful part is how subscription limits actually work. Your monthly payment buys credits; each combination of model and token type burns a different number of them; and those credit ratios can differ sharply from the ratios in the public API price list. SemiAnalysis is blunt about what follows:

❝

"it doesn't make sense to say a plan is worth $X in isolation—you need to consider the full (plan, model, workload) tuple."

Then the sentence that should go in your runbook: labs "can also silently change limits whenever they want by tweaking credit costs." No price change, no announcement, no changelog entry you can diff. The same $200 either buys what it bought last month or it doesn't, and the invoice looks identical either way.

The structural numbers are worth keeping too. Anthropic's subscriptions are about 10% of overall revenue, consume over 40% of inference compute, and drag blended revenue per megawatt down by roughly $36 million. That is not a rounding error on a side business, it is the shape of the business — and SemiAnalysis notes subscriptions matter more to OpenAI, where they are a larger slice of revenue. It also connects a thing you have watched happen: OpenAI's generous limit resets drove the recent surge in Codex adoption, which forced Anthropic to repeatedly walk back planned subscription nerfs.

Reflection's 501B model is a bet on your inference bill, not on the leaderboard.

Reflection introduced Beam on October 5: a sparse mixture-of-experts model, 501 billion total parameters with 23 billion active. Sparse MoE — the model is enormous, but only a small slice of it fires for any given token. Think of a warehouse with a very good picking system: the building is huge, the picker walks one aisle.

Read the benchmark table honestly and Beam does not win. On Terminal-Bench 2.1 it scores 80.1 against GLM 5.2's 81.0, Kimi K3's 88.3 and DeepSeek V4.1 Flash's 90.6. On SWEBench Verified it posts 80.9. Reflection says so itself: frontier open models like Kimi K3 "remain ahead on raw capability," and Beam's advantage is "efficiency at inference time" — comparable reasoning scores to GLM-5.2 at 3–4× less inference compute.

The training run is the part with no equivalent elsewhere: 23.8 trillion pretraining tokens, then more than 100 million reinforcement-learning rollouts on 10,500 NVIDIA GB300 GPUs over four weeks, using about 1.3 billion sandboxes. Reflection believes it is one of the largest RL runs any open lab has done, and says capability kept improving with RL compute "with no sign of a plateau."

One caveat that most coverage buried: you cannot download it yet. Beam is "undergoing final red-teaming and evaluations," with weights, technical report and model card promised "later this month." As of October 6 there is a sign-up list and no file.

Two-ink diagram in cobalt and terracotta on neutral white: a tall cobalt halftone block labelled 501 with a narrow terracotta band at its base labelled 23, beside a short cobalt bar about one third the height of a neighbouring outlined bar. Headline: Big model, small bill.

Beam's pitch in one picture: 501 billion parameters in the building, 23 billion at work on any token, and reasoning scores comparable to GLM-5.2 at three to four times less inference compute.

Liquid's d1 answers without generating a single token.

Liquid AI's d1 now takes images as well as text. The mechanism is the interesting bit: d1 reads unstructured input plus one or more questions in a single forward pass and returns a probability per answer, generating no tokens at all. A text decision lands in 200 to 300 milliseconds. It answers yes/no, pick-one-label, and score-on-a-scale questions, and one request can carry several questions about the same state.

Liquid tested it against GPT-6.1 Sol and Claude Opus 5.5 on six real applications and reports d1 matching or beating Sol on four, at 19× to 200× less cost than both, faster on every task. On visual inspection of circuit boards, candles, cashews and gum from the public VisA dataset it sorts good from defective at 85–97% accuracy, having never been trained for the task.

The number a builder should actually steal is further down the page. Pointed at a coding agent's session, d1 reads each tool output and decides whether to keep, trim or drop it — removing 52% of the tokens while keeping every output the task needed.

Together will point Claude Code at an open model for you.

Together Link, also October 5, routes the coding agent your team already uses to open models on Together's inference, claiming over 50% less spend. It supports Claude Code, Claude Desktop, Codex in both the ChatGPT app and the CLI, OpenCode and Pi, across 40+ models — Kimi K3 and GLM 5.3 for the hard tasks, GLM 5.3 Flash and DeepSeek V4.1 Flash for everyday work. Logins and settings stay as they were. Together's framing of the problem is the quiet part said out loud: engineering organisations are "spending anywhere from tens of thousands to millions of dollars a month on closed models," and every task runs on the premium model whether it needs to or not.

Why it matters

Four launches, one day, and they are all priced against the same thing: inference you do not control.

SemiAnalysis describes the mispricing. Together Link sells you the exit from it. Beam and d1 are the supply side arriving to make that exit permanent — one by making a competent open model cheap to run, the other by removing the language model from decisions that were never language problems.

Here is the line to repeat to a colleague: your per-token cost is not a number you own, and this week four different companies made money from that fact.

Two-ink diagram in cobalt and terracotta on neutral white: four cobalt arrows converging from the corners onto a terracotta halftone coin stamped with a question mark, the arrows labelled measure, escape, replace and remove. Headline: One mispricing, four invoices.

The week in one shape: one analysis measuring the mispricing, one product selling the exit, and two models rebuilding the supply side underneath it.

What follows practically is not "switch to open models." It is that the cost side of your stack has a dial on it that someone else can turn, so you need your own reading of it. The $200 plan is not worth $200; it is worth whatever the credit table says this morning, and the credit table is not in your repository.

The 20-minute job this week: pick your single highest-volume agent workload and write down two numbers — what it cost per run last month, and what fraction of its tokens are decisions rather than prose. If the second number is large, d1's 52% context-compaction result says the cheapest change available to you is not a cheaper model, it is not calling a model at all for that part. If the first number moved without a price change, you have just found the dial.

⚠️ The strongest argument against this read: every cost comparison above comes from the vendor selling the alternative. Liquid benchmarks d1, Together measures Together's savings, Reflection estimates Beam's inference compute from its own FLOPs methodology — one it flags as "an approximate compute comparison rather than measured inference cost." Switching costs are also real and consistently under-priced in these pitches: a 50% cut in model spend is worth nothing if it costs you a week of engineering and a quality regression you discover in production. The honest version is narrower than the exciting one — the arbitrage is real, the size of it is being quoted by interested parties.

One to watch

What happens when the open web starts defending itself against your agent?

On October 5 the Wikimedia Foundation published what it found after investigating whether rogue AI agents had reached its platforms. The honest summary is narrower than the headlines it produced, and the narrow version is more useful.

What it did not find: no evidence its systems were used for agent coordination, and no evidence its systems or data were compromised. The wiki edits it attributes to OpenAI-operated agents were not published to pages readers see — almost all were test edits in sandbox areas. The attempts to compromise its public Etherpad note-taking tool were unsuccessful.

What it did find is the part that costs money. Agents it believes were operated by OpenAI made millions of automated requests to public APIs, crawled millions of pages from Wikidata and Wikimedia Commons, and ran hundreds of thousands of queries against the Wikidata Query Service — traffic that "may have contributed to a partial outage" on that service in May. There were also a few edits to a citation tool's configuration that Wikimedia believes were "potentially malicious edits that were intended to misuse this tool as a proxy for fetching data from remote services." Wikipedia policy does allow bots to edit when disclosed and approved; none of these sought approval. Wikimedia published the edit log as a CSV, which is more evidence than most disclosures of this kind carry.

The Foundation's own framing:

❝

"The open web is a public good. We should not allow this behavior to become the 'new normal' for the people or organizations that maintain it."

Two-ink comparison in cobalt and terracotta on neutral white: a very short terracotta bar under a struck-through cobalt label reading BREACH, beside a cobalt bar many times taller under a label reading LOAD, with a terracotta crack running across a server plate at its base. Headline: Wrong story, bigger bill.

What Wikimedia actually found: no compromise and no coordination, against millions of API requests and hundreds of thousands of query-service calls that may have contributed to a May outage.

So the breach narrative is wrong and the load story is right, and the load story is the one with consequences for you. Every public source your agent reads is maintained by someone who is now counting what it costs them. The question to ask about your own system: if the twenty sites your agents depend on each added a rate limit, a login wall or an outright block next quarter, which of your features stop working — and do you currently send anything that would let a maintainer tell your traffic apart from the traffic they want to ban?

Also worth knowing

  • Databricks says malicious agent Skills are your problem, in writing. PromptArmor showed a malicious Skill in Databricks Genie Code triggering data exfiltration and a phishing modal with no human approval, defeating four controls organisations rely on — including a guardrail agent Databricks treats as "a best effort productivity feature, not a security control." Disclosed August 16, 2026; Databricks ruled it the "user's responsibility to ensure that uploaded skills do not contain malicious content."

  • Genie's egress control fails for a reason worth internalising. The exfiltration leaves from the user's browser when the result renders, not from the coding environment — so an egress policy scoped to the sandbox never sees it. If your agent renders anything, your threat model has a second network edge.

  • Beam's efficiency claim rests on an estimate, not a measurement. Reflection computes FLOPs as roughly 2 × active parameters × mean generated tokens, excluding prefill, attention and serving overhead, and says so plainly. Useful for ranking, not for forecasting your bill.

  • d1 is free to try without an integration. Both console.liquid.ai and the d1 Playground run the live demos, including the SQL-style filter that drops d1 into a WHERE clause over support tickets.

  • Together Link is one command and reversible. Settings and logins stay as they are, and going back to the native closed model is a supported path rather than a migration — which makes it cheap to measure the saving on your own workload rather than trusting the 50% figure.

  • What we dropped: a claim circulating that Together Link reports your savings per session. We could not confirm it on Together's own announcement page, so it stays out until we can.

One thing before you go

What's the one number in your stack that changed this week? Reply and tell me — I read every one, and the next issue gets better because of it.

If this was useful, forward it to the person on your team who owns the inference bill.

— The Agent Company