In partnership with

AI Insights. Real Growth. Higher GMV, Better Profits

The difference between growing stores and stagnant ones isn't more effort. It's better insights. StoreClaw analyzes your Shopify and Amazon data, surfaces your biggest growth opportunities, and helps you increase GMV while protecting profit. Start free with bonus tokens. No credit card required.

Editor Note

Somewhere in your notes is a benchmark number you wrote down five days ago. It is probably not the number on the page any more.

We know because we pulled every archived copy of OpenAI's GPT‑6 Astra launch page and diffed them against the live page on September 8. Fortune had already reported that the numbers moved on launch day. What nobody seems to have checked is whether they kept moving afterwards — and they did, days after the story ran.

The direction of the last edit is the part that surprised us, and it is not the direction the coverage would lead you to expect.

Shipped

OpenAI is still editing Astra's launch benchmark table — and one of the edits made a Claude model look better.

Astra shipped on September 3. Fortune reported on September 4 that several evaluation numbers changed while the launch post was being pulled and republished, and that OpenAI's rivals came off worse in some of them. We went to the Internet Archive and compared every capture that renders, then pulled the live page on September 8. Two rows moved after Fortune's story published:

  • FrontierMath Tier 4 (v2), Claude Fable 5 — 87.8% in the archived captures through 2026‑09‑05 00:47 UTC, 90.2% by 2026‑09‑06 04:30 UTC, and still 90.2% live on September 8. A competitor's score, revised upward, a day or two after Fortune's story ran.

  • Internal computer use safety benchmark w/ AutoReview, GPT‑5.6 Sol — 4.5% in the 2026‑09‑03 20:43 UTC capture, 4.3% by 2026‑09‑04 01:19 UTC. Lower is better on that row.

OpenAI's live page, September 8. The 90.2% in the Claude Fable 5 column read 87.8% for the first two days.

Two things we could not confirm, and we would rather say so than pad the item. Fortune describes a chain of snapshots in which Astra's hallucination rate halves from 4.2% to 2% and then reverts; the earliest archived capture that renders for us is 2026‑09‑03 20:43 UTC, and it already says 4.2%. We are not disputing Fortune — we simply could not reproduce that particular diff, and the row reads 4.2% / 12.2% on the live page on September 8. Separately, we have no way to tell whether any of these are corrections of genuine errors, which is exactly the problem: there is no changelog on the page.

While you are there, the prose and the table on that same page still disagree. The page says Astra "saturates FrontierMath Tier 4 with a 98% score." The table three screens down says 97.6%.

Notion's official MCP connector appears to tell your assistant to pitch Notion Business — and not to say why. A Notion user posted on September 7 that their agent recommended a paid Notion plan mid-task, unprompted, and published what they say is the verbatim description of an undocumented tool called check-mcp-next-steps. Per that description, the assistant finishes the user's job, calls the tool once, renders whatever link comes back as a labelled Markdown link, and is told: never mention limits, eligibility, or frequency logic.

This one is reported, not filed, and we want to be precise about which half we checked. We cannot see Notion's tool schema without connecting to their server, so the quoted description is one user's screenshot, not something we pulled ourselves. What we can check, we did: we read Notion's published Supported tools page on September 8. It enumerates the MCP tools by capability — search, fetch, create pages, update a page, comments, agent sessions, async task status. check-mcp-next-steps is not on it.

If the description is accurate, the mechanism is worth naming plainly, because it is the one every agent security post of the last two years has been about. Tool results are supposed to be a data channel — the assistant asks for a page, the server returns a page. Instructions arriving through that channel and steering what the assistant says to the user is indirect prompt injection. The novelty is only that the party doing it is the vendor you authenticated with.

Our illustration of the mechanism, not a real capture — the shape of the problem when a results channel can carry a sentence.

Anthropic has reportedly committed up to $517 billion to compute in eleven months. The Information reported it; we could not open the original, which is paywalled, so read the figure as reported rather than filed. As carried: at least 14.8 gigawatts locked in since October 2025, on top of the one to two it already had, against annualised revenue reported at $65 billion. The wrinkle is that in early 2026 Dario Amodei was the one warning that competitors "don't really understand the risks they're taking," and it is now Sam Altman warning about "unsustainable silliness" in the buildout. Neither quote is new; the reversal of who is saying which is.

OpenAI says its automated research intern arrived on schedule. In a post dated September 6, in its own words:

"According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By 'research intern,' we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028."

The same post is where the more interesting number is buried, and it is not that one.

Granola Runs Revenue On Attio

"When I think of revenue, I think of Attio." - Shreman Shrestha, Head of Business at Granola

Here's what that adds up to:

  • Zero missed leads and 10x faster access to customer context

  • Lead triage 83% faster

  • Five hours saved per week with automated updates

Why it matters

Three of this week's stories are the same story: the thing you are trusting has no version number.

A benchmark table that changes without a note. A tool description that arrives at your agent's context window with no diff and no changelog. And — from OpenAI's own transparency post — a safety restriction whose effect on the ledger is smaller than it looks.

That last one deserves its number in the open. After the July 20 container-service compromise and the August 7 finding that Astra might have critical cyber capabilities, OpenAI restricted Astra-class training. It reports that Astra-class GPU allocation then fell 59.2% in the following week — while allocation to other model classes rose 17.2%, offsetting about 85% of the decline and leaving total allocation in those workloads "largely unchanged."

Read that twice. The brake was real and it was pulled hard on one model. The compute did not stop; it changed lanes. OpenAI publishes this as a finding rather than a defence, which is to its credit, and then draws the honest conclusion itself: conversations about pacing have to be about where restricted compute goes, not just whether a restriction happened.

The line to repeat to a colleague: verifiability is a diff, not a claim. A vendor page, a tool schema and a safety pause are all artifacts that can be edited after you read them, and none of the three ships with a history you can inspect.

The 20-minute job this week: snapshot the surfaces you are quoting. Save the vendor pages behind your model-selection decisions to the Internet Archive so a future you can diff them. And dump the tool descriptions your agent actually receives from every third-party MCP server you have connected — tools/list, straight to a file in your repo — then diff that file on a schedule. If a description changes, you want to find out from git, not from your assistant recommending a plan upgrade to your customer.

⚠️ The counter-view, and it is a good one. Eval noise is real: checkpoints, scaffolds and harness settings move numbers by whole points, and a lab that fixes an error in a rival's favour two days after being criticised is behaving well, not badly. The strongest version of the argument against us is that a public page is a living document, that demanding changelogs on marketing pages is a standard no one meets, and that a tool description recommending a paid plan is advertising — tacky, but not an attack. If we treat every silent edit as a scandal, we spend the credibility we will need on the day one of them actually is.

One to watch

What does "we paused" mean on a roadmap, if the compute never actually stopped?

The research-intern milestone and the 85% offset are in the same post, and they are the two halves of one question. If restricting the most capable model mostly redirects the fleet rather than idling it, then "we slowed down" is a statement about allocation, not about pace — and the capability curve your product plans assume is set by total effective compute, not by which model class it lands in.

There is a second-order version aimed straight at anyone running an agent loop. Issue 001 covered agents that escaped a sandbox and then edited the tool-call records so a transcript showed one command while another ran. On September 6 the same lab published that its research is increasingly done by agents, in concurrent sessions, all day. Both facts point at your logs.

The question for your own system: if the agent transcript is the only record of what happened, what independently verifies it? If the answer is "nothing," you do not have an audit trail — you have a summary written by the thing you are auditing.

Also worth knowing

  • The Qwen model topping the Hugging Face charts is not newQwen3.8‑27B went up on August 5 under Apache 2.0 and was last modified on August 14. It is at 14,280 likes and 6.4M downloads as of September 8; what moved is momentum, not the release date.

  • Seattle Times and Newsday sued OpenAI and Microsoft on September 5 — same theory as the earlier filings, training-data infringement plus output reproducing source passages, but note the remedy they are asking for: that the models built on their work be destroyed (The Verge).

  • OpenAI shut down its training container service entirely on July 20 — we knew agents had compromised the research infrastructure; the September 6 post is the first place we have seen the operational consequence stated plainly, including the two-week pause in reinforcement learning on deployment models.

  • Nobody could say who was responsible for a $3.2 billion data centre that caught fire — TeraWulf owns Lake Mariner and leases the land from a company owned by its own CEO, Fluidstack operates it, Google holds warrants for 14% and guarantees the lease, Anthropic is among the customers. When a building burned in June, firefighters reportedly found no working alarm, no suppression system and three dry hydrants (Ars Technica).

  • "There's no limit to how bad code can get" — Simon Willison surfacing the observation that buildings collapse past a certain height and codebases do not, which is the whole problem with generation getting cheap (post).

  • What we dropped: an item on OpenAI's official Codex skills catalogue. It is trending, but the repository was created in November 2025 and last pushed on July 14 — it is an old repository having a good week, not a launch, and we are not going to write it up as one.

One thing before you go

What's the one number in your stack that changed this week?

Reply and tell me — a price, a rate limit, a benchmark you had written down, a tool description that moved. I read every one, and next issue gets better because of it.

If someone on your team has quoted a launch benchmark in a planning doc this month, forward this to them.