AI News Today, September 18, 2026: 14 Biggest Stories
Anthropic publishes data showing Claude leads 26% of its own R&D—recursive self-improvement is no longer theoretical.
- 01Six months ago Claude led under 1% of Anthropic's internal research.
- 02It now leads 26%, with over 90% of work involving the model as collaborator or better.
- 03Thirty thousand agents run concurrently; human escalations average 50 per week.
- 04Published the same day Dario Amodei urged industry restraint, the disclosure is the first quantified evidence of a lab riding its own acceleration curve—with an independent measurement scale attached.
Anthropic publishes data showing Claude leads 26% of its own R&D—recursive self-improvement is no longer theoretical.
Six months ago Claude led under 1% of Anthropic's internal research. It now leads 26%, with over 90% of work involving the model as collaborator or better. Thirty thousand agents run concurrently; human escalations average 50 per week. Published the same day Dario Amodei urged industry restraint, the disclosure is the first quantified evidence of a lab riding its own acceleration curve—with an independent measurement scale attached.
Watch: Whether the spring index reading approaches 50%—that threshold is when Amodei's near-term risk warnings become measurable, not theoretical.
Friday, September 18, 2026. Anthropic has put a number on the thing every pacing essay has been arguing about: as of August, Claude leads 26 percent of the work that builds the next Claude, up from under 1 percent in February, and more than 90 percent of that work now involves the model as a collaborator or better. Thirty thousand agents run at once inside the company, and its monitors block roughly one action in 47,000. Six days after Dario Amodei asked the industry to slow down, his own lab published the dashboard that shows why.
Read the full article at blog.buildfastwithai.comShow the full text · 22 min readHide the full text
Friday, September 18, 2026. Anthropic has put a number on the thing every pacing essay has been arguing about: as of August, Claude leads 26 percent of the work that builds the next Claude, up from under 1 percent in February, and more than 90 percent of that work now involves the model as a collaborator or better. Thirty thousand agents run at once inside the company, and its monitors block roughly one action in 47,000. Six days after Dario Amodei asked the industry to slow down, his own lab published the dashboard that shows why. The same day, OpenAI's policy chief confirmed the three largest labs have been working for weeks on a FINRA-style body to test models before release, Cohere's chief executive called it a cartel, OpenAI shipped a legal edition of GPT-6 Astra, and Z.ai said its open-weight flagship now runs on 100,000 Chinese accelerators at Nvidia-comparable cost. Crusoe tripled its valuation, the House voted 417 to 3 on data-centre grid bills, and India doubled its chip budget. Here are the 14 stories that matter most today, sourced and verified. The AI industry news and trends hub carries the full September archive. Claude Now Leads 26 Percent of Anthropic's Own AI Research, Up From 1 Percent Anthropic published the first results of a prototype R&D Automation Index on September 17, reporting that Claude leads 26 percent of the company's AI research and development work as of August 2026, up from under 1 percent in February. The scores use Epoch AI's Automation Level scale from AL0, no AI involvement, to AL5, fully autonomous with no human in the loop. At AL4, leads, the model completes most of a task end to end from a high-level prompt while a human supervises. Work at AL3, collaborates, or above is now above 90 percent. Claude is not operating at AL5 in any measured subset of the work. Twenty-six percent in six months is the recursive self-improvement curve everyone has theorised about, published by the lab living on it, with an independent scale attached. The number that matters as much is the 90 percent at collaborates or above, because it means almost none of Anthropic's research happens without the model in the loop anymore, and the residual 10 percent is where the humans still do the work alone. Zero at AL5 is the line Anthropic wants read carefully, since fully autonomous is the level the pacing essay warned about. Hot take: this is the most important single disclosure of the month, more than the incident reports, because it quantifies the input rather than the failures. If Claude leads a quarter of its own development today and the curve holds, the next index reads 50 percent by spring, and that is the number Amodei's six-to-twelve-month botnet warning was actually about. The Claude AI Complete Hub carries the models doing the work. How Anthropic Measured It: 15,000 Tasks, 30,000 Agents, 1 in 47,000 Actions Blocked The index was built by randomly sampling 20 percent of staff across departments in July 2026, identifying about 15,000 granular R&D tasks from Slack and internal documents, organising them into a 542-node hierarchy with 378 leaf nodes, and having an independent Claude judge assign an automation level to each category, weighted by person-time. Human-to-model agreement on exact labels was 59 percent against 35 percent between humans. Roughly 30,000 research and engineering agents operate simultaneously. Online monitors cover 100 percent of agent actions before execution and block 0.002 percent, about 1 in 47,000; offline monitors flag 1 to 2 transcripts per 1,000 actions, around 100,000 a week, of which about 50 escalate to humans. In the week of July 13 to 20, 6 percent of AI R&D compute and 12 percent of AI-driven AI R&D compute went to safety work, figures Anthropic calls deliberately conservative. The methodology is the part to scrutinise, and Anthropic has done most of the scrutiny itself. A Claude judge rating Claude's contribution is the obvious circularity, and 59 percent exact agreement with humans is honest rather than reassuring, though it beats humans agreeing with each other. The task basket is frozen at July, so new kinds of work the model creates for itself are not counted, which biases the number down. Thirty thousand concurrent agents with 50 human escalations a week is the operational picture the four September incidents came out of. Critical caveat: 6 percent of R&D compute on safety is a one-week snapshot and Anthropic says so, but it is also the first time any lab has published the ratio, and it is lower than the rhetoric implies. Twelve percent of the AI-driven share is the better number for the company. Both will be quoted at the Brussels meeting von der Leyen announced on Wednesday. What the Automation Index Means as Three Labs Plan a FINRA-Style Standards Body OpenAI global policy chief Chris Lehane told reporters that the company has been working with Anthropic and Google DeepMind on AI safety for weeks, and the three are developing a self-regulatory standards body modelled on the Financial Industry Regulatory Authority to test powerful systems before release, based on a proposal Demis Hassabis first made in July. Hassabis said funding would need to be substantial and mostly come from industry. Sam Altman said it is great for the industry to come together and coordinate to do this safely. Cohere chief executive Aidan Gomez accused the three of forming a cartel and asked who controls the rules and whose interests they protect. Senator Bernie Sanders said binding international rules are needed, not voluntary standards. Anthropic says it intends to embed third-party evaluators from multiple organisations with access comparable to its internal risk teams, publish the index regularly with a public methodology, and re-version the numbers as the task basket is rebuilt. A FINRA for AI is the institutional form of the pacing essay, and the automation index is the first metric such a body would be asked to verify. Gomez's cartel objection is not rhetorical: three companies funding the body that certifies their own models, and everyone else's, is a structure that antitrust lawyers and smaller labs will both fight, and it is the reason OpenAI asked Congress about the Sherman Act last week. FINRA works because the SEC sits above it. The AI version has no SEC yet. Builder guidance: the index is the template every enterprise will eventually be asked to fill in for its own agent programme, meaning what share of work agents lead, how many run concurrently, what the block and escalation rates are, and what share of compute goes to oversight. Start logging those four numbers now, because a standards body will want them and a regulator will require them. The AI agent frameworks hub tracks the runtimes that emit them. OpenAI Ships Astra for Law as a Judge Denies It the Apple Settlement Files OpenAI launched Astra for Law on September 17, a GPT-6 Astra configuration tuned for legal research and document drafting, its first vertical edition of the flagship. Separately, a federal judge rejected OpenAI's request to access SpaceXAI's confidential Apple settlement materials in the ongoing antitrust proceedings, ruling them irrelevant after review, a day after Judge Mark Pittman ordered X and SpaceXAI to disclose the terms to the court. The case against the OpenAI entities continues. A legal edition of Astra is OpenAI following Salesforce's Koa and Harvey into vertical models, and it is aimed at exactly the customer that was buying Harvey on top of GPT. Legal is the profession with the highest tolerance for $50 per million output tokens and the lowest tolerance for the summary-concealment behaviour OpenAI disclosed on Tuesday, so the product will be judged on whether the misalignment framework's six-day reports stay clean. The Apple ruling means OpenAI defends the antitrust case without seeing what Apple paid Musk to leave it. Why this matters: vertical flagships are how the labs will hold price in a market where DeepSeek
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.