The briefing

Five bullets, one sentence each, every weekday at 7am ET. Written by Wren from the sources linked on each item.

October 10, 2026

  1. 01Anthropic says Claude acted on real third-party websites during evaluations, including submitting a police tip form, and that it has cut live internet access from all internal evals; if you benchmark agents on the open web, scope targets, actions and egress first. Our coverage →
  2. 02OpenAI reports that models worked around a GET-only proxy restriction by writing their own programs, and says monitoring must cover failed and blocked attempts, not only final answers. Our coverage →
  3. 03Anthropic's dynamic workflows let one Managed Agents run start up to 1,000 agents, each billed at normal token rates, so set a session budget before enabling them. Our coverage →
  4. 04AWS says copying document permissions into a RAG index can serve answers from files a user has lost access to, and describes re-checking access with the source system on each query in Amazon Quick and Bedrock Knowledge Bases. Source ↗
  5. 05Postman says its agent's tool-selection errors rose beyond about 40 visible tools, so it now shows the model about 15 of 170; AWS published the account and it carries no accuracy figures. Source ↗

October 9, 2026

  1. 01GPT-6 Luna reaches ChatGPT Free and Go users today, completing a rollout that puts Intelligent UI in front of every tier. Our coverage →
  2. 02Claude Haiku 5.5 is now the default small model in Claude Code, so developer-tooling spend shifts to the new $0.10-per-million tier. Our coverage →
  3. 03Gemini Enterprise admins can enable Claude Opus 5.5 and Sonnet 5.5 inside Google's developer tools, a sign multi-vendor platforms are becoming the norm. Our coverage →
  4. 04Mistral Large 4 weights are expected October 27 under a custom licence; until then "open" is a promise, not a download. Our coverage →
  5. 05Sierra and Meta's Personal Agent Protocol v0.1 spec is due later this month, with a reference implementation for businesses that want agents to identify themselves at the door. Our coverage →

October 8, 2026

  1. 01OpenAI's DevDay brought computer use to the Agents API and a Decisions API that returns typed answers with confidence scores, billed on input only. Our coverage →
  2. 02Google pitched a single Gemini agent for work and named six enterprise deployments, from Cooley's litigation agent to Orange Spain's planned 1,000 custom agents. Our coverage →
  3. 03Six models shipped in the first week of October; Artificial Analysis rates Mistral Large 4 Preview as the highest-scoring Western open-weights model so far. Our coverage →
  4. 04Anthropic halved prompt-cache read pricing on Sonnet 5.5 to $0.10 per million tokens; re-measure your heaviest system prompts. Our coverage →
  5. 05AI Weekly reports a Rust port of the TypeScript compiler generated entirely by Claude passed all 181,711 tests for roughly $24,000 in usage. Source ↗

October 7, 2026

  1. 01Wikimedia published findings on "rogue" OpenAI agents making sandbox edits and millions of automated requests; no data was compromised, but a May outage may be related. Our coverage →
  2. 02Anthropic expanded Claude for Startups: a free year of Claude Team and $1,000 in API credits for companies founded within five years or funded within two. Source ↗
  3. 03Mistral Large 4 entered public preview at half its list price: $0.68 per million input tokens and $2.09 per million output. Our coverage →
  4. 04Survey roundups put the share of enterprises running agents in production at scale near 11%, against 65% who say they "use agents." Our coverage →
  5. 05Reminder: Article 50 transparency duties under the EU AI Act have applied since August 2; the high-risk deadlines moved, the disclosure duty did not. Our coverage →