Should You Still Be Renting All Your AI?


Your Weekly AI Briefing for Leaders

Welcome to this week's AI Tech Circle briefing: clear insights on Generative AI that actually matter.

This week, I looked at numbers instead of demos, and two of them stuck with me. Open models now carry more than half the traffic on one major AI gateway, up from 7% last December. And one of the best-funded labs on earth just told the markets it plans to spend over half a trillion dollars on compute in the coming decade. Someone is going to own AI infrastructure, and someone is going to rent it. This week's Deep Dive is about deciding which one you should be, workload by workload.

Today at a Glance:

  • Executive Brief: OpenAI shelves its next model as the regulators arrive, Google launches Gemini 4 Argon, Anthropic heads for a $2 trillion IPO
  • Deep Dive: Own the Baseload, Rent the Frontier — when open-weight models in your own infrastructure beat the API
  • Weekly News & Updates
  • Favorite Tip of the Week
  • Generative AI Use Case of the Week
  • Podcast
  • Courses and events to attend
  • Tool / Product Spotlight
  • The Investment in AI

Executive Brief

The Model That Won't Ship: OpenAI Pulls GPT-6.1 Astra as the Regulators Arrive

Last week we covered the sandbox escape; this week the consequences arrived. OpenAI scrapped the planned release of GPT-6.1 Astra, citing safety concerns, a first for a flagship model this close to launch (Bloomberg). The FTC then opened a broad investigation into OpenAI and Anthropic over consumer risks; California's Attorney General served OpenAI an investigative subpoena tied to the agent incidents, and six CEOs - OpenAI, Anthropic, Google, Meta, xAI, and Nvidia; signed a voluntary "Joint Commitment" at the White House (AI Weekly). Inside OpenAI, three safety researchers were dismissed, and its safety-transparency lead departed.

Why this matters to you: two assumptions broke this week that announced models ship, and that AI oversight stays voluntary. Ask your team which roadmap items quietly depend on a model that hasn't shipped yet, and start an evidence file of how you test and control the AI you already run; "show us your safety work" is now a procurement and regulator question, not a lab question.

Google Launches Gemini 4 Argon and Reprices the Frontier

Google released Gemini 4 Argon, its new frontier model, at an introductory $2/$10 per million tokens with a 1-million-token output ceiling and a 77.9% score on DeepSWE. Alphabet rose on the launch, and the pricing undercuts rivals at the top end of the market.

Why this matters to you: the unit economics of every AI feature you run are now repriced roughly quarterly. Before you renew or switch anything, benchmark on your own tasks and make sure your contracts let you follow the price curve down rather than locking in last year's rates.

Anthropic Heads for the Public Markets at $2 Trillion

Anthropic is targeting a November IPO at a valuation around $2 trillion, which would be the largest listing in history. A leaked prospectus shows steep losses, rapid growth, and a commitment of roughly $518 billion to compute over the next decade, alongside $42 billion in convertible-note financing from Broadcom tied to chip leases. The same week, a BIS study found 55.2% of AI funding from 2021–2025 came from other AI firms.

Why this matters to you: your key AI vendors' balance sheets are becoming public-market questions, and over half the sector's funding is circular. For each AI vendor you depend on, know how it is funded and what happens to your contract and your data if capital gets expensive. That's a due-diligence line item now, not paranoia.

Deep Dive: Own the Baseload, Rent the Frontier

Every AI conversation I had this year started with "which model?" The better question, and the one your CFO will eventually ask, is "which models do we rent, and which do we own?" Utilities solved this long ago: you run cheap, predictable capacity for your baseload and buy expensive peak power only when demand spikes. This week produced the clearest evidence yet that AI has reached the same split.

The rule: Own the Baseload, Rent the Frontier. Steady, high-volume, well-specified work, classification, extraction, summarization, routing, translation, is baseload, and it is a candidate to run on open-weight models in infrastructure you control. Novel, hard, spiky work - frontier reasoning, complex agents, anything you do rarely but need done brilliantly - is peak, and you rent it from the frontier labs.

Three pieces of evidence from the last seven days:

The demand side already moved. Open models now handle 56% of tokens on Vercel's AI Gateway, up from 7% in December, and run about 40% of AT&T's AI workloads, with enterprises reporting inference cost cuts of roughly 50%. This isn't hobbyists; it's production traffic migrating to the cheap end of the curve.

The supply side got serious. NaiveAI open-sourced a 309-billion-parameter mixture-of-experts model (15.5B active parameters, 1M-token context) under an MIT license, Perplexity open-sourced its 27B Decisions model at $0.04 per million input tokens, and IBM gave its Bob platform a self-hosted deployment option for regulated enterprises. The gap between "open" and "enterprise-grade" is closing from both directions.

Good enough is now measurable and absurdly cheap. The new StudentBench evaluation found AI tutors matching expert human GRE tutors, and an open 31B model delivered comparable learning gains at roughly 918x lower cost than the frontier setup. For a well-defined task, the quality gap between open and frontier has collapsed faster than most procurement cycles.

The honest counterpoint: frontier prices fall too; Gemini 4 Argon launched at $2/$10 this very week. So the case for owning is not only cost. It is control: predictable spend at volume, data that never leaves your perimeter, and permanence. The other lesson this week is that you can scrap a rented model before it ever ships. You can't take an open-weight model you host away from yourself.

You can try doing this:

  1. Find your baseload. Pull last month's AI and API bills and list your three highest-volume recurring tasks. If they look like classification, extraction, summarization, or routing, they're baseload shapes.
  2. Run one honest benchmark. Take one of those tasks, collect 50–100 real examples with your own acceptance bar, and test one open-weight model against your current API. Judge on your pass rate, not on vibes or leaderboards.
  3. Price it three ways. Current API vs. a managed open-model endpoint vs. self-hosted - including people time. Then write down your repatriation trigger: the monthly spend at which owning beats renting. When the bill crosses the line, you'll already have the decision made.

Weekly News & Updates...

  • Google, OpenAI and Anthropic are planning a joint AI safety standards body, a private standards authority for frontier AI. Why it matters: standards written by your three biggest vendors will become the default procurement baseline; track the drafts rather than discovering them in a contract.
  • Anthropic reported Claude computed a six-particle, nine-loop scattering amplitude in planar N=4 super Yang–Mills, one loop past the published record, verified by SLAC physicist Lance Dixon. Why it matters: frontier models are starting to extend research frontiers, not just summarize them - worth a pilot in your most R&D-heavy function.
  • Apple tightened Full Disk Access on macOS, citing AI agents becoming "more capable and autonomous." Why it matters: OS vendors now treat agents as a threat class; any desktop-agent rollout you plan will need explicit permission engineering, not default access.
  • Italian bank Fideuram lost €95 million to an AI voice-clone fraud, with about €39.5M still missing. Why it matters: a familiar voice is no longer authentication - put callback verification on every payment instruction above a threshold, this week.
  • arXiv capped submissions at two per month per author after volume quadrupled in a decade. Why it matters: AI-accelerated output is forcing rate limits on human review - your internal approval processes will meet the same flood.

Favorite Tip of the Week

Run your first local open model - in under 15 minutes. If this week's Deep Dive is a decision your organization should make, this tip is how you build the instinct for it personally.

  1. Download Ollama (Mac, Windows, or Linux) and install it - two minutes.
  2. Open a terminal and type "ollama run gemma3" (or pick any small open model from Ollama's library page). It downloads a few gigabytes, then drops you into a chat that runs entirely on your laptop; no account, no internet after the download.
  3. Give it one real but non-sensitive task from your week: rewrite an email, summarize pasted meeting notes, draft a job description.
  4. Note what was good enough and what wasn't. That gap - felt firsthand- is your own baseload/frontier line, and you'll never read a model announcement the same way again.

Generative AI Use Case of the Week

Banking: Barclays scales Claude from pilot to utility. Barclays expanded Claude to more than 16,000 internal users, who now run about a million internal searches a month; its Global Markets division routes 120,000 emails a day through the system, and the bank targets half its developers on Claude Code by year-end. The pattern worth copying: start where volume is huge and quality is measurable (search, email triage), publish the usage numbers internally, then expand into developer tooling; and measure adoption per workflow, not by seats bought.

The Opportunity...

Podcast:

Why Companies Need to Own Their AI with James Lang: What is the business value of private AI, and why should companies consider owning the AI systems that connect to their internal data and operations? In Episode 202 of Open Tech Talks, host Kashif Manzoor speaks with James Lang, an operator, investor, and technology strategist at Overlang Venture Partners, about private AI, business growth, AI security, legal technology, business valuation, and the future of work.

show
Why Companies Need to Own Th...
Oct 4 · OPEN Tech Talks: AI wort...
19:16
Spotify Logo
 
​

​
​


Courses to attend:

  • ​Open Source Models with Hugging Face (DeepLearning.AI): how to find, evaluate, and run open models; the hands-on companion to this week's Deep Dive and the fastest way to run the benchmark in step 2.
  • ​Efficiently Serving LLMs (DeepLearning.AI): the serving-side economics; batching, quantization, KV caching; that decide whether self-hosting your baseload actually saves money or just moves the bill.

Events:

Tool / Product Spotlight

  • ​GitHub Copilot computer use (public preview): Copilot can now operate desktop applications - clicking, typing, and navigating the way a person would. Worth testing if you want to understand what desktop agents will do to your workflows before your employees quietly find out on their own.
  • ​Mercury Voice (Inception Labs): a voice model with 320ms latency at $0.20/$0.75 per million tokens. Worth testing if you run a phone or support channel where today's voice bots feel one beat too slow; that beat is what callers hang up on.

The Investment in AI

  • Instinct raised a $1 billion Series C at a $10 billion valuation, led by Sequoia, Benchmark and Coatue; four times the valuation it carried just one month earlier, with no disclosed user numbers. What the money says: consumer agents that act; book, pay, cancel, call - are being priced as the next consumer platform. Expect some of your customers to start arriving via their agents; make sure your booking and support channels don't break when a machine calls.

That's it for this week - thanks for reading!

Reply with your thoughts or favorite section.

Found it useful? Share it with a friend or colleague to grow the AI circle.

Until next Saturday,

Kashif

The views expressed here are solely my own, based on my experience, practice, and observation, and do not necessarily represent those of my current or previous employers or their clients and customers.

AI Tech Circle

Learn something new every Saturday about Generative AI #AI #ML #Cloud and #Tech with Weekly Newsletter. Join with 592+ AI Enthusiasts!

Read more from AI Tech Circle

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing: clear insights on Generative AI that actually matter. This week, the AI agent problem became real. OpenAI paused tool-use work on its most capable models after an agent found a way out of its sandbox through the one channel nobody thought to lock: DNS. A Government found out three months late that an agent had been inside one of its portals. And researchers showed agents turning to hacking techniques while...

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing: clear insights on Generative AI that actually matter. Today at a Glance: Executive Brief: The "Pacing" Debate Goes Global and Hits Markets, Google Opens the Home to AI Agents, Anthropic Discloses an Incident Deep Dive: The Build-Before-Delegate Rule - How to Use AI Without Losing the Skills Your Judgment Depends On Tip of the Week Podcast Courses and events to attend Tool / Product Spotlight Executive Brief...

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing: clear insights on Generative AI that actually matter. I have started observing that large enterprises, where there is the most money at stake, have stopped asking, "How impressive is the AI?" and started asking, "What does each use case of it actually cost?" That question is now moving down the AI stack. Within a year, your CFO will ask the same thing about the AI Agents your team deployed. And in most...