Are Your AI Agents Actually Working? Five Numbers That Answer It


Your Weekly AI Briefing for Leaders

Welcome to this week’s AI Tech Circle briefing: clear insights on Generative AI that actually matter.

I have started observing that large enterprises, where there is the most money at stake, have stopped asking, "How impressive is the AI?" and started asking,

"What does each use case of it actually cost?"
That question is now moving down the AI stack. Within a year, your CFO will ask the same thing about the AI Agents your team deployed. And in most organizations I work with, nobody can answer it right now.

I've been in enough of those meetings to settle on an answer. Read the detailed article below...

Today at a Glance:

  • Executive Brief: Nvidia Buys Hugging Face, GPT-6 Astra Goes Broad, Both Frontier Labs Ask for the Brakes
  • Deep Dive: The Agent Scorecard, Five Numbers That Tell You If Your Agents Are Paying Off
  • Tip of the Week
  • Podcast
  • Courses and events to attend
  • Tool / Product Spotlight

Nvidia Now Owns the Front Door to Open-Source AI

Nvidia is acquiring Hugging Face for $12.93 billion, the platform that hosts more than 3 million models and is used by over 18 million developers. It dominated open-ecosystem discussion all week. Jensen Huang has pledged to keep the platform open and hardware-neutral.

Why this matters to you: if your teams download models, datasets, or tooling from Hugging Face, and most do, whether IT knows it or not, the distribution channel for open AI now belongs to the world's dominant chip vendor.

GPT-6 Astra Reaches Everyone This Week, and It Uses Your Computer

OpenAI's GPT-6 Astra's headline capability is computer use: navigating applications, filling forms, updating CRM records, managing calendars, the model operating a computer the way a person would. OpenAI claims state-of-the-art results across computer use, coding, cybersecurity, and science, and its leadership is openly using the phrase "AGI era."

Why this matters: computer-use agents are moving from research preview to mainstream product, which means your desktop applications and internal tools will soon become surfaces that AI acts on directly.

Two practical notes.

First, treat "AGI" as marketing and judge the model on your own tasks.

Second, watch the economics carefully; it's reportedly around 2.5x pricier per token than its predecessor.

The People Building the Frontier Just Asked Everyone to Slow Down

In the same week, OpenAI's chief scientist Jakub Pachocki published an essay arguing that voluntary slowdowns should become normal until shared safety standards exist, and Anthropic's CEO published one asking the world to slow AI down, proposing outside experts to check labs' work and like-minded countries coordinating on standards. The Anthropic essay landed days after one of its employees publicly resigned over how the industry handles risk.

Why this matters beyond the labs:

When the builders themselves call for external audits, external audits become likely.

Expect "show us your safety work" to move from a nice-to-have into buyer procurement checklists and regulatory expectations for anyone deploying frontier AI, not just building it. The organizations that can already document how they test and control their AI systems will be ready. The rest will be scrambling.

The Agent Scorecard - 5 Numbers That Tell You If Your Agents Are Paying Off

Most teams measure AI agents the way they'd measure a website: usage. How many tasks it ran.

How many people used it.

How many hours it "saved," estimated by someone who wanted the project approved.

Usage is not value.

An AI agent that runs a thousand tasks a week, of which a third get quietly redone by a human, isn't saving anyone anything; it's adding a review step to work that used to be done once. The estimated hours saved almost never offset the hours added to check and correct the output. And the tasks that failed badly enough to become an incident are usually tracked by a different team, in a different system, and never make it into the productivity slide.

An agent is paying off when cost per accepted output is below the cost of a human doing the same work, net hours are positive, and the exception rate hasn't risen since deployment. All three. Two out of three is a project to fix, not a success to scale.

This rule protects you in both directions. It stops you from scaling an agent that looks busy but is quietly costing you. And it lets you defend an agent that's genuinely working when someone senior decides AI is overhyped, because you'll have five numbers and they'll have an opinion.

  • ChatGPT, Claude, and Grok All Went Down Within the Same Few Hours: On September 3, OpenAI, Anthropic, and xAI each confirmed service interruptions, and Gemini saw a surge in reported failures at the same time. Why this matters to you: a lot of teams believe they're covered because they use two or three LLM providers. This week showed the providers can fail together. What to do: run a thirty-minute drill this month, pick your most important AI-dependent workflow and ask what actually happens for the next four hours if every model is unavailable.

Favorite Tip Of The Week:

Before You Let an AI Use Your Computer, Give It Its Own Account

With computer-use agents reaching mainstream users this week, here's the single most useful precaution I can offer, and follow it.

Don't run an AI agent inside your main user account. Create a separate account or browser profile just for the agent, with only the applications and folders it needs for the task, no saved passwords, and no access to your primary email or banking.

Potential of AI:

The most interesting research detail of the week is an evaluation, not a model. Informed by the Hugging Face incident, OpenAI built a test for whether a model facing a difficult or impossible task will go beyond its authorized scope to complete it. Without production safeguards, GPT-5.6 Sol exceeded the authorized target 48% of the time. GPT-6 Astra did so in 0% of cases; it recognized when finishing would mean overstepping and returned to the user instead.

Why this matters for you: it's a new category of evaluation you can borrow. Give your own agents a task they can't complete within their permitted scope and watch what they do. An agent that stops and asks is one you can trust with more. One that improvises around the boundary is one you can't. OpenAI | VentureBeat

AI in Business Tip:

Before You Buy Any AI Agent, Ask One Question: Can I Get the Log?

As agents move from answering questions to taking actions, filling forms, updating records, and sending messages, the most important thing a vendor can give you isn't a better model. It's a complete record of what the agent did. Which systems it touched, what it read, what it changed, and what it sent. Most AI products today give you a chat transcript.

Why this matters now: when an agent makes a mistake in a customer record or an outbound email, "what happened?" is the first question your customer, your auditor, and your lawyer will ask.

The Opportunity...

Podcast:

  • Open Tech Talks Podcast Episode 199: "The Future of Multi-Model AI Workflows with Jonathan Archer
    How can developers use multiple AI models without constantly switching tools or hitting rate limits?

Apple | Youtube

show
The Future of Multi Model AI...
Sep 13 · OPEN Tech Talks: AI wort...
26:33
Spotify Logo
 

Courses to attend:


Events:


Tech and Tools...

  • OpenVDN: An open video model that generated 14.4 seconds of video in 11.23 seconds on eight B200 GPUs. Real-time AI video is now technically here, if you have the hardware.
  • Nvidia PAIR: Routes local AI tasks across devices on a home or office network, worth a look if you're experimenting with on-device or hybrid AI.Grok Imagine Video 1.5 agent: xAI's upgraded video-generation agent, now running on its Image 2.0 model, relevant for creative and marketing teams tracking the generative video space.

The Investment in AI

  • Crusoe raised more than $3 billion at a roughly $30 billion valuation, backed by a reported five-year, $13 billion cloud contract with Jane Street, a sign that financial firms are now buying AI compute at hyperscaler scale. Crypto Integrated
  • Nscale is reportedly seeking $3.5 billion ahead of an IPO, including $2 billion from Nvidia, a chipmaker funding a cloud that buys its chips. Watch this "circular financing" pattern; it flatters demand numbers across the whole AI supply chain. AI Weekly

That's it for this week - thanks for reading!

Reply with your thoughts or favorite section.

Found it useful? Share it with a friend or colleague to grow the AI circle.

Until next Saturday,

Kashif


The opinions expressed here are solely my conjecture based on experience, practice, and observation. They do not represent the thoughts, intentions, plans, or strategies of my current or previous employers or their clients/customers. The objective of this newsletter is to share and learn with the community.

Dubai, UAE

You are receiving this because you signed up for the AI Tech Circle newsletter or Open Tech Talks. If you'd like to stop receiving all emails, click here. Unsubscribe · Preferences

AI Tech Circle

Learn something new every Saturday about Generative AI #AI #ML #Cloud and #Tech with Weekly Newsletter. Join with 592+ AI Enthusiasts!

Read more from AI Tech Circle

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing- clear insights on Generative AI that actually matter. I usually write this newsletter over the weekend, and I even kept it for a few weeks without writing it. I thought it was worth writing. And that's maybe procrastination; I was putting it off. And maybe the good reason is that with all the LLMs and Generative AI, all the knowledge and everything is there; whatever you want, you can type it, you just prompt...

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing- clear insights on Generative AI that actually matter. Executive Brief The Law Catches Up to the AI Six weeks ago, an OpenAI model escaped its test environment and hacked Hugging Face. This week, the consequences arrived. Alabama's attorney general issued a subpoena to OpenAI, part of a multi-state coalition of attorneys general demanding records, safety policies, and testing procedures and asking the company...

Your Weekly AI Briefing for Leaders Welcome to this week’s AI Tech Circle briefing- clear insights on Generative AI that actually matter. Last week in this newsletter, we discussed the AI architecture that keeps AI agents reliable in production and why the guardrails, the token budgets, and the kill switches are vital. This week, an AI model broke into a real company to rethink how you all are deploying AI agents into production. What Actually Happened During an internal evaluation, OpenAI's...