Wednesday, July 22, 2026

The Telco AI $60 Billion "Opportunity"

Google Cloud published a piece in RCR Wireless this morning arguing that agentic AI represents a sixty billion dollar opportunity for telecom operators. It is a well constructed argument, the engineering description is accurate, and the case studies are real. It is also, read carefully, an argument about cost avoidance wearing the vocabulary of growth. I want to be precise about this, because eight days ago I published a piece arguing that the industry's central discipline problem is its refusal to separate the two, and this article is the cleanest illustration of the problem I have seen since.

Start with the numbers. The sixty billion figure comes from Appledore Research, and Appledore is explicit about what it is measuring: operational cost savings by 2030. The McKinsey research cited alongside it reports a thirty to seventy percent reduction in troubleshooting tickets and a fifty five to ninety percent reduction in network operations centre costs. Deutsche Telekom's RAN Guardian identified 237,000 network events in early 2026 and compressed major incident handling from hours to about sixty seconds. Bell Canada's AI Ops platform achieved a twenty five percent reduction in customer reported issues and a faster mean time to repair. Vodafone's agents protect millions in annual operating expenditure. Every one of those is a genuine achievement. Not one of them is revenue. The article's own evidence base is, without exception, my first money flow: AI that reduces cost, which I described as real, happening, the largest near term financial impact of AI on operators, and emphatically not a new line of business.

The word doing the concealing is "opportunity". An opportunity, in the way a board hears it, is something you invest in to get money back that you were not getting before. A cost saving is something you invest in to stop spending money you were already spending. The two justify different capital, different organisational patience, and different governance. Operators that hear sixty billion and staff a growth programme will find, three years in, that they have built a very good efficiency programme and told their investors the wrong story about it. That is not a hypothetical failure mode. It is the failure mode the industry has run repeatedly, and the reason I keep insisting the flows be kept apart on the page before they are kept apart in the budget.

There is a second problem with the sixty billion, which is that it is sitting next to a Deloitte figure of a hundred and fifty billion in "total value" and the two are quietly being read as the same kind of number. They are not. One is a cost line, the other is a mixed construct that includes cost, defended revenue and speculative new revenue in a single total. Adding vendor and consultancy figures that measure different things is the same error I flag on RAN energy savings, where individually plausible percentages get stacked into a number no operator has ever achieved. Treat the sixty billion as the honest number, because at least you can tell what it counts.

Now to the architecture, which is where the article is most interesting and most incomplete. Google Cloud's prescription is a fabric of hyper specialised micro agents, billing agents, inventory agents, RAN guardians, communicating through standardised orchestration protocols, validated against a digital twin, bounded by what it calls a deterministic governance framework with explicit decision boundaries and clean handoff to human engineers. This is good engineering. It is also, for at least the sixth time in a month, a description of containment rather than coordination. Every agent in that picture belongs to the operator. Every protocol is internal. Every boundary is a boundary between the operator's machine and the operator's human. Nothing in the design describes what happens when an agent that the operator does not own, and cannot inspect, arrives with a request.

The digital twin makes the gap unusually visible. A twin is a high fidelity replica of your own network, and it is exactly the right tool for testing a configuration change before you ship it. It is useless for the case that actually matters commercially, because you cannot build a twin of the counterparty. When an enterprise's AI agent negotiates for a guaranteed slice, the operator's agent is not reasoning about a system it can simulate. It is reasoning about an intent it must infer, an authority it must verify, and a commitment it must be able to audit afterwards. That is not a simulation problem. It is a problem of shared topology, shared ontology, explicit authority boundaries and durable audit trails, which is the meta model of the agentic plane I have been arguing for since the spring, and which a runtime does not supply no matter how good the runtime is.

I should say plainly that a hyperscaler making this argument is not a criticism of the hyperscaler. Google Cloud is selling a stack that does what it says it does, and the operators quoted are getting real results from it. My argument is with how operators will read it. There is a version of the next two years in which the industry retools its operations beautifully, takes out a very large amount of cost, calls the result an AI business, and arrives in 2030 with the same revenue line and a smaller headcount. That would not be a failure of technology. It would be a failure to name what was bought.

So the test I would apply to this article, and to every agentic AI business case that lands on a telco investment committee this quarter, is the one I set out eight days ago. Which flow is this, cost, defence, or new revenue? If the supporting evidence is entirely tickets, incidents and NOC headcount, the answer is cost, and the paper should say so in its first sentence rather than its appendix. What binding constraint does an external buyer pay to remove? If nobody outside the company pays anything, there is no buyer, and the word opportunity is unearned. And where on the capacity, platform, outcome ladder does this sit? An operator that automates its own operations has not stepped onto the ladder at all, because the ladder is about what you sell, and nobody is buying your NOC.

Sixty billion dollars of avoided cost is worth having. It is worth a serious programme, serious money and serious executive attention. It is worth all of that as what it is. The operators that will be interesting in 2030 are not the ones that saved the most. They are the ones that could still tell you, at the end of it, which of the three flows each dollar came from.

Monday, July 13, 2026

AI monetization for operators: separating revenue from cost avoidance

Every operator earnings call now features AI prominently. Listen closely, however, and most of what is described as "AI monetization" is nothing of the sort. It is cost avoidance cosplaying a revenue costume.

This distinction matters because the two require different investment logic, different organizational capabilities, and different patience horizons. Operators that blur them will misallocate capital. Operators that separate them have a chance at building genuine new B2B revenue lines — narrower than the hype suggests, but investable.

Three money flows, not one

AI touches operator economics through three distinct channels, and the discipline starts with refusing to aggregate them.

1. AI that reduces cost. Autonomous network operations, agentic customer care, energy optimization, predictive maintenance. This is real, it is happening, and it is the largest near-term financial impact of AI on operators. It is also not revenue. A dollar of opex avoided is valuable, but it does not create a new line of business, and it does not justify the "operators as AI companies" narrative. It justifies a leaner operator.

2. AI that defends existing revenue. Enterprises deploying AI workloads have new connectivity requirements: deterministic performance, low latency to inference endpoints, secure private connectivity to GPU capacity, data-gravity-aware networking. Operators that serve these requirements protect and modestly grow their core B2B connectivity business. This is differentiated connectivity for the AI era — important, defensible, but fundamentally an evolution of what operators already sell.

3. AI that creates new revenue. This is the category everyone wants to talk about and the one that deserves the most scrutiny. It exists, but it is narrower than most strategy decks admit.

The four credible new revenue lines

Having spent the last two years working on AI infrastructure with operators and vendors on both sides of the Atlantic, I see four B2B revenue opportunities that survive contact with commercial reality. They are not equal — they differ in demand maturity, margin profile and time horizon, and they should be funded accordingly.

Sovereign AI capacity — GPU-as-a-service and AI factories — is the most immediate and the most misunderstood. Demand is real and policy-driven, concentrated in regulated sectors; the margin profile is low-to-mid, because the business is capex-heavy and carries utilization risk; and the revenue is available now. The demand side is genuine: governments, healthcare systems, defense, financial services and public administrations in Europe increasingly cannot — or will not — run inference on US hyperscaler infrastructure under foreign jurisdiction. Operators hold assets that map remarkably well to this demand: national data center footprints, energy contracts, security clearances, sovereign trust, and enterprise sales relationships.

Telefónica's recent national rollout of edge-based GPU-as-a-service in Spain is instructive. The underlying edge platform was architected years earlier — I led the team that built and productized it — and for years the business case was marginal on enterprise use cases alone. What changed was not the technology. It was the arrival of sovereign AI demand, which finally gave the infrastructure a paying anchor tenant profile. The lesson generalizes: edge and distributed compute investments become fundable when sovereignty is the demand driver, not the garnish.

The caution: this is a capex-intensive, utilization-sensitive business competing against hyperscalers with structurally lower unit costs. Operators win where sovereignty, data residency and proximity are binding constraints — and lose everywhere else. The addressable market is the regulated slice of national demand, not "the AI market."

Edge inference is real but earlier than its promoters claim. Demand exists where latency or data gravity bind; margins are mid-range; and the horizon is two to five years before this becomes a broad product line. The use cases that pay today are those where physics or data gravity make centralized inference impossible: industrial vision, real-time media production, autonomous operations in ports and factories. I have seen these work commercially. But the buyer set is narrow, and each engagement still resembles a system integration project more than a product sale. This becomes a scalable product line when agentic AI workloads distribute themselves across infrastructure tiers — the architecture I have described elsewhere as the AI Grid. That shift is underway, not arrived.

Data and trust services are the sleeper. Deepfake detection on voice calls, branded and verified calling, identity assurance for AI agents, provenance services. These are small revenue lines today, but they are high-margin, they monetize immediately, they sit directly on operator trust assets that hyperscalers cannot replicate, and demand grows with every AI-enabled fraud headline. For a B2B operator, this category has the best margin-to-capex ratio of the four.

Network APIs are the line whose trajectory has changed most in the past two years. The strategic logic has always been sound — AI agents will need to programmatically request network resources, quality on demand, location, verification — and the commercial signals are finally following: revenues are growing, aggregation initiatives have consolidated distribution, and enterprise visibility is rising with every agentic deployment that needs verified identity or guaranteed quality. It remains the earliest-stage of the four, and the AI agent wave — rather than developer evangelism — is what gives it genuine demand pull. I would invest now to be positioned, while sizing near-term revenue expectations with discipline; the inflection is likely in the second half of the decade.

The monetization ladder

Across all four lines, there is a ladder that determines margin and defensibility:

Sell capacity → sell platform → sell outcomes. Capacity here includes every consumption-metered unit: GPU hours, tokens, gigabits.

Selling raw capacity — GPU hours, token-metered inference, connectivity — is rung one: necessary, low-margin, commoditizing from day one. Tokens deserve a specific caution here: metering in tokens rather than GPU-hours changes the billing unit, not the business. An operator selling tokens against someone else's models and someone else's stack is still selling capacity, at prices that will be set by the most efficient infrastructure provider in the market. Selling a platform — inference-as-a-service with orchestration, security, compliance tooling — is rung two, where margins improve and switching costs appear. Selling outcomes — a fraud-detection rate, a production workflow, a compliant AI deployment for a hospital group — is rung three, where the economics finally resemble a services business worth building.

Operators historically stall at rung one. The reasons are organizational, not technological: product management that thinks in network elements rather than buyer problems, sales forces compensated on connectivity, and business cases that demand payback before the platform layer has time to mature. The operators that climb the ladder will be those that treat AI monetization as a product management and go-to-market transformation, not an infrastructure deployment.

What the buyer actually pays for

A final discipline. In every commercially successful case I have worked on, the enterprise buyer was not paying for "AI." They were paying for a constraint to be removed: data that could not leave the country, latency that broke the use case, a fraud pattern that was costing millions, a compliance requirement that blocked deployment. Price the constraint, not the technology. The moment an operator's AI proposition cannot name the constraint it removes, it is a science project.

Three questions before approving any operator AI business case

  1. Which of the three money flows is this — cost, defense, or new revenue? If the answer mixes them, send it back.
  2. What binding constraint does the buyer pay to remove, and why is an operator structurally better placed to remove it than a hyperscaler or an integrator? Sovereignty, proximity and trust are acceptable answers. "We have a network" is not.
  3. Where does this sit on the capacity–platform–outcome ladder, and what is the credible path up? Rung-one economics with rung-three ambitions is where operator AI investments go to die.

The AI B2B opportunity for operators is real. It is also smaller, slower and more demanding of commercial discipline than the current narrative suggests. The winners will not be the operators with the most GPUs. They will be the ones that can tell the difference between a cost saving, a defended revenue and a new business — and fund each accordingly.

Thursday, July 9, 2026

Operators Lean In On AI Grid Location


Earlier this week I argued that the AI Grid debate needs to move on from where you place a GPU to whether geographically dispersed compute can behave as a single fabric. I stand by that. But a story that has been building across the press this week is a useful reminder that the location question, the one I have been answering the same way for two years, is now being settled in public by the people who actually own the radio networks. And they are settling it against the tower.

The reporting is consistent. Light Reading describes Nokia and Nvidia's AI-RAN proposition running into telco resistance. Verizon, Vodafone, Orange and, notably for me, Telus have all raised doubts about putting graphics processing units into the radio access network. AT&T's chief technology officer has cast public doubt on the case for AI compute at the far edge. The enthusiasm for GPU-in-the-RAN comes from two operators, T-Mobile US and SoftBank, and almost no one else. Much of the rest of the industry is looking at Intel's newer CPUs for its open RAN rollouts rather than filling cell sites with accelerators.

I want to be precise about what this does and does not prove. It does not prove that AI in the RAN is a bad idea. Applying machine learning to scheduling, link adaptation and energy management inside the baseband is real, it is shipping, and Ericsson's AI-in-RAN software subscription is a reasonable way to bring it into existing hardware. What the operators are rejecting is narrower and more specific. They are rejecting the proposition that the cell site should become a general-purpose AI inference venue, stuffed with GPUs, monetised by hosting third-party workloads at the edge of the network. That is the proposition I have said for two years does not survive contact with power, cooling, space, security and, above all, the absence of a monetisation model.

My position has been that AI Grid deployment begins at the central office and the mobile switching office, not the cell site, because every physical and commercial constraint favours the aggregation point over the tower. The reasoning was never controversial to anyone who has stood in both kinds of building. A central office has power feeds, environmental control, physical security and fibre already in place. A cell site has a cabinet, a limited power budget and a landlord. When Verizon, Vodafone, Orange and Telus decline to put GPUs at the far edge, they are not making a new argument. They are confirming an old one, and they are confirming it with capital allocation decisions rather than conference slides, which is the only confirmation that counts.

There is a workstream reason this caught my eye. Telus appearing on the skeptics' list is consistent with what I see in the market: operators that are serious about autonomous operations are also the ones being disciplined about where AI compute physically lands. Those two forms of discipline are related. An operator that thinks clearly about the economics of edge inference tends to think clearly about the economics of everything else in the network.

The AI-RAN enthusiasm gap also matters for how we read vendor claims. When a technology has two vocal operator champions and a longer list of vocal operator skeptics, that is the signature of a capability that has been field-validated in specific conditions but not commercially validated across the market. I have made this distinction before and it applies cleanly here. SoftBank's agentic AI-RAN demonstrations and T-Mobile's Nvidia-backed edge trials are real engineering. They are not yet evidence that the model generalises to operators with different cost structures, different energy prices and different enterprise demand. Treat a two-operator enthusiasm as a pilot signal, not a market verdict.

So where does this leave the fabric argument I made last week? Exactly where I left it, and stronger. The operators are removing the least defensible node from the AI Grid, the cell site as inference host, which clears the ground for the argument that actually matters. Once you accept that heavy inference will not live at the tower, the interesting question becomes how you knit central offices, regional data centres and a small number of genuinely latency-bound edge sites into one addressable pool. The industry spent this week deciding where the compute will not go. That is progress. The harder decision, who owns the fabric that arbitrates across the places it will go, is still open, though.

Tuesday, July 7, 2026

AI Grid: Fabric vs Location Considerations

The debate about where AI compute belongs in a telecom network has been framed as a location question from the start. Do you put the GPUs at the cell site, the central office, the regional data centre, or the hyperscale campus? I have argued consistently that the honest answer begins at the central office and the mobile switching office, because power, cooling, fibre, physical security and latency sufficiency all favour those sites over the tower. That position has not changed. But two announcements from Asia this week suggest the more consequential question is no longer where the compute sits. It is whether the compute behaves as one pool regardless of where it sits.

NTT Docomo disclosed a nationwide testbed it calls GPU over APN. It pools graphics processing units spread across eight locations in five Japanese cities and presents them to a workload as a single platform, connected over the all-photonics network that NTT Group has been building under its IOWN programme. Docomo describes it as the realisation of its AI-Centric ICT Platform concept, part of what the group now labels AIOWN, its AI-native infrastructure. Strip away the acronyms and the claim is precise and significant: distributed GPUs, addressed as if co-located, over deterministic optical transport.

KT made the point from the other direction. Its new chief executive committed 18 trillion won, roughly 11.7 billion dollars, over three years, including 3.26 billion for one gigawatt of AI data centre capacity and a plan to connect that centralised infrastructure with edge sites serving low-latency workloads such as autonomous vehicles and industrial robotics. One operator is making dispersed compute act centralised. The other is extending centralised compute out to the edge. Both are describing the same thing from opposite ends, which is a compute fabric rather than a compute site.

This matters because it decouples two decisions the industry keeps conflating. Where you place a GPU is a question about power, land and cost. Where you run a workload is a question about latency, data gravity and sovereignty. As long as placement and execution are the same decision, every AI deployment becomes a real estate argument. Once a photonic fabric can make placement invisible to the workload, the two decisions separate. Training and heavy batch inference go where power and space are cheap. Latency-bound inference lands close to the user. The fabric arbitrates between them. This does not contradict the case for the central office, it absorbs it: the central office still wins for latency-bound edge inference, but that win is now a node in a graph rather than an isolated site.

I have some history with this problem. In 2018 I wrote about building at Telefonica what was probably the industry's first fully programmable multi-access edge computing platform, and the hardest part was never the compute. It was making distributed compute addressable, governable and billable as a coherent resource rather than a scatter of isolated sites. The technology around it has moved on considerably, but the unsolved problem is the same one Docomo is now attacking with photonics.

A practitioner's caution is in order. A fabric that makes national-scale GPUs behave as one pool is a testbed today, not a product. Docomo demonstrated it in a lab-grade programme. KT's edge connection is a plan, not a live deployment. Deterministic optical transport carrying commercial service level agreements under contended traffic is a materially harder thing than a controlled demonstration, and I would treat "as if co-located" the way I treat vendor energy savings figures: directionally real, quantitatively unproven at scale. The distance between a testbed that works and a fabric that carries production workloads is precisely the distance Open RAN spent five years crossing.

Still, the framing is the takeaway. The AI Grid conversation needs to move from siting to fabric. The operators that win will be the ones who can treat geographically dispersed compute as a single addressable resource, orchestrate workloads across it against real constraints, and price and settle access to it. That is a transport, orchestration and settlement problem before it is a property problem. The question is no longer which building holds the GPUs. It is who owns the fabric that makes the buildings irrelevant.

Monday, July 6, 2026

The Agent Runtime Is Not the Agent Model

DTW Ignite in Copenhagen made one thing clear: the vendor community has decided that the path to autonomous networks runs through agent runtimes. NVIDIA introduced NemoClaw blueprints and the OpenShell secure runtime to give long-running agents policy guardrails and sandboxed access to telecom systems. AdaptKey is piloting security-hardened agents for self-healing 5G operations. ServiceNow is bringing Project Arc to the NOC, orchestrating incident response from alert to work order. NTT DATA is building anomaly agents that escalate to research agents for telemetry analysis. Synthetic data rounds out the stack, a pragmatic answer to the fact that more than half of operators say their most valuable network data is too sensitive to use.

This is genuine progress and I do not want to minimize it. Containment, auditability and policy enforcement are necessary conditions for letting agents touch production networks. An agent that cannot be sandboxed cannot be trusted, and an agent whose actions cannot be audited cannot be certified. The runtime layer has to be built.

Containment is not coordination

But look carefully at what these announcements govern: individual agents, operating within a single operator's domain, executing workflows that a human has scoped in advance. This is vertical governance. It answers the question of whether an agent is allowed to perform an action. It does not answer the question that autonomous networks will actually pose at scale: when two agents are each permitted to act, and their permitted actions conflict, who decides?

Consider a scenario that is closer than most operators think. An enterprise logistics agent requests guaranteed throughput for a fleet of delivery robots. Simultaneously, a network energy agent, operating under its own perfectly valid mandate, is shutting down capacity in the same cluster to meet a sustainability target. Both agents are sandboxed. Both are auditable. Both are compliant with their policies. The runtime layer sees two well-behaved agents. The network sees a contradiction.

This is the problem I described in my previous post on network APIs. APIs were designed for developer access, not for agent-to-agent negotiation. Runtimes inherit the same blind spot. They secure the execution of each agent without providing any shared representation of the agentic plane itself.

What the meta-model requires

For agents to negotiate rather than collide, the industry needs a meta-model of the agentic plane: a topology of which agents exist and where they sit, an ontology so that an enterprise agent and a network agent mean the same thing by capacity, latency or priority, explicit authority boundaries defining what each agent may commit on behalf of its principal, shared state models so that negotiations reference the same view of the network, and audit trails that span negotiations rather than individual actions. None of the DTW announcements address this layer. They cannot, because it is not a product any single vendor can ship. It is a model the industry must agree on, the way it once agreed on network information models for OSS.

There is a familiar pattern here. The industry built firewalls before it built routing protocols for the internet's trust boundaries, and it spent two decades paying for the sequencing. We are building the firewalls of the agentic era first. The operators and standards bodies that formalize the agentic plane meta-model will define how enterprise AI and network AI transact for the next decade. The ones that stop at the runtime will discover that a network full of safely contained agents is not an autonomous network.

Wednesday, July 1, 2026

DTW Ignite 2026: The API Is Not Enough

I returned from DTW Ignite in Copenhagen with one conviction: the interface between enterprise applications and network infrastructure is about to change in a way the industry has not yet designed for.

Network APIs were never really about autonomous networks. That framing conflates two separate problems. APIs — CAMARA, GSMA Open Gateway, the decades of network exposure work that preceded them — were designed to let developers discover and consume network resources from outside the operator domain. Quality on Demand, location services, device status, number verification: clean REST interfaces exposed through a developer portal so that a programmer writing a B2B application could request a network capability and pay for it. Real progress on a real problem. But the problem was developer access, not network autonomy.

What is coming next is different in kind, not degree.

Enterprise AI agents are beginning to consume network infrastructure directly — not through a developer writing an integration, but autonomously, in real time, as part of executing a business objective. An industrial automation agent that needs guaranteed low-latency connectivity for a robotics fleet. A financial services agent that needs to provision a secure, isolated network path for a time-sensitive transaction. A logistics agent that needs to dynamically reserve bandwidth across multiple carrier domains as a shipment moves between jurisdictions. In none of these cases is there a developer in the loop. The agent has an intent, it needs network resources to fulfil it, and it needs to negotiate those resources with the network — now, at machine speed, without human mediation.

That negotiation cannot happen through a developer portal. It cannot happen through a static API catalogue with a PDF explaining what each endpoint does. The enterprise agent and the network need to speak to each other, and neither CAMARA nor MCP — whatever their respective merits — were designed for that conversation.

The network side of this exchange needs to be represented by network AI agents of its own: agents that can expose available capacity in real time, understand the constraints and commitments already in place, reason over competing demands, and negotiate resource allocation in a way that respects the network's operating boundaries. That is not a developer API. That is an autonomous counterparty.

And for those network agents to function — to negotiate reliably, to be governed, to be audited, to avoid conflicting with each other across RAN, transport, core, and the operational layers of OSS and BSS — they need something the industry is not yet building: a meta-model of the agentic plane itself.

Operators building autonomous networks are doing the right foundational work. Network topology models. Data ontologies. Decision layers. Closed-loop control architectures. These give the automation layer a complete and current picture of the environment it is operating in. But agents operating on that network need an equivalent model of themselves. Every agent with an identity, a capability scope, an authority boundary, a state, a dependency graph, and an audit trail. An abstract topology and ontology of agents, sitting alongside the topology and ontology of the network.

Without that model, what looks like autonomous negotiation between enterprise AI and network AI is actually uncontrolled interaction between systems that cannot see each other. An enterprise agent requesting bandwidth does not know what the network agent is authorised to commit. The network agent does not know what other network agents have already promised. No shared representation, no conflict detection, no governance.

The developer exposure problem is largely solved, or at least well understood. The agent-to-agent negotiation problem has barely been framed. That is the conversation the industry needs to have, and Copenhagen convinced me we are not having it yet.