Showing posts with label AI grid. Show all posts
Showing posts with label AI grid. Show all posts

Sunday, August 23, 2026

Ericsson Says the Telco Edge Was Too Early

Ericsson has put a name to the failure of the last edge computing cycle. Joe Constantine, the company's Americas chief strategy and technology officer, told Fierce Network in a piece published by Diana Goovaerts on 21 August that mobile edge computing was early rather than wrong, and his summary of what went wrong is short. "Ten years ago, MEC was a supply side concept." What has changed, in his account, is that a demand driver has arrived. "We have AI inferencing. It's the growth opportunity, an application that MEC did not have at the time. So, if you look at this, we believe that the industry, the technology and the market is vastly different today from 10 years ago." He supports the traffic case with Ericsson's own forecasts, that global mobile traffic will triple between 2023 and 2029 with AI as the key driver, and that uplink traffic will grow 10 times by 2035. He argues the network has changed too, from best effort connectivity to 5G built for time-critical communication and capable of 15 millisecond latency and "five nines of reliability." Fierce notes in the same piece that TM Forum's chief executive has told it that operators should not bet the farm on the edge.

The diagnosis is correct and it is unusual to hear it from a vendor

Supply side concept is the right post-mortem, and it is the shortest accurate description of that decade anyone has offered on the record. Operators built edge capacity because they had property, power and a latency story, then went looking for someone who wanted it. The platforms worked. The buyers did not appear. When I wrote about a fully programmable multi-access edge platform in November 2018, the services I could name were faster file uploads, console games without a console, and editing a document without downloading it. Those were real improvements to existing experiences. None of them was a business that an enterprise procurement department was going to sign for. Inference is a materially better answer than that, because it is a workload with a measurable cost that somebody is already paying somewhere else. That is a genuine change in the argument and it should be credited before anything else is said about it.

15 milliseconds is a central office number

The interesting discipline in Constantine's case is that his own numbers settle the location question, and they settle it away from the radio. A 15 millisecond budget is a loose one. It is comfortably met from a metro exchange or a mobile switching office serving hundreds of sites, and it does not require compute at the tower. That matters because the edge conversation still routinely conflates two very different capital programmes. Putting accelerators in a few hundred central offices is a brownfield project on estate that is already zoned, powered, cooled, fibred and physically secured. Putting them at tens of thousands of cell sites is a different business with a different power bill, a different maintenance model and a different landlord. I have argued on the AI Grid that deployment starts at central offices and mobile switching offices for exactly these reasons, and the 2018 platform was built into the central office for the same ones. Nothing in the Ericsson case contradicts that. If 15 milliseconds is the requirement, the requirement is an exchange.

The economic claim is asserted rather than costed

"Routing all this traffic to a centralized cloud isn't just only slow, it's economically not even sustainable" is the load-bearing sentence of the whole argument, and it arrives without a number attached. There is no published cost per inference at a telco edge site set against the same inference in a hyperscale region, at comparable utilisation, including the operator's power, cooling, refresh and remote-hands costs. Until that comparison exists, the economic case for distribution rests on the intuition that moving bits is expensive and moving them less must therefore be cheaper. That intuition ignores utilisation, which is what actually decides the economics of accelerators. A central region runs its fleet hot across many customers and time zones. A metro edge site runs a smaller fleet against local demand that peaks and troughs. The traffic forecasts are also Ericsson's own, published in its Mobility Report, and a 10 times uplink projection reaching to 2035 is a forecast rather than an observation. It should be read as one.

Physical AI names a use case, not a counterparty

Asked what actually requires the edge, Constantine points to robotics and physical AI. "Any system that moves, so drones, vehicles, humanoids, robots talking to other robots, all these will become an autonomous system that needs a network to think with it." His example is an autonomous car whose onboard sensors can see what is in front of it but not around the corner, with network sensing supplying the rest. The technical case is sound. The commercial case has the same shape as the one that failed. A drone fleet, a warehouse robot estate and a vehicle platform are owned by companies that are not the operator, and the thing being sold to them is not raw compute, which they can buy cheaper elsewhere, but contextual awareness that the network holds and they do not.

Selling that is a coordination problem before it is a latency problem. Robots talking to other robots across an ownership boundary need an agreed model of what is being asserted, whose authority stands behind it, and what record both sides will accept afterwards. That is the gap I argued no runtime supplies, and that NGMN has now enumerated at length for coordination inside a single operator's own domains, which is the easier version of the same problem. An operator can install the accelerators in its exchanges this year. It cannot unilaterally produce the model that lets a vehicle manufacturer's autonomy stack trust and pay for what the network says about the road. The compute is the part of this that money can buy quickly, which is why it is the part vendors describe in detail.

What to watch

3 things. First, where the accelerators physically land. If operator edge AI deployments in the next 18 months are announced at central offices and metro sites, the location argument is settled and the cell site edge can be retired from the conversation. Second, a published cost per inference or cost per token at a telco edge site against a hyperscale baseline, from anybody, on the record. That single figure decides whether distributed inference is an economic position or a latency preference. Third, the first contract in which an enterprise pays an operator for network context rather than for compute or connectivity, because that is the class of deal the 2010s never produced and it is the only one that would show the demand side has actually changed.

Wednesday, August 19, 2026

NGMN Confirms Agentic AI Needs Improvements

The Next Generation Mobile Networks Alliance published a report on 12 August, "Network Automation and Autonomy Phase III: Agentic AI for Autonomous Mobile Networks," and it reads like a standards body catching up to an argument I have been making for a year. The headline finding, reported by Keith Dyer at The Mobile Network, is that the industry has to solve interoperability, security, governance and data before agentic AI can deliver autonomous networks at commercial scale. The detail underneath the headline is what matters. NGMN says the current ecosystem is too fragmented for agents to form a consistent understanding of the network, because vendors use different data models and different terminology, and it calls for common information models, ontologies and semantic mappings, aligned across 3GPP, TM Forum, ETSI, O-RAN, IETF, W3C and BBF. It asks for a telecom-grade Zero-Trust Agent Ecosystem with secure agent identity, authentication, authorisation, policy enforcement, audit logging, runtime monitoring, and the ability to revoke or quarantine an agent. Read that list again. Topology, ontology, authority boundaries, state, audit trails. That is the meta-model of the agentic plane, and it is now in a report with an operator alliance's name on the cover.

I want to be precise about the provenance here, because the sequence is the point. I argued at DTW Ignite that the API is not enough, that the interfaces we built for developer access were never designed to let an autonomous agent understand what it is acting on. I argued a few weeks later that the agent runtime is not the agent model, that a place to execute an agent and put guardrails around it supplies none of the topology, ontology, authority and audit that coordination between agents actually needs. NGMN has now enumerated eight cross-organisation challenge areas, and they map almost one for one onto that case. Lack of standards for agent knowledge and context sharing. Incomplete information modelling for networks. Lack of end-to-end security, identity and trust for autonomous functions. Lack of operator-friendly governance and lifecycle tools. Under-addressed economic and organisational readiness. When a body chaired by Orange's group CTO writes the same list I published, the argument stops being a contrarian read and becomes the consensus. That is a good day for the thesis. It is a better day to point out the part the report does not solve.

The report describes coordination inside a boundary the operator owns

Look closely at the architecture NGMN proposes. A supervisory agent coordinates specialist agents, one each for RAN, core, transport, cloud and security, to diagnose a service problem, weigh corrective actions, resolve conflicts, and verify that service was restored. The worked examples are fault management, service assurance and RAN optimisation. In the RAN case, a cross-domain agent delegates intent to a RAN optimisation agent, which supervises a sub-agent in the infrastructure layer, and when an optimisation action collides with network-wide energy management the cross-domain agent arbitrates the trade-off. This is a serious and correct piece of engineering. It is also, in every example, coordination across domains that a single operator owns. RAN, core, transport, cloud and security are five administrative domains inside one company's walls. The hard thing NGMN is describing is getting an operator's own silos, built by different vendors with different data models, to hand each other a shared and trustworthy understanding of state. That is worth doing, and it is exactly why a runtime and a set of guardrails were never going to be enough. But it is one owner reconciling with itself.

The problem I have been pointing at sits one boundary further out, and NGMN's own scope is the clearest evidence that it is a distinct and harder problem. When an enterprise's AI wants to negotiate with an operator's network AI, for a service level, a slice, a capacity commitment, a remediation, there is no shared owner to impose a common ontology, no single authority to issue the identities, and no single audit log that both parties will accept as the record of what was agreed and who was answerable. Everything NGMN specifies, the semantic alignment, the zero-trust identities, the audit trail, is defined for agents operating within the operator's estate. Across an ownership boundary, the same requirements do not disappear. They get harder, because now two parties have to agree on the model before either can trust an action, and neither controls the other. NGMN has validated that the meta-model is necessary. It has not, and by its own framing could not, close the gap where the two hardest words in this whole subject live, which are coordination and accountability between parties who do not report to the same CTO.

Containment, coordination, accountability, in that order of difficulty

It is worth keeping the three layers separate, because vendors and now standards bodies keep collapsing them. Containment is the Zero-Trust Agent Ecosystem: the identity, the authorisation, the runtime monitoring, the kill switch that lets you revoke or quarantine an agent inside your own network. That is genuinely useful and NGMN is right to specify it in detail. Coordination is the shared ontology and semantic mapping that lets one agent understand what another agent is doing, and this is the layer NGMN is now trying to build across an operator's internal domains. Accountability is the ability to produce, on demand, the identity of an agent, the authority under which it acted, and a trail that a regulator and an outside counterparty would both sign, across a boundary you do not own. The industry has spent a year shipping the first layer and calling it the third. NGMN has now, to its credit, put real weight on the second. The third is still open, and it is the one I argued the EU AI Act made expensive when its enforcement powers switched on at the start of this month, because enforcement does not stop at the edge of your administrative domain.

There is a discipline point here that operators should not lose in the enthusiasm of seeing a favourite thesis ratified. NGMN itself flags it: "While technology groups focus on architecture, few have formal guidance on the required telco organisational changes," and it lists economic and organisational readiness as an under-addressed challenge in its own right. The meta-model is not a product you procure. It is a set of agreements about topology, ontology, authority and audit that has to be built and then adopted, and the report's call for alignment across seven standards organisations is a fair description of how long that takes. Cross-SDO harmonisation of information models is measured in years, not quarters. When I helped set network autonomy targets in operator programmes, the technology was rarely the thing that slipped. The organisational change and the cross-domain agreements were. Anyone booking autonomous-network savings against a framework paper rather than a ratified model is doing the same thing I warn against when vendors stack their individual efficiency claims into a total no network has ever achieved. Read this report as the industry agreeing on the destination. Do not read it as arrival.

What to watch

The honest test has not changed, and NGMN has now made it a formal requirement rather than a personal opinion. For any action an agent takes on its own, can you produce its identity, the authority under which it acted, and an audit trail that both a regulator and an outside counterparty would accept. If the answer lives entirely inside one vendor's runtime, you have containment and you should not mistake it for accountability. NGMN has described the plane the runtime runs on, in more detail than anyone in the industry has committed to print, and it has correctly located the work in standards and organisation rather than in any single box. The next phase belongs to whoever builds the ontology and the identity fabric that hold up not just across an operator's own domains, but across the boundary to the enterprise on the other side of the negotiation. That boundary is where the value is, and it is the one the report stops at. Watch for whether the alignment NGMN asks for actually happens across those seven bodies, and watch, as always, for a pilot inside one operator's walls being described as autonomy across everyone's.

Tuesday, July 28, 2026

The Price Of RAN Security In Europe: 40B Euros... And More

Seven of Europe's largest operators commissioned GSMA Intelligence to price the removal of designated high-risk vendors from their networks, and the number came back this week. Under the European Commission's proposed Cybersecurity Act 2, stripping out equipment from suppliers such as Huawei would cost the region's operators between thirty and forty billion euros in direct costs, split across mobile at up to twenty two billion, fixed at around five billion, and transport at up to twelve billion. The figure that matters more, though, is the second one. GSMAi finds that the same measure would push mobile equipment prices up by roughly twenty four percent, because forcing designated vendors out of the market shrinks the field of suppliers, and a smaller field charges more. Deutsche Telekom, Fastweb plus Vodafone, Meo, Orange, Telefonica, United Group and Vodafone Group supplied the underlying data.

I have spent part of my career making the opposite bet. The entire economic premise of Open RAN, the work I led at Telefonica and have written about for years, is that opening the interfaces widens the supplier base, and a wider base drives the cost of the radio down. You can read my longer assessment of where that project actually stands on the state of Open RAN, but the direction of the argument was never in doubt: more suppliers, more competition, lower unit cost. What the Cybersecurity Act 2 proposal does is run that logic in reverse. It removes suppliers, it narrows competition, and the price goes up. The twenty four percent is not a side effect anyone chose. It is arithmetic. You cannot both mandate a smaller supplier pool and expect the survivors to hold their prices.

What makes this more than a European regulatory footnote is where it occurs. It lands on top of a repricing that is already under way for an entirely different reason. I have been arguing for some weeks that the AI build-out is bidding up the exact components a modern radio depends on, high bandwidth memory, advanced substrates, power semiconductors, because AI factories and the RAN supply chain now compete for the same silicon. The consequence I keep pointing to is that even an operator which never buys a single GPU still pays more for its radios, because the AI boom has repriced the inputs. Now set the GSMAi finding beside that. You have two independent forces, one driven by AI demand and one driven by security regulation, both pushing the unit cost of the radio in the same direction, at the precise moment operators are being told to find capital for AI.

Beware of stacked figures though, so let me be precise. Do not add the two numbers. The forty billion is a one-time direct cost of replacement. The twenty four percent is a forward price trajectory on new equipment. They measure different things over different horizons, and stacking them produces a headline number no operator will ever see on an invoice, which is the same error I flag when vendors add their individual RAN energy savings into a total no network has ever achieved. The direction is clear, though.The radio is getting more expensive from two sides at once, and only one of those sides is optional.

This sharpens a position I have held on the location question. If the radio is being repriced upward by both AI demand and regulation, the case for distributing expensive compute out to every cell site gets weaker, not stronger. Costly, repriced silicon is exactly the kind of asset you concentrate where you can amortise it, at the central office and the switching centre, not the kind you scatter across tens of thousands of sites on the promise of a latency budget that most workloads do not need. I made that argument on its own terms when I wrote that operators are leaning in on AI Grid location and set out the fabric versus location considerations underneath it. The cost side now reinforces the physics side. Scarcity concentrates capital.

There is a monetisation discipline here as well, and it is the one I set out when I argued for separating revenue from cost avoidance. High-risk vendor removal is not a modernisation and it is not a growth programme. It is a defensive cost, imposed from outside, that produces no new revenue line and buys no new capability. An operator can, of course, use a forced swap-out as the occasion to modernise, and some will. But the spend itself belongs on the cost side of the ledger, named honestly, and not folded into an AI or transformation story where it does not belong. The moment a swap-out driven by security regulation starts appearing in a slide about AI readiness, someone has mixed the flows again.

At last, this is RAN. While going forward these elements will become more programmatic, it is still unclear whether a 2G, 3G or 4G radio from a high risk vendor actually represents a security risk. The highest compromise risk is in the Core, the OSS and ancillary systems. It is easier to install IMEI catchers and antenna spoofs than to hack into a deployed live system.

Europe may well decide the security case is worth forty billion euros. That is a legitimate political choice. What the choice is not is free, and it is not neutral for competition. A framework designed to reduce dependence on one class of vendor achieves it by reducing dependence on vendors generally, which is to say by making the market smaller and the equipment dearer. Operators should book that outcome for exactly what it is: a defensive, imposed, largely non-recoverable cost, landing on an equipment line that the AI build-out was already repricing without any help from regulators. It is certainly not random that the leading RAN vendors are predominantly europeans, with emerging Korean and Japanese options.

The political decision would transfer value from networks to vendors. It is unclear whether that is sustainable if the pool of vendors does not increase substantially.

Monday, July 27, 2026

Verizon Earning: From Copper to Fibre to Edge

 

Verizon disclosed on its second quarter earnings call last week a dark fibre agreement with Google worth well in excess of a billion dollars, and chief executive Dan Schulman was explicit that it is the first of several, with further deals expected by year end worth multiple billions of dollars in revenue over the coming years. The headlines went to the number and to the counterparty. The more instructive detail came later in the same call, where Schulman described retrofitting thousands of central offices, the copper now being decommissioned, into edge data centres for low latency AI inferencing. One operator, one earnings call, two of the arguments I have been making all month, and they are not the same argument.

Let's start with the fibre. This is not cost avoidance dressed as growth, which is the trap I wrote about when Google Cloud called agentic AI a sixty billion dollar opportunity and every figure underneath the headline turned out to be a saved operating cost. It is also not the speculative new revenue that strategy decks reach for. It is my second money flow, AI creating fresh demand for what the operator already sells, landed on the income statement as contracted revenue. Schulman called it incremental, long duration, high quality, and drawn from some of the most demanding infrastructure customers in the world. He is right to be pleased. Route and real estate are exactly the assets an operator holds that a hyperscaler cannot conjure at will, and the AI build out is short of both.

Now look at what Verizon actually sold. Dark fibre is unlit glass. Google puts its own optics on each end, chooses its own wavelengths, runs its own capacity, and owns everything above the physical layer. Verizon is the landlord of the route and nothing more. On the capacity, platform, outcome ladder I set out two weeks ago, this is not even rung one, it is the ground the ladder stands on. That is not a criticism. A contracted, long duration, low churn landlord business against demand this strong is a genuinely good thing to own, and it is more defensible than most of what operators like to call platforms. The discipline is only this: name it correctly. The moment next year's deck describes a dark fibre lease as an AI platform business, the margin expectation that travels with the word platform will arrive, and a landlord business will not carry it.

The central office retrofit is the disclosure I want to emphasize. For two years I have argued that AI grid compute belongs first at the central office and the mobile switching office, not at the cell site, because power, cooling, fibre, real estate and security all favour the building the operator already runs. I made the fabric versus location case in early July and watched operators lean into central office siting a few days later. Verizon has now put capital behind it on an earnings call. The copper decommission is what makes it work: retiring the old plant frees the floor space and, more importantly, the power feed and the fibre entrance, which are the two constraints that actually bind at an inference site. I built what was probably the first fully programmable multi access edge platform at Telefonica in 2018, and the lesson from that programme was that the physics was never the obstacle. The obstacle was a paying tenant. Low latency inference is the tenant the central office was always waiting for.

The reason to read the two disclosures together is that they resolve a question people keep posing as a choice. Fabric or location, route or venue, is the AI grid a transport problem or a siting problem. Verizon's answer, in one call, is both, and the call even tells you which is which today. The fabric is the contracted revenue, available now, sold in its rawest form. The location is the capital project, the copper coming out and the racks going in, its revenue still ahead of it. An operator that understood only the first would sell glass to hyperscalers and miss the building. An operator that understood only the second would light up central offices with no anchor tenant, which is precisely the mistake the edge computing industry made for a decade. Verizon is doing both, and the sequencing is correct.

So I would put the deal through the same three questions I put to every operator AI business case. Which money flow is this. It is flow two, defended and grown connectivity revenue, honestly labelled, not flow three in disguise. What binding constraint does the buyer pay to remove. Google is paying for route diversity and a dedicated physical layer it controls end to end, away from shared congestion, which is a real constraint and an operator asset. And where does it sit on the ladder. Rung one for the fibre, with a credible path upward only if the central office retrofit becomes a platform the operator actually operates, rather than a colocation cage it merely rents to the same hyperscalers. The fibre deal is booked. The ladder question is still open, and it will be answered in the buildings, not on the routes.

I wrote a fortnight ago that the operators who will be interesting in 2030 are not the ones with the most GPUs but the ones who can still tell you which of the three flows each dollar came from. Verizon has just given the cleanest demonstration yet of the discipline: a fabric dollar and a location dollar, named separately, on the same call. The test now is whether it keeps them separate all the way up the ladder, or whether the word platform arrives before the platform does.

Monday, July 13, 2026

AI monetization for operators: separating revenue from cost avoidance

Every operator earnings call now features AI prominently. Listen closely, however, and most of what is described as "AI monetization" is nothing of the sort. It is cost avoidance cosplaying a revenue costume.

This distinction matters because the two require different investment logic, different organizational capabilities, and different patience horizons. Operators that blur them will misallocate capital. Operators that separate them have a chance at building genuine new B2B revenue lines — narrower than the hype suggests, but investable.

Three money flows, not one

AI touches operator economics through three distinct channels, and the discipline starts with refusing to aggregate them.

1. AI that reduces cost. Autonomous network operations, agentic customer care, energy optimization, predictive maintenance. This is real, it is happening, and it is the largest near-term financial impact of AI on operators. It is also not revenue. A dollar of opex avoided is valuable, but it does not create a new line of business, and it does not justify the "operators as AI companies" narrative. It justifies a leaner operator.

2. AI that defends existing revenue. Enterprises deploying AI workloads have new connectivity requirements: deterministic performance, low latency to inference endpoints, secure private connectivity to GPU capacity, data-gravity-aware networking. Operators that serve these requirements protect and modestly grow their core B2B connectivity business. This is differentiated connectivity for the AI era — important, defensible, but fundamentally an evolution of what operators already sell.

3. AI that creates new revenue. This is the category everyone wants to talk about and the one that deserves the most scrutiny. It exists, but it is narrower than most strategy decks admit.

The four credible new revenue lines

Having spent the last two years working on AI infrastructure with operators and vendors on both sides of the Atlantic, I see four B2B revenue opportunities that survive contact with commercial reality. They are not equal — they differ in demand maturity, margin profile and time horizon, and they should be funded accordingly.

Sovereign AI capacity — GPU-as-a-service and AI factories — is the most immediate and the most misunderstood. Demand is real and policy-driven, concentrated in regulated sectors; the margin profile is low-to-mid, because the business is capex-heavy and carries utilization risk; and the revenue is available now. The demand side is genuine: governments, healthcare systems, defense, financial services and public administrations in Europe increasingly cannot — or will not — run inference on US hyperscaler infrastructure under foreign jurisdiction. Operators hold assets that map remarkably well to this demand: national data center footprints, energy contracts, security clearances, sovereign trust, and enterprise sales relationships.

Telefónica's recent national rollout of edge-based GPU-as-a-service in Spain is instructive. The underlying edge platform was architected years earlier — I led the team that built and productized it — and for years the business case was marginal on enterprise use cases alone. What changed was not the technology. It was the arrival of sovereign AI demand, which finally gave the infrastructure a paying anchor tenant profile. The lesson generalizes: edge and distributed compute investments become fundable when sovereignty is the demand driver, not the garnish.

The caution: this is a capex-intensive, utilization-sensitive business competing against hyperscalers with structurally lower unit costs. Operators win where sovereignty, data residency and proximity are binding constraints — and lose everywhere else. The addressable market is the regulated slice of national demand, not "the AI market."

Edge inference is real but earlier than its promoters claim. Demand exists where latency or data gravity bind; margins are mid-range; and the horizon is two to five years before this becomes a broad product line. The use cases that pay today are those where physics or data gravity make centralized inference impossible: industrial vision, real-time media production, autonomous operations in ports and factories. I have seen these work commercially. But the buyer set is narrow, and each engagement still resembles a system integration project more than a product sale. This becomes a scalable product line when agentic AI workloads distribute themselves across infrastructure tiers — the architecture I have described elsewhere as the AI Grid. That shift is underway, not arrived.

Data and trust services are the sleeper. Deepfake detection on voice calls, branded and verified calling, identity assurance for AI agents, provenance services. These are small revenue lines today, but they are high-margin, they monetize immediately, they sit directly on operator trust assets that hyperscalers cannot replicate, and demand grows with every AI-enabled fraud headline. For a B2B operator, this category has the best margin-to-capex ratio of the four.

Network APIs are the line whose trajectory has changed most in the past two years. The strategic logic has always been sound — AI agents will need to programmatically request network resources, quality on demand, location, verification — and the commercial signals are finally following: revenues are growing, aggregation initiatives have consolidated distribution, and enterprise visibility is rising with every agentic deployment that needs verified identity or guaranteed quality. It remains the earliest-stage of the four, and the AI agent wave — rather than developer evangelism — is what gives it genuine demand pull. I would invest now to be positioned, while sizing near-term revenue expectations with discipline; the inflection is likely in the second half of the decade.

The monetization ladder

Across all four lines, there is a ladder that determines margin and defensibility:

Sell capacity → sell platform → sell outcomes. Capacity here includes every consumption-metered unit: GPU hours, tokens, gigabits.

Selling raw capacity — GPU hours, token-metered inference, connectivity — is rung one: necessary, low-margin, commoditizing from day one. Tokens deserve a specific caution here: metering in tokens rather than GPU-hours changes the billing unit, not the business. An operator selling tokens against someone else's models and someone else's stack is still selling capacity, at prices that will be set by the most efficient infrastructure provider in the market. Selling a platform — inference-as-a-service with orchestration, security, compliance tooling — is rung two, where margins improve and switching costs appear. Selling outcomes — a fraud-detection rate, a production workflow, a compliant AI deployment for a hospital group — is rung three, where the economics finally resemble a services business worth building.

Operators historically stall at rung one. The reasons are organizational, not technological: product management that thinks in network elements rather than buyer problems, sales forces compensated on connectivity, and business cases that demand payback before the platform layer has time to mature. The operators that climb the ladder will be those that treat AI monetization as a product management and go-to-market transformation, not an infrastructure deployment.

What the buyer actually pays for

A final discipline. In every commercially successful case I have worked on, the enterprise buyer was not paying for "AI." They were paying for a constraint to be removed: data that could not leave the country, latency that broke the use case, a fraud pattern that was costing millions, a compliance requirement that blocked deployment. Price the constraint, not the technology. The moment an operator's AI proposition cannot name the constraint it removes, it is a science project.

Three questions before approving any operator AI business case

  1. Which of the three money flows is this — cost, defense, or new revenue? If the answer mixes them, send it back.
  2. What binding constraint does the buyer pay to remove, and why is an operator structurally better placed to remove it than a hyperscaler or an integrator? Sovereignty, proximity and trust are acceptable answers. "We have a network" is not.
  3. Where does this sit on the capacity–platform–outcome ladder, and what is the credible path up? Rung-one economics with rung-three ambitions is where operator AI investments go to die.

The AI B2B opportunity for operators is real. It is also smaller, slower and more demanding of commercial discipline than the current narrative suggests. The winners will not be the operators with the most GPUs. They will be the ones that can tell the difference between a cost saving, a defended revenue and a new business — and fund each accordingly.

Thursday, July 9, 2026

Operators Lean In On AI Grid Location


Earlier this week I argued that the AI Grid debate needs to move on from where you place a GPU to whether geographically dispersed compute can behave as a single fabric. I stand by that. But a story that has been building across the press this week is a useful reminder that the location question, the one I have been answering the same way for two years, is now being settled in public by the people who actually own the radio networks. And they are settling it against the tower.

The reporting is consistent. Light Reading describes Nokia and Nvidia's AI-RAN proposition running into telco resistance. Verizon, Vodafone, Orange and, notably for me, Telus have all raised doubts about putting graphics processing units into the radio access network. AT&T's chief technology officer has cast public doubt on the case for AI compute at the far edge. The enthusiasm for GPU-in-the-RAN comes from two operators, T-Mobile US and SoftBank, and almost no one else. Much of the rest of the industry is looking at Intel's newer CPUs for its open RAN rollouts rather than filling cell sites with accelerators.

I want to be precise about what this does and does not prove. It does not prove that AI in the RAN is a bad idea. Applying machine learning to scheduling, link adaptation and energy management inside the baseband is real, it is shipping, and Ericsson's AI-in-RAN software subscription is a reasonable way to bring it into existing hardware. What the operators are rejecting is narrower and more specific. They are rejecting the proposition that the cell site should become a general-purpose AI inference venue, stuffed with GPUs, monetised by hosting third-party workloads at the edge of the network. That is the proposition I have said for two years does not survive contact with power, cooling, space, security and, above all, the absence of a monetisation model.

My position has been that AI Grid deployment begins at the central office and the mobile switching office, not the cell site, because every physical and commercial constraint favours the aggregation point over the tower. The reasoning was never controversial to anyone who has stood in both kinds of building. A central office has power feeds, environmental control, physical security and fibre already in place. A cell site has a cabinet, a limited power budget and a landlord. When Verizon, Vodafone, Orange and Telus decline to put GPUs at the far edge, they are not making a new argument. They are confirming an old one, and they are confirming it with capital allocation decisions rather than conference slides, which is the only confirmation that counts.

There is a workstream reason this caught my eye. Telus appearing on the skeptics' list is consistent with what I see in the market: operators that are serious about autonomous operations are also the ones being disciplined about where AI compute physically lands. Those two forms of discipline are related. An operator that thinks clearly about the economics of edge inference tends to think clearly about the economics of everything else in the network.

The AI-RAN enthusiasm gap also matters for how we read vendor claims. When a technology has two vocal operator champions and a longer list of vocal operator skeptics, that is the signature of a capability that has been field-validated in specific conditions but not commercially validated across the market. I have made this distinction before and it applies cleanly here. SoftBank's agentic AI-RAN demonstrations and T-Mobile's Nvidia-backed edge trials are real engineering. They are not yet evidence that the model generalises to operators with different cost structures, different energy prices and different enterprise demand. Treat a two-operator enthusiasm as a pilot signal, not a market verdict.

So where does this leave the fabric argument I made last week? Exactly where I left it, and stronger. The operators are removing the least defensible node from the AI Grid, the cell site as inference host, which clears the ground for the argument that actually matters. Once you accept that heavy inference will not live at the tower, the interesting question becomes how you knit central offices, regional data centres and a small number of genuinely latency-bound edge sites into one addressable pool. The industry spent this week deciding where the compute will not go. That is progress. The harder decision, who owns the fabric that arbitrates across the places it will go, is still open, though.

Tuesday, July 7, 2026

AI Grid: Fabric vs Location Considerations

The debate about where AI compute belongs in a telecom network has been framed as a location question from the start. Do you put the GPUs at the cell site, the central office, the regional data centre, or the hyperscale campus? I have argued consistently that the honest answer begins at the central office and the mobile switching office, because power, cooling, fibre, physical security and latency sufficiency all favour those sites over the tower. That position has not changed. But two announcements from Asia this week suggest the more consequential question is no longer where the compute sits. It is whether the compute behaves as one pool regardless of where it sits.

NTT Docomo disclosed a nationwide testbed it calls GPU over APN. It pools graphics processing units spread across eight locations in five Japanese cities and presents them to a workload as a single platform, connected over the all-photonics network that NTT Group has been building under its IOWN programme. Docomo describes it as the realisation of its AI-Centric ICT Platform concept, part of what the group now labels AIOWN, its AI-native infrastructure. Strip away the acronyms and the claim is precise and significant: distributed GPUs, addressed as if co-located, over deterministic optical transport.

KT made the point from the other direction. Its new chief executive committed 18 trillion won, roughly 11.7 billion dollars, over three years, including 3.26 billion for one gigawatt of AI data centre capacity and a plan to connect that centralised infrastructure with edge sites serving low-latency workloads such as autonomous vehicles and industrial robotics. One operator is making dispersed compute act centralised. The other is extending centralised compute out to the edge. Both are describing the same thing from opposite ends, which is a compute fabric rather than a compute site.

This matters because it decouples two decisions the industry keeps conflating. Where you place a GPU is a question about power, land and cost. Where you run a workload is a question about latency, data gravity and sovereignty. As long as placement and execution are the same decision, every AI deployment becomes a real estate argument. Once a photonic fabric can make placement invisible to the workload, the two decisions separate. Training and heavy batch inference go where power and space are cheap. Latency-bound inference lands close to the user. The fabric arbitrates between them. This does not contradict the case for the central office, it absorbs it: the central office still wins for latency-bound edge inference, but that win is now a node in a graph rather than an isolated site.

I have some history with this problem. In 2018 I wrote about building at Telefonica what was probably the industry's first fully programmable multi-access edge computing platform, and the hardest part was never the compute. It was making distributed compute addressable, governable and billable as a coherent resource rather than a scatter of isolated sites. The technology around it has moved on considerably, but the unsolved problem is the same one Docomo is now attacking with photonics.

A practitioner's caution is in order. A fabric that makes national-scale GPUs behave as one pool is a testbed today, not a product. Docomo demonstrated it in a lab-grade programme. KT's edge connection is a plan, not a live deployment. Deterministic optical transport carrying commercial service level agreements under contended traffic is a materially harder thing than a controlled demonstration, and I would treat "as if co-located" the way I treat vendor energy savings figures: directionally real, quantitatively unproven at scale. The distance between a testbed that works and a fabric that carries production workloads is precisely the distance Open RAN spent five years crossing.

Still, the framing is the takeaway. The AI Grid conversation needs to move from siting to fabric. The operators that win will be the ones who can treat geographically dispersed compute as a single addressable resource, orchestrate workloads across it against real constraints, and price and settle access to it. That is a transport, orchestration and settlement problem before it is a property problem. The question is no longer which building holds the GPUs. It is who owns the fabric that makes the buildings irrelevant.

Monday, July 6, 2026

The Agent Runtime Is Not the Agent Model

DTW Ignite in Copenhagen made one thing clear: the vendor community has decided that the path to autonomous networks runs through agent runtimes. NVIDIA introduced NemoClaw blueprints and the OpenShell secure runtime to give long-running agents policy guardrails and sandboxed access to telecom systems. AdaptKey is piloting security-hardened agents for self-healing 5G operations. ServiceNow is bringing Project Arc to the NOC, orchestrating incident response from alert to work order. NTT DATA is building anomaly agents that escalate to research agents for telemetry analysis. Synthetic data rounds out the stack, a pragmatic answer to the fact that more than half of operators say their most valuable network data is too sensitive to use.

This is genuine progress and I do not want to minimize it. Containment, auditability and policy enforcement are necessary conditions for letting agents touch production networks. An agent that cannot be sandboxed cannot be trusted, and an agent whose actions cannot be audited cannot be certified. The runtime layer has to be built.

Containment is not coordination

But look carefully at what these announcements govern: individual agents, operating within a single operator's domain, executing workflows that a human has scoped in advance. This is vertical governance. It answers the question of whether an agent is allowed to perform an action. It does not answer the question that autonomous networks will actually pose at scale: when two agents are each permitted to act, and their permitted actions conflict, who decides?

Consider a scenario that is closer than most operators think. An enterprise logistics agent requests guaranteed throughput for a fleet of delivery robots. Simultaneously, a network energy agent, operating under its own perfectly valid mandate, is shutting down capacity in the same cluster to meet a sustainability target. Both agents are sandboxed. Both are auditable. Both are compliant with their policies. The runtime layer sees two well-behaved agents. The network sees a contradiction.

This is the problem I described in my previous post on network APIs. APIs were designed for developer access, not for agent-to-agent negotiation. Runtimes inherit the same blind spot. They secure the execution of each agent without providing any shared representation of the agentic plane itself.

What the meta-model requires

For agents to negotiate rather than collide, the industry needs a meta-model of the agentic plane: a topology of which agents exist and where they sit, an ontology so that an enterprise agent and a network agent mean the same thing by capacity, latency or priority, explicit authority boundaries defining what each agent may commit on behalf of its principal, shared state models so that negotiations reference the same view of the network, and audit trails that span negotiations rather than individual actions. None of the DTW announcements address this layer. They cannot, because it is not a product any single vendor can ship. It is a model the industry must agree on, the way it once agreed on network information models for OSS.

There is a familiar pattern here. The industry built firewalls before it built routing protocols for the internet's trust boundaries, and it spent two decades paying for the sequencing. We are building the firewalls of the agentic era first. The operators and standards bodies that formalize the agentic plane meta-model will define how enterprise AI and network AI transact for the next decade. The ones that stop at the runtime will discover that a network full of safely contained agents is not an autonomous network.

Wednesday, July 1, 2026

DTW Ignite 2026: The API Is Not Enough

I returned from DTW Ignite in Copenhagen with one conviction: the interface between enterprise applications and network infrastructure is about to change in a way the industry has not yet designed for.

Network APIs were never really about autonomous networks. That framing conflates two separate problems. APIs — CAMARA, GSMA Open Gateway, the decades of network exposure work that preceded them — were designed to let developers discover and consume network resources from outside the operator domain. Quality on Demand, location services, device status, number verification: clean REST interfaces exposed through a developer portal so that a programmer writing a B2B application could request a network capability and pay for it. Real progress on a real problem. But the problem was developer access, not network autonomy.

What is coming next is different in kind, not degree.

Enterprise AI agents are beginning to consume network infrastructure directly — not through a developer writing an integration, but autonomously, in real time, as part of executing a business objective. An industrial automation agent that needs guaranteed low-latency connectivity for a robotics fleet. A financial services agent that needs to provision a secure, isolated network path for a time-sensitive transaction. A logistics agent that needs to dynamically reserve bandwidth across multiple carrier domains as a shipment moves between jurisdictions. In none of these cases is there a developer in the loop. The agent has an intent, it needs network resources to fulfil it, and it needs to negotiate those resources with the network — now, at machine speed, without human mediation.

That negotiation cannot happen through a developer portal. It cannot happen through a static API catalogue with a PDF explaining what each endpoint does. The enterprise agent and the network need to speak to each other, and neither CAMARA nor MCP — whatever their respective merits — were designed for that conversation.

The network side of this exchange needs to be represented by network AI agents of its own: agents that can expose available capacity in real time, understand the constraints and commitments already in place, reason over competing demands, and negotiate resource allocation in a way that respects the network's operating boundaries. That is not a developer API. That is an autonomous counterparty.

And for those network agents to function — to negotiate reliably, to be governed, to be audited, to avoid conflicting with each other across RAN, transport, core, and the operational layers of OSS and BSS — they need something the industry is not yet building: a meta-model of the agentic plane itself.

Operators building autonomous networks are doing the right foundational work. Network topology models. Data ontologies. Decision layers. Closed-loop control architectures. These give the automation layer a complete and current picture of the environment it is operating in. But agents operating on that network need an equivalent model of themselves. Every agent with an identity, a capability scope, an authority boundary, a state, a dependency graph, and an audit trail. An abstract topology and ontology of agents, sitting alongside the topology and ontology of the network.

Without that model, what looks like autonomous negotiation between enterprise AI and network AI is actually uncontrolled interaction between systems that cannot see each other. An enterprise agent requesting bandwidth does not know what the network agent is authorised to commit. The network agent does not know what other network agents have already promised. No shared representation, no conflict detection, no governance.

The developer exposure problem is largely solved, or at least well understood. The agent-to-agent negotiation problem has barely been framed. That is the conversation the industry needs to have, and Copenhagen convinced me we are not having it yet.

Thursday, May 21, 2026

Non-RT RIC, AI-RAN, and the AI Grid: Three Different Bets on the Future of the RAN


I have been asked a few times lately what the difference is between the Non-Real Time RIC and AI-RAN. The question itself tells you something. Both sit under the broad "AI in the RAN" umbrella, marketed aggressively by the same vendors, debated in the same conference sessions. But they are fundamentally different in architecture, ambition, and business model. And neither is quite the same as what NVIDIA formally branded the AI Grid at GTC 2026 — which is where the most important and most misread opportunity actually sits.

The Non-RT RIC: the pragmatic bet

The Non-RT RIC is an O-RAN defined software layer that sits in the Service Management and Orchestration layer above the RAN, not inside it. Control loops over one second. rApps for energy saving, traffic steering, slice assurance, automated optimization. Think of it as the evolution of Self-Organizing Networks, re-platformed on open interfaces with a proper application model and a genuinely lower barrier to entry — cloud-native and OSS skills are sufficient. No RAN silicon expertise required.

This is precisely why the early commercial traction is here, not in AI-RAN. AT&T is deploying Ericsson's SMO and Non-RT RIC to replace two legacy C-SON systems. TELUS has launched an RIC platform alongside its Open RAN rollout. Swisscom is deploying one for multi-technology network management. These are not trials. These are production decisions.

AI-RAN: real performance gains, speculative revenue

AI-RAN embeds AI natively into the RAN stack itself .The AI-RAN Alliance — founded in February 2024, now at 109 member companies — defines it across three working groups: AI-for-RAN, AI-and-RAN, and AI-on-RAN.

AI-for-RAN is the most mature: using AI to optimize the RAN itself — the scheduler, link adaptation, beamforming, interference management. T-Mobile and Ericsson have been trialing an AI-driven scheduler and link adaptation engine on a live 5G Advanced network since Q2 2025, targeting commercial deployment in Q3 2026. Nokia and NVIDIA, backed by a $1 billion equity partnership, are testing GPU-accelerated AI-RAN with BT, Elisa, NTT DOCOMO, and Vodafone.

AI-and-RAN is where the narrative gets more ambitious — and more speculative. The idea is that RAN sites become shared compute infrastructure, running both network workloads and enterprise AI workloads on the same hardware. The tower becomes a distributed AI compute node. New revenue streams. Operators escape the utility trap.

AI-on-RAN is the monetization layer for the above. The commercial mechanisms are still being defined. That tells you where the maturity is.

The AI Grid: follow NVIDIA's sequencing, not its marketing

At GTC 2026, NVIDIA formally introduced the AI Grid as a reference design — geographically distributed AI infrastructure, using the telco footprint to run inference workloads closer to users. The numbers are interesting: early Comcast benchmarks showed inference cost reductions of up to 76% versus centralized deployments. HPE, SpectroCloud, and others have already announced implementations aligned to the reference architecture.

I have used this concept in my own work for years to describe the evolution from isolated MEC deployments into a coherent, programmable distributed inference fabric. Good to see NVIDIA put a formal architecture behind it. But the marketing obscures a critical sequencing question.

NVIDIA's own GTC announcements noted that many operators are starting by lighting up existing wired edge sites — central offices and mobile switching offices — as AI Grids they can monetize today. The cell site layer is a later phase. AT&T's CTO Igal Elbaz has been direct about questioning the value of pushing compute all the way to the far edge to save one or two milliseconds of latency. T-Mobile's SVP of network infrastructure defined her AI edge strategy as what is at a data center at a mobile switching office. Verizon's CTO has flagged the cost and complexity of far-edge GPU deployments.

These are the three largest US operators. They are not being conservative for the sake of it. The economics are straightforward: central offices and mobile switching offices already have power, cooling, connectivity, and physical security. They aggregate traffic from hundreds of cell sites. The sub-500ms latency threshold that NVIDIA's own reference design targets is achievable from a well-positioned CO. It does not require a GPU at the tower — not for the use cases that have a business case today.

I have seen this movie before with MEC. The industry led with its most ambitious architectural vision, ran the infrastructure investment ahead of the demand, and recovered slowly. The AI Grid does not have to repeat that pattern.

What to actually do

Start with the Non-RT RIC. The contracts are being signed, the ecosystem is opening, the business case is defensible.

On AI-RAN, wait for AI-for-RAN where your vendors have credible near-term roadmaps. Treat AI-and-RAN at the cell site as a long term speculative option — worth tracking, too early to fund at scale.

On the AI Grid, follow NVIDIA's own sequencing rather than the brochure. Central offices and mobile switching offices first. Build the orchestration and service layer from there outward. Expand to the far edge when the use cases and economics justify it — not because a GPU manufacturer's demand forecast requires it.

The cell site AI Grid is a compelling long-term vision. The central office AI Grid is deployable today. In this industry, deployable usually wins.