Showing posts with label AI inference. Show all posts
Showing posts with label AI inference. Show all posts

Monday, July 13, 2026

AI monetization for operators: separating revenue from cost avoidance

Every operator earnings call now features AI prominently. Listen closely, however, and most of what is described as "AI monetization" is nothing of the sort. It is cost avoidance cosplaying a revenue costume.

This distinction matters because the two require different investment logic, different organizational capabilities, and different patience horizons. Operators that blur them will misallocate capital. Operators that separate them have a chance at building genuine new B2B revenue lines — narrower than the hype suggests, but investable.

Three money flows, not one

AI touches operator economics through three distinct channels, and the discipline starts with refusing to aggregate them.

1. AI that reduces cost. Autonomous network operations, agentic customer care, energy optimization, predictive maintenance. This is real, it is happening, and it is the largest near-term financial impact of AI on operators. It is also not revenue. A dollar of opex avoided is valuable, but it does not create a new line of business, and it does not justify the "operators as AI companies" narrative. It justifies a leaner operator.

2. AI that defends existing revenue. Enterprises deploying AI workloads have new connectivity requirements: deterministic performance, low latency to inference endpoints, secure private connectivity to GPU capacity, data-gravity-aware networking. Operators that serve these requirements protect and modestly grow their core B2B connectivity business. This is differentiated connectivity for the AI era — important, defensible, but fundamentally an evolution of what operators already sell.

3. AI that creates new revenue. This is the category everyone wants to talk about and the one that deserves the most scrutiny. It exists, but it is narrower than most strategy decks admit.

The four credible new revenue lines

Having spent the last two years working on AI infrastructure with operators and vendors on both sides of the Atlantic, I see four B2B revenue opportunities that survive contact with commercial reality. They are not equal — they differ in demand maturity, margin profile and time horizon, and they should be funded accordingly.

Sovereign AI capacity — GPU-as-a-service and AI factories — is the most immediate and the most misunderstood. Demand is real and policy-driven, concentrated in regulated sectors; the margin profile is low-to-mid, because the business is capex-heavy and carries utilization risk; and the revenue is available now. The demand side is genuine: governments, healthcare systems, defense, financial services and public administrations in Europe increasingly cannot — or will not — run inference on US hyperscaler infrastructure under foreign jurisdiction. Operators hold assets that map remarkably well to this demand: national data center footprints, energy contracts, security clearances, sovereign trust, and enterprise sales relationships.

Telefónica's recent national rollout of edge-based GPU-as-a-service in Spain is instructive. The underlying edge platform was architected years earlier — I led the team that built and productized it — and for years the business case was marginal on enterprise use cases alone. What changed was not the technology. It was the arrival of sovereign AI demand, which finally gave the infrastructure a paying anchor tenant profile. The lesson generalizes: edge and distributed compute investments become fundable when sovereignty is the demand driver, not the garnish.

The caution: this is a capex-intensive, utilization-sensitive business competing against hyperscalers with structurally lower unit costs. Operators win where sovereignty, data residency and proximity are binding constraints — and lose everywhere else. The addressable market is the regulated slice of national demand, not "the AI market."

Edge inference is real but earlier than its promoters claim. Demand exists where latency or data gravity bind; margins are mid-range; and the horizon is two to five years before this becomes a broad product line. The use cases that pay today are those where physics or data gravity make centralized inference impossible: industrial vision, real-time media production, autonomous operations in ports and factories. I have seen these work commercially. But the buyer set is narrow, and each engagement still resembles a system integration project more than a product sale. This becomes a scalable product line when agentic AI workloads distribute themselves across infrastructure tiers — the architecture I have described elsewhere as the AI Grid. That shift is underway, not arrived.

Data and trust services are the sleeper. Deepfake detection on voice calls, branded and verified calling, identity assurance for AI agents, provenance services. These are small revenue lines today, but they are high-margin, they monetize immediately, they sit directly on operator trust assets that hyperscalers cannot replicate, and demand grows with every AI-enabled fraud headline. For a B2B operator, this category has the best margin-to-capex ratio of the four.

Network APIs are the line whose trajectory has changed most in the past two years. The strategic logic has always been sound — AI agents will need to programmatically request network resources, quality on demand, location, verification — and the commercial signals are finally following: revenues are growing, aggregation initiatives have consolidated distribution, and enterprise visibility is rising with every agentic deployment that needs verified identity or guaranteed quality. It remains the earliest-stage of the four, and the AI agent wave — rather than developer evangelism — is what gives it genuine demand pull. I would invest now to be positioned, while sizing near-term revenue expectations with discipline; the inflection is likely in the second half of the decade.

The monetization ladder

Across all four lines, there is a ladder that determines margin and defensibility:

Sell capacity → sell platform → sell outcomes. Capacity here includes every consumption-metered unit: GPU hours, tokens, gigabits.

Selling raw capacity — GPU hours, token-metered inference, connectivity — is rung one: necessary, low-margin, commoditizing from day one. Tokens deserve a specific caution here: metering in tokens rather than GPU-hours changes the billing unit, not the business. An operator selling tokens against someone else's models and someone else's stack is still selling capacity, at prices that will be set by the most efficient infrastructure provider in the market. Selling a platform — inference-as-a-service with orchestration, security, compliance tooling — is rung two, where margins improve and switching costs appear. Selling outcomes — a fraud-detection rate, a production workflow, a compliant AI deployment for a hospital group — is rung three, where the economics finally resemble a services business worth building.

Operators historically stall at rung one. The reasons are organizational, not technological: product management that thinks in network elements rather than buyer problems, sales forces compensated on connectivity, and business cases that demand payback before the platform layer has time to mature. The operators that climb the ladder will be those that treat AI monetization as a product management and go-to-market transformation, not an infrastructure deployment.

What the buyer actually pays for

A final discipline. In every commercially successful case I have worked on, the enterprise buyer was not paying for "AI." They were paying for a constraint to be removed: data that could not leave the country, latency that broke the use case, a fraud pattern that was costing millions, a compliance requirement that blocked deployment. Price the constraint, not the technology. The moment an operator's AI proposition cannot name the constraint it removes, it is a science project.

Three questions before approving any operator AI business case

  1. Which of the three money flows is this — cost, defense, or new revenue? If the answer mixes them, send it back.
  2. What binding constraint does the buyer pay to remove, and why is an operator structurally better placed to remove it than a hyperscaler or an integrator? Sovereignty, proximity and trust are acceptable answers. "We have a network" is not.
  3. Where does this sit on the capacity–platform–outcome ladder, and what is the credible path up? Rung-one economics with rung-three ambitions is where operator AI investments go to die.

The AI B2B opportunity for operators is real. It is also smaller, slower and more demanding of commercial discipline than the current narrative suggests. The winners will not be the operators with the most GPUs. They will be the ones that can tell the difference between a cost saving, a defended revenue and a new business — and fund each accordingly.

Thursday, July 9, 2026

Operators Lean In On AI Grid Location


Earlier this week I argued that the AI Grid debate needs to move on from where you place a GPU to whether geographically dispersed compute can behave as a single fabric. I stand by that. But a story that has been building across the press this week is a useful reminder that the location question, the one I have been answering the same way for two years, is now being settled in public by the people who actually own the radio networks. And they are settling it against the tower.

The reporting is consistent. Light Reading describes Nokia and Nvidia's AI-RAN proposition running into telco resistance. Verizon, Vodafone, Orange and, notably for me, Telus have all raised doubts about putting graphics processing units into the radio access network. AT&T's chief technology officer has cast public doubt on the case for AI compute at the far edge. The enthusiasm for GPU-in-the-RAN comes from two operators, T-Mobile US and SoftBank, and almost no one else. Much of the rest of the industry is looking at Intel's newer CPUs for its open RAN rollouts rather than filling cell sites with accelerators.

I want to be precise about what this does and does not prove. It does not prove that AI in the RAN is a bad idea. Applying machine learning to scheduling, link adaptation and energy management inside the baseband is real, it is shipping, and Ericsson's AI-in-RAN software subscription is a reasonable way to bring it into existing hardware. What the operators are rejecting is narrower and more specific. They are rejecting the proposition that the cell site should become a general-purpose AI inference venue, stuffed with GPUs, monetised by hosting third-party workloads at the edge of the network. That is the proposition I have said for two years does not survive contact with power, cooling, space, security and, above all, the absence of a monetisation model.

My position has been that AI Grid deployment begins at the central office and the mobile switching office, not the cell site, because every physical and commercial constraint favours the aggregation point over the tower. The reasoning was never controversial to anyone who has stood in both kinds of building. A central office has power feeds, environmental control, physical security and fibre already in place. A cell site has a cabinet, a limited power budget and a landlord. When Verizon, Vodafone, Orange and Telus decline to put GPUs at the far edge, they are not making a new argument. They are confirming an old one, and they are confirming it with capital allocation decisions rather than conference slides, which is the only confirmation that counts.

There is a workstream reason this caught my eye. Telus appearing on the skeptics' list is consistent with what I see in the market: operators that are serious about autonomous operations are also the ones being disciplined about where AI compute physically lands. Those two forms of discipline are related. An operator that thinks clearly about the economics of edge inference tends to think clearly about the economics of everything else in the network.

The AI-RAN enthusiasm gap also matters for how we read vendor claims. When a technology has two vocal operator champions and a longer list of vocal operator skeptics, that is the signature of a capability that has been field-validated in specific conditions but not commercially validated across the market. I have made this distinction before and it applies cleanly here. SoftBank's agentic AI-RAN demonstrations and T-Mobile's Nvidia-backed edge trials are real engineering. They are not yet evidence that the model generalises to operators with different cost structures, different energy prices and different enterprise demand. Treat a two-operator enthusiasm as a pilot signal, not a market verdict.

So where does this leave the fabric argument I made last week? Exactly where I left it, and stronger. The operators are removing the least defensible node from the AI Grid, the cell site as inference host, which clears the ground for the argument that actually matters. Once you accept that heavy inference will not live at the tower, the interesting question becomes how you knit central offices, regional data centres and a small number of genuinely latency-bound edge sites into one addressable pool. The industry spent this week deciding where the compute will not go. That is progress. The harder decision, who owns the fabric that arbitrates across the places it will go, is still open, though.

Tuesday, July 7, 2026

AI Grid: Fabric vs Location Considerations

The debate about where AI compute belongs in a telecom network has been framed as a location question from the start. Do you put the GPUs at the cell site, the central office, the regional data centre, or the hyperscale campus? I have argued consistently that the honest answer begins at the central office and the mobile switching office, because power, cooling, fibre, physical security and latency sufficiency all favour those sites over the tower. That position has not changed. But two announcements from Asia this week suggest the more consequential question is no longer where the compute sits. It is whether the compute behaves as one pool regardless of where it sits.

NTT Docomo disclosed a nationwide testbed it calls GPU over APN. It pools graphics processing units spread across eight locations in five Japanese cities and presents them to a workload as a single platform, connected over the all-photonics network that NTT Group has been building under its IOWN programme. Docomo describes it as the realisation of its AI-Centric ICT Platform concept, part of what the group now labels AIOWN, its AI-native infrastructure. Strip away the acronyms and the claim is precise and significant: distributed GPUs, addressed as if co-located, over deterministic optical transport.

KT made the point from the other direction. Its new chief executive committed 18 trillion won, roughly 11.7 billion dollars, over three years, including 3.26 billion for one gigawatt of AI data centre capacity and a plan to connect that centralised infrastructure with edge sites serving low-latency workloads such as autonomous vehicles and industrial robotics. One operator is making dispersed compute act centralised. The other is extending centralised compute out to the edge. Both are describing the same thing from opposite ends, which is a compute fabric rather than a compute site.

This matters because it decouples two decisions the industry keeps conflating. Where you place a GPU is a question about power, land and cost. Where you run a workload is a question about latency, data gravity and sovereignty. As long as placement and execution are the same decision, every AI deployment becomes a real estate argument. Once a photonic fabric can make placement invisible to the workload, the two decisions separate. Training and heavy batch inference go where power and space are cheap. Latency-bound inference lands close to the user. The fabric arbitrates between them. This does not contradict the case for the central office, it absorbs it: the central office still wins for latency-bound edge inference, but that win is now a node in a graph rather than an isolated site.

I have some history with this problem. In 2018 I wrote about building at Telefonica what was probably the industry's first fully programmable multi-access edge computing platform, and the hardest part was never the compute. It was making distributed compute addressable, governable and billable as a coherent resource rather than a scatter of isolated sites. The technology around it has moved on considerably, but the unsolved problem is the same one Docomo is now attacking with photonics.

A practitioner's caution is in order. A fabric that makes national-scale GPUs behave as one pool is a testbed today, not a product. Docomo demonstrated it in a lab-grade programme. KT's edge connection is a plan, not a live deployment. Deterministic optical transport carrying commercial service level agreements under contended traffic is a materially harder thing than a controlled demonstration, and I would treat "as if co-located" the way I treat vendor energy savings figures: directionally real, quantitatively unproven at scale. The distance between a testbed that works and a fabric that carries production workloads is precisely the distance Open RAN spent five years crossing.

Still, the framing is the takeaway. The AI Grid conversation needs to move from siting to fabric. The operators that win will be the ones who can treat geographically dispersed compute as a single addressable resource, orchestrate workloads across it against real constraints, and price and settle access to it. That is a transport, orchestration and settlement problem before it is a property problem. The question is no longer which building holds the GPUs. It is who owns the fabric that makes the buildings irrelevant.

Monday, July 6, 2026

The Agent Runtime Is Not the Agent Model

DTW Ignite in Copenhagen made one thing clear: the vendor community has decided that the path to autonomous networks runs through agent runtimes. NVIDIA introduced NemoClaw blueprints and the OpenShell secure runtime to give long-running agents policy guardrails and sandboxed access to telecom systems. AdaptKey is piloting security-hardened agents for self-healing 5G operations. ServiceNow is bringing Project Arc to the NOC, orchestrating incident response from alert to work order. NTT DATA is building anomaly agents that escalate to research agents for telemetry analysis. Synthetic data rounds out the stack, a pragmatic answer to the fact that more than half of operators say their most valuable network data is too sensitive to use.

This is genuine progress and I do not want to minimize it. Containment, auditability and policy enforcement are necessary conditions for letting agents touch production networks. An agent that cannot be sandboxed cannot be trusted, and an agent whose actions cannot be audited cannot be certified. The runtime layer has to be built.

Containment is not coordination

But look carefully at what these announcements govern: individual agents, operating within a single operator's domain, executing workflows that a human has scoped in advance. This is vertical governance. It answers the question of whether an agent is allowed to perform an action. It does not answer the question that autonomous networks will actually pose at scale: when two agents are each permitted to act, and their permitted actions conflict, who decides?

Consider a scenario that is closer than most operators think. An enterprise logistics agent requests guaranteed throughput for a fleet of delivery robots. Simultaneously, a network energy agent, operating under its own perfectly valid mandate, is shutting down capacity in the same cluster to meet a sustainability target. Both agents are sandboxed. Both are auditable. Both are compliant with their policies. The runtime layer sees two well-behaved agents. The network sees a contradiction.

This is the problem I described in my previous post on network APIs. APIs were designed for developer access, not for agent-to-agent negotiation. Runtimes inherit the same blind spot. They secure the execution of each agent without providing any shared representation of the agentic plane itself.

What the meta-model requires

For agents to negotiate rather than collide, the industry needs a meta-model of the agentic plane: a topology of which agents exist and where they sit, an ontology so that an enterprise agent and a network agent mean the same thing by capacity, latency or priority, explicit authority boundaries defining what each agent may commit on behalf of its principal, shared state models so that negotiations reference the same view of the network, and audit trails that span negotiations rather than individual actions. None of the DTW announcements address this layer. They cannot, because it is not a product any single vendor can ship. It is a model the industry must agree on, the way it once agreed on network information models for OSS.

There is a familiar pattern here. The industry built firewalls before it built routing protocols for the internet's trust boundaries, and it spent two decades paying for the sequencing. We are building the firewalls of the agentic era first. The operators and standards bodies that formalize the agentic plane meta-model will define how enterprise AI and network AI transact for the next decade. The ones that stop at the runtime will discover that a network full of safely contained agents is not an autonomous network.

Wednesday, July 1, 2026

DTW Ignite 2026: The API Is Not Enough

I returned from DTW Ignite in Copenhagen with one conviction: the interface between enterprise applications and network infrastructure is about to change in a way the industry has not yet designed for.

Network APIs were never really about autonomous networks. That framing conflates two separate problems. APIs — CAMARA, GSMA Open Gateway, the decades of network exposure work that preceded them — were designed to let developers discover and consume network resources from outside the operator domain. Quality on Demand, location services, device status, number verification: clean REST interfaces exposed through a developer portal so that a programmer writing a B2B application could request a network capability and pay for it. Real progress on a real problem. But the problem was developer access, not network autonomy.

What is coming next is different in kind, not degree.

Enterprise AI agents are beginning to consume network infrastructure directly — not through a developer writing an integration, but autonomously, in real time, as part of executing a business objective. An industrial automation agent that needs guaranteed low-latency connectivity for a robotics fleet. A financial services agent that needs to provision a secure, isolated network path for a time-sensitive transaction. A logistics agent that needs to dynamically reserve bandwidth across multiple carrier domains as a shipment moves between jurisdictions. In none of these cases is there a developer in the loop. The agent has an intent, it needs network resources to fulfil it, and it needs to negotiate those resources with the network — now, at machine speed, without human mediation.

That negotiation cannot happen through a developer portal. It cannot happen through a static API catalogue with a PDF explaining what each endpoint does. The enterprise agent and the network need to speak to each other, and neither CAMARA nor MCP — whatever their respective merits — were designed for that conversation.

The network side of this exchange needs to be represented by network AI agents of its own: agents that can expose available capacity in real time, understand the constraints and commitments already in place, reason over competing demands, and negotiate resource allocation in a way that respects the network's operating boundaries. That is not a developer API. That is an autonomous counterparty.

And for those network agents to function — to negotiate reliably, to be governed, to be audited, to avoid conflicting with each other across RAN, transport, core, and the operational layers of OSS and BSS — they need something the industry is not yet building: a meta-model of the agentic plane itself.

Operators building autonomous networks are doing the right foundational work. Network topology models. Data ontologies. Decision layers. Closed-loop control architectures. These give the automation layer a complete and current picture of the environment it is operating in. But agents operating on that network need an equivalent model of themselves. Every agent with an identity, a capability scope, an authority boundary, a state, a dependency graph, and an audit trail. An abstract topology and ontology of agents, sitting alongside the topology and ontology of the network.

Without that model, what looks like autonomous negotiation between enterprise AI and network AI is actually uncontrolled interaction between systems that cannot see each other. An enterprise agent requesting bandwidth does not know what the network agent is authorised to commit. The network agent does not know what other network agents have already promised. No shared representation, no conflict detection, no governance.

The developer exposure problem is largely solved, or at least well understood. The agent-to-agent negotiation problem has barely been framed. That is the conversation the industry needs to have, and Copenhagen convinced me we are not having it yet.

Thursday, May 21, 2026

Non-RT RIC, AI-RAN, and the AI Grid: Three Different Bets on the Future of the RAN


I have been asked a few times lately what the difference is between the Non-Real Time RIC and AI-RAN. The question itself tells you something. Both sit under the broad "AI in the RAN" umbrella, marketed aggressively by the same vendors, debated in the same conference sessions. But they are fundamentally different in architecture, ambition, and business model. And neither is quite the same as what NVIDIA formally branded the AI Grid at GTC 2026 — which is where the most important and most misread opportunity actually sits.

The Non-RT RIC: the pragmatic bet

The Non-RT RIC is an O-RAN defined software layer that sits in the Service Management and Orchestration layer above the RAN, not inside it. Control loops over one second. rApps for energy saving, traffic steering, slice assurance, automated optimization. Think of it as the evolution of Self-Organizing Networks, re-platformed on open interfaces with a proper application model and a genuinely lower barrier to entry — cloud-native and OSS skills are sufficient. No RAN silicon expertise required.

This is precisely why the early commercial traction is here, not in AI-RAN. AT&T is deploying Ericsson's SMO and Non-RT RIC to replace two legacy C-SON systems. TELUS has launched an RIC platform alongside its Open RAN rollout. Swisscom is deploying one for multi-technology network management. These are not trials. These are production decisions.

AI-RAN: real performance gains, speculative revenue

AI-RAN embeds AI natively into the RAN stack itself .The AI-RAN Alliance — founded in February 2024, now at 109 member companies — defines it across three working groups: AI-for-RAN, AI-and-RAN, and AI-on-RAN.

AI-for-RAN is the most mature: using AI to optimize the RAN itself — the scheduler, link adaptation, beamforming, interference management. T-Mobile and Ericsson have been trialing an AI-driven scheduler and link adaptation engine on a live 5G Advanced network since Q2 2025, targeting commercial deployment in Q3 2026. Nokia and NVIDIA, backed by a $1 billion equity partnership, are testing GPU-accelerated AI-RAN with BT, Elisa, NTT DOCOMO, and Vodafone.

AI-and-RAN is where the narrative gets more ambitious — and more speculative. The idea is that RAN sites become shared compute infrastructure, running both network workloads and enterprise AI workloads on the same hardware. The tower becomes a distributed AI compute node. New revenue streams. Operators escape the utility trap.

AI-on-RAN is the monetization layer for the above. The commercial mechanisms are still being defined. That tells you where the maturity is.

The AI Grid: follow NVIDIA's sequencing, not its marketing

At GTC 2026, NVIDIA formally introduced the AI Grid as a reference design — geographically distributed AI infrastructure, using the telco footprint to run inference workloads closer to users. The numbers are interesting: early Comcast benchmarks showed inference cost reductions of up to 76% versus centralized deployments. HPE, SpectroCloud, and others have already announced implementations aligned to the reference architecture.

I have used this concept in my own work for years to describe the evolution from isolated MEC deployments into a coherent, programmable distributed inference fabric. Good to see NVIDIA put a formal architecture behind it. But the marketing obscures a critical sequencing question.

NVIDIA's own GTC announcements noted that many operators are starting by lighting up existing wired edge sites — central offices and mobile switching offices — as AI Grids they can monetize today. The cell site layer is a later phase. AT&T's CTO Igal Elbaz has been direct about questioning the value of pushing compute all the way to the far edge to save one or two milliseconds of latency. T-Mobile's SVP of network infrastructure defined her AI edge strategy as what is at a data center at a mobile switching office. Verizon's CTO has flagged the cost and complexity of far-edge GPU deployments.

These are the three largest US operators. They are not being conservative for the sake of it. The economics are straightforward: central offices and mobile switching offices already have power, cooling, connectivity, and physical security. They aggregate traffic from hundreds of cell sites. The sub-500ms latency threshold that NVIDIA's own reference design targets is achievable from a well-positioned CO. It does not require a GPU at the tower — not for the use cases that have a business case today.

I have seen this movie before with MEC. The industry led with its most ambitious architectural vision, ran the infrastructure investment ahead of the demand, and recovered slowly. The AI Grid does not have to repeat that pattern.

What to actually do

Start with the Non-RT RIC. The contracts are being signed, the ecosystem is opening, the business case is defensible.

On AI-RAN, wait for AI-for-RAN where your vendors have credible near-term roadmaps. Treat AI-and-RAN at the cell site as a long term speculative option — worth tracking, too early to fund at scale.

On the AI Grid, follow NVIDIA's own sequencing rather than the brochure. Central offices and mobile switching offices first. Build the orchestration and service layer from there outward. Expand to the far edge when the use cases and economics justify it — not because a GPU manufacturer's demand forecast requires it.

The cell site AI Grid is a compelling long-term vision. The central office AI Grid is deployable today. In this industry, deployable usually wins.


Monday, March 16, 2026

The philosophical problem with agentic AI


Jensen Huang’s address at GTC gave me a lot to think about. So much so that I decided to drive to Sana Cruz for a taste of the ocean. I had to wait 30 minutes to get the table I wanted, just by the beach, in the sun but with a little shade… as I mistype table on my iPad, I am thankful for the autocorrect to sanitize my  somewhat boozy prose, while mostly appreciating the elegantly subtle blue underlying of the word batle, prompting me to consider “is that really what you meant to write, or do you meant table”?

I like that. I like that more than the blue pencil with the little star that insistently offers an AI assisted rewrite. Oh, sure, I am not a native English writer, so my grammar is somewhat tainted by the other 3 languages I might think in at any point in time. If I compound St Patrick and this weekend’s VI nations rugby results for France, you will understand if my writing is not the usual corporate polish. 

Having said that, I was at GTC for the first time, I listen to Jensen’s performance and I was left enlightened and a bit worried. By now, the headline and the sound bite out there must be the $1 Trillion line of sight on chip revenues for Nvidia over the next couple of years. Obviously, it is an extraordinary number. Unfathomable. Impossible to imagine for most of us. Almost impossible to think that we, collectively would spend 125$ ( at 8 billion people) of Nvidia stuff over the next couple of years. Surely that’s impossible. 

Unless this is not about need, but about demand.  Unless that demand is accelerated, compounded, exponentially nurtured beyond its natural curve. 

Essentially, what I retained from the presentation was that the larger the model, the more the interactions, the larger the demand, the faster and more the tokens have to be created to satisfy it. (I am sure AI could rewrite this sentence more elegantly, but screw it). The measurement unit becomes token per Watt,as it is a limiting factor for a given data center and tokens per second as it is the limiting factor for a given service. Jensen even alluded to the fact that they will factor in token per month grants in engineering packages as it becomes a productivity factor. 

The thesis for the 1T$ revenue relies on demand exploding and the emergence of low latency, high I/O token market. Low latency, high I/O is understandable. Multimodal, video models, requiring real time inferencing from vehicles, robots and generally physical AI will drive it. The demand explosion, though, even factoring in the integration of compute and AI in to its, devices, edges… if we look at adoption curves and industrial capacity is decades away,  not in 2 years. Unless…

Unless we are not the demand. Us, consumers, enterprises, industries, governments… Agentic AI and Clawdbot are just showing how, beyond automation, agency becomes a compounding factor. Agents, that you create, for specific purpose are understandable, useful controllable. 

Agents, that interpret your intent, create other agents to enact their interpretation, have access to your digital life, credit card, HR, accounts receivables, invoices, orders, security cameras, GPS movements better be accountable, auditable, controllable. Agents that create fleets of agents to parcel out their workload is where I have doubts. The d’explosion in demand relies on the hypothesis that we will let agents create agents consume tokens to satisfy our needs.

No doubt, we will have agents to control, audit, police agents, but it feels wrong to delegate tasks just because you can or for the concept of efficiency.

This is where the the philosophical debate clashes with the economic model. I learned that hard times create hard men. Hard men create easy times. Easy times create easy men. Easy men create hard times. We might have evolved from this adage, but I feel that, being a kinetic, rather than a literal learner, I’ve learned from trying. I’ve learned from friction. To this day, I write on my notebook with a pen. I don’t forget anything I write. I forget most of what I type. It feels to me that friction is an integral part of the learning experience. More, it is an integral part of the human experience. The taste for effort, trying the hard things, failing is not only what most mankind experience on a daily basis, it is also, at least for me a great  condition to happiness. I am infinitely happier labouring and succeeding than an automated, frictionless, efficient experience. Even with a better result.

As my children are about to enter the workforce, I am confronted daily to the question “what is a safe, fulfilling carrer?”. It used to be that medicine, law, engineering guaranteed a safe economic path. Nowadays, it looks like most entry level intellectual effort can easily, efficiently be replaced, and that agentic AI will only accelerate that trend. How are they supposed to master a domain they won’t be able to tinker and stumble? Maybe I am just an old fart and just like calculators and computers did not replace engineers, a higher level of abstraction will necessitate higher levels of intellectual efforts ? But this feels different. 

Particularly if compute keeps accelerating and artificial intelligence surpasses human intelligence, then what? What is the imperative to learn, labour, try, suffer, if is not necessary? Where do you draw the line between agents that help and augment and agents that enable and replace?

Until then, I’ll keep labouring and burdening you with poorly written posts, but somewhat original or at least unique, because they’re mine. I enjoy this table, i waited 30 minutes for because I chose it and waited for it. I am not sure it would have tasted better should my personal AI butler had booked it for me on my way there.


Wednesday, March 11, 2026

AI is a new G

returned from MWC 2026 with an uneasy feeling.

The telecommunications industry has long been defined by its generational leaps—each "G" marking a profound shift in capabilities, use cases, and societal impact.

2G brought reliable digital voice and SMS, enabling mass mobile communication. 3G introduced mobile data and picture messaging, laying the foundation for internet on the go. 4G powered the explosion of social media, apps, and always-on connectivity. 5G delivered massive bandwidth, fueling high-definition video streaming.

These evolutions followed a predictable cadence governed by 3GPP standards, with operators methodically upgrading infrastructure, spectrum, and devices in multi-year cycles. Parallel to this, the network itself transformed through virtualization: from SDN separating control and data planes, to disaggregating hardware from software, and evolving VNFs (Virtual Network Functions) into cloud-native CNFs (Cloud-native Network Functions). These shifts improved flexibility, scalability, and cost efficiency but remained incremental within the familiar "G" framework.

AI is entering telecom in silos—AI-RAN for spectrum and energy optimization, agentic AI in OSS for autonomous operations and predictive assurance, customer service copilots for intent-based support—delivering proven cost savings (e.g., 25-40% OPEX reductions in network ops, up to 35% energy efficiency). Yet these domain-specific wins rarely connect into a unified, end-to-end intelligence layer. Data stays fragmented across RAN, core, edge, OSS/BSS, leading to duplicated efforts, incomplete visibility, and "agent sprawl" risks. Industry sources highlight how silos impede multi-agent ecosystems and true autonomous networks.

This misconception manifests in several ways:

Viewing AI as incremental tech add-ons — Operators often pursue isolated pilots (e.g., AI-RAN trials, genAI copilots, or agentic OSS agents) expecting quick wins without addressing deeper structural issues.

Underplaying organizational and cultural complexity — AI demands far more than engineering upgrades. It requires breaking down legacy silos (RAN/IT/OSS/BSS), fostering cross-functional agility, upskilling thousands in ML ops/data governance, and driving cultural shifts to trust agentic systems. Cultural resistance, job security fears, and fragmented skills often stall progress, with many projects failing to move beyond pilots (only ~30% of genAI use cases reach production in some analyses). Organizational challenges—including change management and silo-breaking—as top barriers, yet leadership frequently delegates AI to a separate function rather than owning it as a CEO imperative.

Misjudging the scale of change needed — Unlike past "G" evolutions (hardware/spectrum-driven, standardized via 3GPP), AI is a software-defined, data-hungry, adaptive intelligence layer that reshapes workflows, decision-making, operating models, and even business identity (from connectivity provider to intelligent platform). Treating it as "just tech" ignores the need for unified data fabrics, intent-based orchestration, governed multi-agent ecosystems, and radical process redesign—efforts that can take years, not quarters, and demand massive internal rewiring.

New vendors (hyperscalers, specialized AI-RAN players, agentic platforms) disrupt legacy supplier models, while operating models evolve toward intent-driven, cloud-native, agent-orchestrated environments requiring cross-functional agility and new skills. Massive CAPEX uncertainty surrounds compute (GPUs, accelerators), high-bandwidth memory, power, and cooling—often in the hundreds of billions globally—amid unclear ROI timelines and risks like underutilization. AI excels at cost management through optimization, but revenue-generating services (e.g., enterprise AI platforms, GPUaaS, network APIs for AI workloads, personalized offerings) remain nascent for most operators. This imbalance—cost wins without broad revenue upside, vendor shifts, and compute investment risks—demands an AI strategy that starts with organization and operational models, not technology.

This underestimation risks turning AI from a greenfield opportunity into added complexity: persistent silos, agent sprawl, duplicated investments, and missed revenue potential. Proven cost optimizations are real, but without holistic transformation, operators may achieve efficiency gains while remaining commoditized pipes in an AI-driven world. Warning to operators: AI is not "plug-and-play." Underestimating its demands—starting with organization, leadership alignment, operating model redesign, and cultural renewal before heavy technology scaling—will lead to stalled initiatives, wasted CAPEX (especially on compute/infra), and competitive disadvantage. Frontrunners recognize AI as a radical reinvention requiring bold, enterprise-wide commitment; the rest risk being left behind as the intelligence generation unfolds.

Tuesday, February 10, 2026

Where Do Network Operators Go From Here? A View Ahead of MWC 2026

With Mobile World Congress just around the corner in Barcelona, the telecom sector finds itself at another inflection point. The headlines are familiar: ongoing layoffs across major operators, C-level reshuffles, persistent ARPU erosion, and debt structures that constrain organic investment. Vendors are already talking up 6G roadmaps while AI dominates conversations—both for aggressive OPEX reduction and tentative new revenue paths. Yet the near-term reality feels more evolutionary than revolutionary.

The recent wave of workforce reductions is not, in my view, primarily an AI story—at least not yet. It reflects the long tail of a structural shift that began over a decade ago: the gradual but relentless transition from proprietary telco platforms to cloud-native architectures. We are finally seeing the full operational benefits of user/control-plane separation, hardware/software disaggregation, widespread network virtualization, and centralized policy orchestration. These changes deliver greater automation, elastic scaling, and dramatically shorter development and validation cycles. The outcome is clear: managing a modern mobile network no longer requires the headcount levels of the previous era. Painful as the adjustment is, it is the inevitable consequence of borrowing proven cloud-native principles. Cost discipline is essential, but it is not a growth strategy. The more pressing question is how operators convert more reliable, elastic, and automated networks into sustainable revenue expansion.

Private Networks: Successes Exist, but They Remain Hard-Won

Private cellular networks continue to polarize opinion. Some portray them as a commercial disappointment; others point to hundreds of documented use cases. The reality sits firmly in between. Genuine deployments delivering positive returns do exist, particularly in verticals with high-value connectivity requirements and tolerance for tailored solutions. Energy (smart grids and remote monitoring), healthcare (indoor coverage in hospitals and clinics), large venues (stadiums and event spaces), mining (autonomous haulage and safety systems), and ports (crane automation and terminal logistics) stand out as segments where demand is tangible and economics can work. The common thread in successful cases is not technology alone but deployment philosophy: cloud-native designs that run on commodity hardware, leverage centralized intelligence, and minimize site-specific customization. When executed this way, private networks become scalable and margin-accretive rather than bespoke projects that drain resources. Operators who treat private 5G as an extension of their public edge and orchestration capabilities—rather than isolated silos—are better positioned to capture repeatable value.

Data: The Next Realistic Monetization Frontier

Beyond connectivity and private networks, operators sit on an underutilized asset: vast quantities of network-derived and network-transported data. Until recently most of this information has been siloed for internal analytics, dashboards, and regulatory reporting. That picture is beginning to change. Monetization remains nascent compared with the advertising-driven models of social platforms, yet the opportunity is material. API gateways that expose selected network and user context (location aggregates, mobility patterns, congestion signals, roaming events) represent only the surface layer. Consider a few practical illustrations:
  • Ride-hailing platforms could benefit from near-real-time insight into clusters of international roamers converging in a city district—an indicator of an upcoming conference, trade show, or major event. Pre-positioning drivers becomes more efficient, improving service levels and reducing wait times.
  • eSIM and travel-focused virtual operators could package value-added bundles—discounted car rentals, hotel reservations, restaurant bookings, or attraction tickets—targeted at detected travelers arriving in high-demand locations.
  • Navigation services (Google Maps, Waze, and equivalents) could gain from telco-sourced, fine-grained congestion and flow data that augments probe-vehicle inputs, especially in areas with sparse device coverage or during atypical events. Privacy and regulatory compliance are non-negotiable hurdles, as are competitive dynamics with hyperscalers and data aggregators. Success will depend on responsible data handling, anonymization at scale, clear value propositions for enterprise partners, and commercial models that avoid commoditization. Operators that can evolve from pure connectivity providers toward curated data intermediaries—leveraging their unique position across physical infrastructure, subscriber scale, and real-time network telemetry—stand to capture incremental revenue without requiring entirely new network builds. As we head to MWC 2026, the conversation will likely revolve around AI acceleration, 6G timelines, and edge monetization. Beneath the buzz, though, the fundamentals remain: disciplined cost management, selective private-network wins, and thoughtful exploration of data opportunities. What are you seeing in your markets? Are private networks crossing the chasm in specific verticals? And where do you place data monetization on the priority list for the next 18–24 months? I welcome your perspectives in the comments.

Thursday, January 29, 2026

Physical AI: How Network Operators Could Leverage Edge Computing for Smarter Robotics

As the telecom landscape evolves, one emerging trend that's catching my eye is Physical AI—the integration of advanced AI into physical devices like robots, enabling them to interact intelligently with the real world. With my background in telco-cloud strategy, I'm particularly intrigued by how network operators could position themselves as key enablers in this space. By providing low-latency edge infrastructure, telcos might unlock new revenue streams while supporting innovative applications that blend robotics, computer vision, and conversational AI.

In a recent analysis, I've been exploring how robots equipped with cameras and speakers could benefit from distributed AI processing at the network edge. This setup allows for real-time scene analysis, object detection, facial recognition, and natural language interactions with humans—all without relying solely on centralized clouds that introduce delays or high costs.

What is Physical AI?

Physical AI refers to AI systems embodied in hardware that perceive, reason, and act in physical environments. Unlike traditional AI that's confined to software, this involves robots or devices that use sensors (like cameras) to understand their surroundings and actuators (like speakers) to respond. The key challenge? Processing massive data streams in real time while maintaining privacy, efficiency, and low latency. This is where telco networks shine, with their distributed edge nodes offering compute power closer to the action.

Edge AI Inference: Powering Perception in Robotics

Operators could facilitate edge-based AI inference, where robots offload complex tasks like scene recognition, object identification, and facial analysis to nearby network edges. For instance, a service robot in a retail store uses its camera to scan the environment: edge inference quickly identifies products on shelves, detects customer faces for personalized greetings (with privacy safeguards), or recognizes obstacles to navigate safely. This sub-10ms processing avoids the pitfalls of cloud round-trips, reducing bandwidth usage and enabling seamless, responsive interactions.

Techniques like federated learning could further enhance this, allowing robots to fine-tune models collaboratively across distributed edges without sharing raw data—ideal for maintaining user privacy in sensitive scenarios.

Generative AI for Natural Language Conversations

Pair that with generative AI models running at the edge for conversational capabilities. Robots with speakers could engage in fluid, context-aware dialogues: a healthcare assistant bot recognizes a patient's face, infers emotional state from scene cues, and generates empathetic responses using natural language processing. Or in manufacturing, a collaborative robot converses with workers in real time—"Hand me the red tool"—while using object recognition to confirm and act.

By offering "AI-as-a-Service" at the edge, operators could provide scalable, usage-based access to these capabilities. Enterprises get high-performance AI without massive capex on private infrastructure, while telcos monetize their pervasive networks.

Real-World Opportunities and Examples

Consider verticals ripe for this:

  • Retail and hospitality: Robots greeting customers by name (via facial rec), recommending items based on scene analysis, and chatting naturally to assist.
  • Healthcare: Companion bots in hospitals using edge inference to monitor patient environments, detect falls, and converse to provide reminders or emotional support.
  • Logistics and manufacturing: Autonomous robots navigating warehouses, identifying inventory via objects/scenes, and collaborating verbally with human teams.
  • Smart cities: Public service bots patrolling areas, recognizing incidents (e.g., litter or crowds), and interacting with citizens through voice.

These use cases could drive B2B partnerships, where operators bundle connectivity with edge AI compute—potentially adding 10-20% to ARPU through premium services.

Considerations for Carriers

To capitalize, carriers might assess their edge footprints for AI readiness, pilot federated models for privacy, and collaborate with robot vendors or AI platforms. Challenges like energy efficiency and standardization remain, but the rewards in a growing Physical AI market make it worth exploring.

Wednesday, January 28, 2026

Distributed AI at the Edge: Opportunities for Telecom Networks in an Evolving AI Landscape

The rapid growth of AI applications is creating new demands on network infrastructure, particularly for low-latency, distributed processing close to end-users and devices. Rather than remaining focused solely on connectivity, telecom networks are increasingly positioning themselves to support distributed AI capabilities—where inference and even lightweight training can occur at the edge. This shift opens interesting possibilities for operators to play a more central role in the broader AI ecosystem. In a recent interview at FYUZ 2025 ( Telecom Infra Project's flagship event in Dublin), I had the opportunity to discuss these dynamics with TelecomTV . The conversation centered on a practical question: How might telco networks evolve from traditional mobile broadband platforms to ones that can meaningfully support distributed AI workloads?

The Emerging Demands on Networks for Distributed AI

AI inference, and in some cases lightweight training at the edge, benefits significantly from response times below 10 milliseconds and access to distributed parallel processing. Centralized cloud architectures face inherent limitations in these scenarios—issues such as data gravity, backhaul congestion, and rising energy requirements often make proximity to the data source or user essential. AI workloads tend to be compute- and power-intensive, and telecom networks already manage substantial energy footprints; integrating AI processing without thoughtful optimization could increase both costs and environmental impact. At the same time, the limitations of static resource allocation become more apparent—networks increasingly need mechanisms for dynamic, policy-aware traffic prioritization, capacity allocation, and workload steering.

How AI-Integrated RAN Can Support Distributed AI Capabilities

One approach carriers are exploring involves integrating AI capabilities directly into the Radio Access Network (AI RAN). This embeds intelligence into the radio layer, enabling distributed inference and lightweight training to take place across the network's existing footprint of base stations, central offices or MSOs, edge nodes, and fiber backhaul. The result is a pervasive mesh of compute resources located close to users and devices.

Distributed inference allows models to be partitioned and processed in parallel at multiple edge points, significantly reducing latency by keeping data local rather than sending it to distant centralized facilities. Where models need fine-tuning based on fresh, real-time data, techniques such as federated learning offer a way to train collaboratively across distributed locations while maintaining data privacy and avoiding the need to aggregate sensitive information centrally.

Internal Opportunities for Carriers

Carriers could apply these distributed AI capabilities to improve their own network operations. For example, predictive maintenance can become more effective when AI models analyze real-time sensor data from base stations to anticipate equipment issues, enabling proactive interventions that help reduce unplanned downtime.

Traffic management stands to benefit as well—distributed inference at the edge can forecast congestion patterns and dynamically adjust routing to preserve service quality during high-demand periods.

Energy optimization is another area of potential gain, with AI learning from usage patterns to make real-time decisions, such as reducing power to underutilized radio resources during quieter hours. In many cases, these internal improvements could deliver operational cost reductions of 20-30% while enhancing overall network reliability, often without requiring large-scale new investments in specialized AI hardware.

Enterprise Potential: The promise of AIaaS

From a business-to-business perspective, distributed AI at the edge could allow operators to offer "AI-as-a-Service" models to enterprises that require low-latency inference but lack the capital or desire to build their own edge infrastructure. Small and medium-sized enterprises across sectors such as manufacturing, retail, logistics, and others often face this constraint. By leveraging the operator's distributed edge, inference tasks can be offloaded on a usage-based basis, making high-performance AI more accessible without heavy upfront expenditure.

Real-world examples help illustrate the potential.

  • In manufacturing, autonomous robotics depend on real-time object detection and path planning; inference performed at the nearest base station can deliver sub-10ms decisions, avoiding production interruptions without the facility needing to deploy its own compute resources.
  • Field technicians in utilities or construction working with augmented reality tools can receive AI-generated diagnostics overlaid on live video feeds—processed at the edge for instant fault identification, such as detecting structural cracks, supporting faster decisions in remote settings.
  • Retail operations can use edge-based smart analytics to interpret camera feeds for customer behavior insights or immediate security alerts, generating millisecond-level responses without on-site servers.
  • In healthcare, wearables transmitting vital signs for anomaly detection (for instance, flagging potential cardiac events) can benefit from low-latency edge processing to deliver timely alerts, particularly valuable in rural or resource-constrained clinics.
  • Cloud gaming environments can also gain from edge-handled AI upscaling of graphics or intelligent NPC behavior, substantially reducing perceived lag for players and smaller studios that lack powerful local hardware.

By structuring these capabilities as on-demand, sliced services, operators could create additional revenue streams while enabling enterprises to adopt AI more broadly without prohibitive capital requirements.

Considerations for Moving Forward

Operators interested in these opportunities might begin by assessing their current latency profiles, edge compute footprint, and level of AI integration. From there, they could prioritize pilot deployments focused on inference before exploring federated training approaches for stronger privacy controls. Partnerships with cloud providers could help develop hybrid models that combine telco edge strengths with broader AI ecosystems. Early monetization might involve introducing "AI-Ready Connectivity" services—low-latency slices, edge GPU access, and intelligent routing designed for enterprises building AI-driven applications.

Telecom networks already offer a distinctive advantage: widespread, low-latency reach to millions of endpoints. Carriers that thoughtfully explore distributed AI capabilities could position themselves as important contributors to the evolving AI infrastructure landscape, potentially unlocking meaningful new value in a growing market.