By Ricardo Velhuco, Field CTO, Ocient
I spent last week in Rio de Janeiro at Conecta LATAM 2026, where four tracks — Telco Transformation LATAM, Fraud & Security in the AI Era, AI & Digital Transformation, and Network Transformation — all kept circling back to the same underlying tension. Every operator is building toward an autonomous network. Yet few are budgeting for the thing underneath it that makes autonomy possible in the first place: the data foundation.

I gave a keynote on exactly that gap, and I want to share the highlights of it here, because the reaction in the hallways afterward told me it landed on a nerve this region is feeling acutely.
What “AI-native” actually means
Start with the destination, because the industry throws “AI-native” around loosely. I’d argue it means a network with five specific properties:
- Autonomous: It changes itself without a ticket.
- Predictive: It forecasts state, not just detects it.
- Self-optimizing: It runs against a configurable objective — cost, energy, latency, coverage.
- Sovereign: Enforced control over which models see which data.
- Accountable: Every autonomous action explainable and replayable, months or years later.
That last one is the property most vendors skip past, and it’s the one that will matter most to regulators and boards alike.
Get those five right, and two products fall into place: AI-as-a-service that operators can sell to enterprises and governments, and an internal AI factory for operators own development. But none of that is possible without a data layer that can support it with full observability across every domain, one integrated view joining data and user planes, full history, real time, closed feedback loops, agentic access, a semantic layer, governance and lineage, and, critically, sustainable economics. That last point isn’t a commercial nicety. If your foundation costs more to grow than your data grows, autonomy is a budget conversation you’ll keep losing.
The number that sets the problem
To make this concrete, I ran the arithmetic on a representative LATAM tier-1 operator. Add up RAN counters, CDRs, IPFIX flow records, 5G signaling, and trace data at full fidelity, and you land at roughly 200 billion records a day — about 9.1 petabytes a year. That sounds enormous until you compare it to the traffic it describes: it’s 0.07% of the actual user traffic crossing the network.
Every operator I’ve worked with made the same six compromises, and I want to be clear about something I said on stage: every one of them was the right call at the time. Sample instead of retain, because full-fidelity storage was priced off spinning disk. Keep days instead of years, because retention cost scaled linearly. Silo by vendor and domain, because each system solved one problem well. Split telemetry, control, and user planes across different tools and teams. Batch through a lake and then a warehouse, because that was the only affordable route to analytics at volume. Query inside one domain only, because cross-domain queries were slow or impractical.
All six were priced off the same constraint. That constraint, the cost of storing and querying data at this scale, has moved. Which means two things follow, and I asked the room to sit with these for a second rather than rushing past them: your sampling rate is your uncertainty rate. Your retention window is your context window. Whatever you didn’t keep, your AI cannot reason about, no matter how good the model is.
The LATAM Lens: Law and threat
This isn’t abstract for this region. Retention here is a statutory question, and it’s uneven — Brazil mandates five years for mobile call records under ANATEL 477/2007, Colombia five years under Decreto 1704, Peru three years, Mexico now twenty-four months under the new LMTR, Chile at least one year for IP assignment, and Argentina with no general mandate currently in force at all. The irony I pointed out on stage: the data classes with the longest mandated retention are the smallest by volume, and the largest by volume — flow and signal data — carry no retention mandate whatsoever.
Sovereignty in LATAM, and worldwide, is a control problem, not a border problem — there’s no localization mandate, but adequacy gates under LGPD and Colombia’s Ley 1581 govern who may see the data, regardless of where it sits. And cybersecurity is already regulated here — ANATEL’s R-Ciber framework requires board-approved cyber policy and incident notification — but log retention isn’t part of it. That’s a gap operators will have to fill themselves before an auditor asks. On the fraud and security track specifically, we talked about why that gap matters: Salt Typhoon showed state-sponsored actors dwelling in telecom networks for years before detection, and Kaspersky logged over a million blocked ransomware attempts against Latin American organizations in under a year. The threat model is long-dwell. A data foundation with days of retention can’t support an investigation that needs years.
The inverse and what it takes to build it
With OcientAIQ™, we are building for the inverse of those six compromises: retain across every domain and layer, reach back across years of history instead of days, join across domains in a single query, prove every answer, keep it sovereign with per-dataset control over where it lives, and serve it as governed, agent-ready data products.
The results we’re seeing in production back it up. For example, a tower-dump query that used to take 41 minutes now runs in 16 seconds; with roughly 80% lower total cost of ownership at petabyte scale than the retention stack it replaced.
If there’s one thing I’d ask every one of you reading this: run the volumetric arithmetic on your own network. Every domain, every data type, sovereign, secure. That’s the foundation autonomous LATAM networks need, and it’s the one that we’ve built with OcientAIQ.