Salesforce says Zero Copy now handles 120T rows; raises a metadata control fight
Salesforce’s engineering blog says its Zero Copy system scaled from 1 trillion to 120 trillion rows monthly by moving from query to file federation. If tha…
Edward Mullen ·

Salesforce’s engineering blog reports scaling zero-copy from 1 trillion to 120 trillion rows per month to serve petabyte-scale AI workloads. If that trajectory is followed broadly, file federation will displace copy-based pipelines as the primary enterprise sharing layer within twelve months, centralizing metadata and routing decisions across clouds under vendor control and creating an opaque third‑party data supply chain.
This is, so far, single-thread reporting — Salesforce Engineering Blog only, no independent confirmation. No one in the reported packet is on the record. The headline number is large and the architectural pivot is clear; what’s not in the blog is how governance, audit rights, and metadata ownership change once a vendor-managed layer becomes the map for where enterprise data lives.
The number is the hook; the architecture is the story In a post on the Salesforce Engineering Blog, the company describes a shift from Query Federation to File Federation, presented as the mechanism behind the jump from 1 trillion to 120 trillion rows per month to support “petabyte-scale AI workloads.” The blog is a vendor source, not independently replicated, which means the claim should be treated as a directional signal about architecture, not a measured industry benchmark. The move from sending queries into remote stores to federating access to files (and their locations) elevates the pointer index—the catalog and routing layer—to first-class infrastructure.
The consequence is architectural: the system that manages pointers becomes the gatekeeper for cross-cloud work, conditioning every downstream AI job on a third-party-controlled index. Salesforce describes this transition in the headline language: "Scaling Zero Copy from 1 Trillion to 120 Trillion Rows with File Federation."
What the blog shows—and what it doesn’t
The post provides one concrete metric—“1 trillion to 120 trillion rows monthly”—and one mechanism—“File Federation”—to justify that scale-up to “petabyte-scale AI workloads.” It does not specify the baseline environment, underlying storage vendors, query profiles, or hardware involved; the absence of apples-to-apples benchmarks, workload mixes, or failure modes leaves open whether the gains come from caching, schema constraints, or routing heuristics rather than the federation model alone. Absent are reproducibility details and any description of how the system behaves when pointers go stale, permissions drift, or schemas evolve mid-run.
Those omissions matter more to a GC or CISO than the headline number. Once access paths become references rather than copies, the metadata—URIs, catalogs, lineage, and access-control bindings—becomes the substrate for audit and incident response. If your legal discovery or compliance audit depends on reconstructing who resolved which pointer when, the logging, retention, and jurisdiction of that metadata decide your exposure, not the raw throughput number.
The choke point moves from data to metadata
Traditional “copy-based” pipelines spray data across object stores and warehouses, creating well-known governance headaches—duplication, drift, and egress bills. Zero-copy federation flips the problem: you reduce duplication but move the dependency to a routing and indexing tier that decides, at run-time, what gets accessed. The result is a nascent control point where policy, identity, lineage, and cost meet, often inside a vendor-managed service.
If that layer sits with a vendor, then so do the choke points that matter during incidents, audits, and contract renewals. The vendor's own description of the work—"Scaling Zero Copy from 1 Trillion to 120 Trillion Rows with File Federation"—is the only explicit phrasing Salesforce provides about the change in architecture.
Why the obvious read—“it’s a performance win”—misses the risk The dominant interpretation will be that this is a cost/performance optimization enabling petabyte-scale AI without duplicating data. That may be true technically, but it obscures the shift in bargaining power.
When pointers and metadata are the product, the owner of the catalog controls routing, observability, and, in practice, which workloads can run under which policies. That is not just a storage decision—it’s discovery scope, incident forensics, and regulator-facing evidence all rolled into an API.
The blog does not say who owns the pointer graph, where logs live, how long they are retained, or how tenancy segregation is assured across jurisdictions. It also does not specify whether customers can export the full catalog and replay resolution histories into their own SIEM or eDiscovery tools, or whether they need the vendor’s service online to validate past access.
Those omissions don’t negate the result; they identify the pressure points for governance and procurement in real deployments. Note the blog's framing: "Scaling Zero Copy from 1 Trillion to 120 Trillion Rows with File Federation," which is the single explicit, verbatim claim we can tie to Salesforce's post.
The counterargument: federation can improve governance—if you own the pointers One obvious objection is that zero-copy designs can strengthen governance by eliminating uncontrolled data sprawl and centralizing policy enforcement.
If the customer operates the catalog on their tenancy, with exportable logs and bring-your-own KMS and identity, you could see better discovery hygiene and tighter blast-radius control than with proliferated copies. The blog does not engage with this design space, so it remains an untested variable from this single source.
What shifts for GCs and CISOs over the next year Treat file federation as a contract term, not just an architectural choice. If you are contemplating zero-copy for AI workloads, the negotiation fulcrum is metadata custody: who holds the pointer graph, where is it stored, what are the retention and deletion guarantees, and can you export and validate it independent of the vendor’s control plane.
Ask for auditor-ready replay of pointer resolutions, evidence that access decisions are deterministic and re-producible, and clarity on how cross-border transfers are logged when federated paths cross regions.
Watch for three market signals. First, RFPs that list “zero-copy file federation” or equivalent metadata-catalog control as a must-have capability—this would indicate normalization of the architecture into procurement, not just engineering.
Second, emergence of third-party marketplaces or brokers for federated metadata routing and lineage, a tell that the index itself is becoming a tradable surface. Third, audit reports and SOC/ISO mappings that explicitly describe pointer custody, replayability, and tenant isolation—evidence that governance for the metadata tier is catching up to the data tier.
The narrow claim we can source—and the broader risk it implies Salesforce’s blog is a vendor assertion, not a peer-reviewed result or independent benchmark. It says a shift from Query Federation to File Federation enabled a jump from 1 trillion to 120 trillion rows monthly to run “petabyte-scale AI workloads.” That is the sum of the published facts. The broader reading—that metadata is now the supply chain—is analysis grounded in how zero-copy architectures work when scaled and monetized.
The company’s headline is doing double duty: it announces a performance story and, implicitly, an architectural boundary change. The consequence is architectural: the system that manages pointers becomes the gatekeeper for cross-cloud work, conditioning every downstream AI job on a third-party-controlled index. Salesforce describes this transition in the headline language: "Scaling Zero Copy from 1 Trillion to 120 Trillion Rows with File Federation."