How agentic coding vendors legally decouple the protected "Customer Data" from the human reasoning trace that produced it — and harvest the latter at industrial scale.
We are technologists, not attorneys. We report two things only: what these systems capture, and what the vendors’ own terms literally say — both verifiable right here. What a court would make of any of it is a question for your counsel, not us. Nothing here is legal advice or creates an attorney-client relationship. Verify every claim in the Evidence Library →
Leading AI vendors quietly redefined what counts as yours.
Over the last 14 months, major AI coding and agentic platforms quietly expanded an old SaaS term — “Usage Data” — far beyond what it used to mean. What was once just uptime logs and crash reports now includes the detailed record of how employees reason, correct, and decide.
Plain-text training has hit diminishing returns. The most valuable asset in an enterprise AI contract is no longer the customer’s data they want to protect. It’s the “reasoning trace” — captured through the telemetry these systems collect by design, and exactly the kind of signal needed to improve the reasoning capabilities of the next generation of models.
In traditional cloud and SaaS agreements, “Usage Data” or “telemetry” meant narrow operational metrics: server uptime, latency, memory diagnostics, crash reports. The actual customer data, designs, content and cognitive flow of how people worked stayed in the protected Customer Data bucket that enterprises negotiate and audit.
That boundary has been redrawn with remarkable consistency. Leading platforms have expanded “Usage Data” to include reasoning traces: which suggestions are accepted or rejected, what gets edited, how decisions are made, and the execution paths that follow. In the most striking case, the definition of Usage Data remained word-for-word identical across a 14-month window while the surrounding terms added an outright ownership claim — “all right, title, and interest” — and flipped model training from opt-in to opt-out. The words held still. The meaning inverted.
Because these agreements carve Usage Data out of the Customer Data category enterprises traditionally negotiate, the standard safeguards — deletion rights, training opt-outs, zero-retention commitments — do not reach it. Where the terms assert ownership, no opt-out for that category exists at any tier or setting, and structurally none can: you cannot opt out of someone else’s property.
De-identification and aggregation — long presented to customers as privacy protections — function as the transfer mechanism. One provider’s privacy policy exempts content already de-identified from its own 30-day deletion commitment. Another retains embeddings and codebase metadata after the underlying plaintext is deleted. A third states that once data is de-identified, it may be used and shared for any purpose — and that its privacy policy no longer applies to it. In each case, the customer’s identity is removed from the record while the reasoning pattern stays with the vendor.
Retention is not theoretical. In active litigation this year, a federal court ordered one provider to preserve and produce approximately 20 million de-identified consumer conversation logs — demonstrating that retained data remains reachable by discovery and acquisition regardless of any training promise.
One precision up front: what we can prove is custody and title — what these standard terms hand the vendor. Whether any given vendor is training on your traces today is, by design, unprovable from the outside. Closed-weight systems mean no customer can inspect how the model was trained. Training is the logical reading of the public record, current efforts in computer science, and the real limits of the “plain text internet.” But the claim that needs no inference is ownership. The contracts put that beyond dispute. Negotiate from an informed position.
Enterprises should negotiate Usage Data definitions with the same rigor they already apply to Customer Data. They should also evaluate architectures that keep telemetry, memory, and reasoning traces under their own control. The Grove Foundation published the Autonomaton Pattern (GRV-001) under Creative Commons BY 4.0 for exactly this purpose: an open architectural standard that makes operator ownership of institutional reasoning a structural property of an agentic system rather than a contractual hope.
Read the Autonomaton Pattern (GRV-001) →All five Approaching-Critical patterns this quarter are Sovereign open-weight — DeepSeek V4, Mistral, Gemma 4, GLM-5.2, and Qwen 3.6 — the models an enterprise can download, run, and keep. The strongest closed API pattern sits a full tier below. The market is repricing structural viability toward the models a company can hold — the same custody logic this alert traces from the other direction.
Read the full Λ Standings →The full analysis — including the cross-vendor matrix, 14-month custody ratchet, provider-level exhibits, and every primary-source citation with verification hashes — appears below. Every load-bearing claim is fact-checked and linked.
Industry-wide standardization of trajectory harvesting. Even in “protected” environments, vendors systematically capture behavioral reasoning traces.
Enterprises pay for intelligence while simultaneously feeding vendors the proprietary reasoning traces that make the models more powerful.
A "real trust boundary" so organizations retain ownership of their traces, evals, memory, and institutional context.
Control over "compute, models, data stack, and alpha" — owning the means of production so learning isn't transferred to the vendor.
Autonomous coding harnesses no longer function merely as productivity accelerators. They are structurally integrated arrays that capture the operational graphs and cognitive trajectories of enterprise knowledge workers — the exact reasoning paths needed to train next-generation PRMs.
Legacy cloud terms treated "telemetry" as server uptime, crash reports, and infrastructure metrics; the interaction payload — code, plans, files — was protected user content. Agentic vendors surgically decoupled the two.
A standing exemption removes any content already "de-identified and disassociated from your account" from the scope of deletion requests. Once a trajectory is stripped of regulated identifiers and reduced to an abstract interaction graph, it is relabeled as de-identified telemetry — and the enterprise's right to delete ceases.
"De-identified" is a database label — not a legal firewall. Once applied, the reasoning trace becomes vendor analytical property, outside deletion rights and ZDR scope.
A federal court ordered the preservation and production of approximately twenty million consumer ChatGPT logs — explicitly including conversations users believed they had permanently deleted.
Proven in open court: de-identified data retained on a vendor server is structurally preserved and legally reachable by subpoena — regardless of any UI deletion assurance.
Five documented beats in fourteen months. Each one transferred more custody of how your company thinks from the enterprise to the vendor.
Every step is timestamped on the vendor’s own site — the April 2025 terms are still linked from the current page. Legal teams do not move this fast, in this direction, by accident.
A file you own is table stakes. A system that cannot exceed it — and a record that grows more valuable every time you use it — that’s the standard.
The same signal the vendors bank as Usage Data, running inside operator-owned software. The model-independent agentic system (autonomaton harness) clustered ten memory records around an emergent theme, checked them against its tracked goals, found none, and proposed a precise amendment to its own goal file — in this case, evidence-linked and staged for human review and approval by design.
The vendor keeps this as opaque telemetry. The autonomaton hands the operator the whole chain — evidence, reasoning, and a veto.
A rounding error two years ago; roughly half of some weeks’ volume now. Enterprises are voting with volume for a sovereign, multi-model stack — where the trace never leaves the building.
Independently confirmed: open Chinese models run 60–90% cheaper than leading US models, and GLM-5.2 landed within one point of Claude Opus 4.8 on a watched agentic benchmark — at roughly a fifth of the cost. The frontier premium buys months, not moats.
By classifying the outcome-verified reasoning trace — every accepted completion, every manually edited variable, every terminal-error correction — as "de-identified telemetry" or "usage metrics," vendors legally sidestep ZDR pledges. ZDR stays valid for raw text, which is now a depreciating asset. The premium resource for PRMs is the interactive logic path itself.
Enterprises are paying premium subscriptions to fund the automation of their own proprietary judgment.
Raw data (bits) is unipolar — it moves frictionlessly. Meaning, context and intent are multipolar — they depend on localized context and human agency. Harvesting interaction telemetry is the systematic capture of that agency.
A single-quarter collapse from monopolistic to competitive — a volume-driven migration toward sovereign, multi-model, open-weight architectures.
Read the fine print. Then change your architecture.