[COMPANY] Book a data audit

Data supply for the AI economy

Turn data you already have into income you don’t have to work for.

We find it inside your business, strip every trace of personal data to a standard that puts it out of scope of the GDPR, package it as a priced dataset, and place it with the buyers who need it. You carry no legal load and no ongoing work. After onboarding, the revenue is passive.

01

What we do

We are specialists in esports and in enterprise operational workflows — and in turning both into something an AI lab can legally train on.

Esports

Competitive play is expert decision-making under time pressure, recorded at a density nothing else in the world produces: every input, every read, every correction, hundreds of times an hour. Almost none of it is retained in a form anyone can train on. We capture it at the source and package it.

Enterprise workflows

The work inside a claims desk, an underwriting queue or a compliance review is expert judgement applied hundreds of times a week and then thrown away. We record the shape of it — which controls, in what order, against which case, to what outcome — as a reusable trace of how the decision was actually reached.

Software we built ourselves

Both run on a capture stack we wrote rather than bought. Personal data is detected and replaced with anonymous tokens on the machine it was captured on, before anything is transmitted — and the same entity keeps the same token across sessions, which is what lets a buyer follow one case end to end without ever learning whose it was.

Workflow traces are the one category the labs cannot buy at any price today. They are paying contractors by the hour to reconstruct, badly, work that your organisation already does correctly and then files away.

02

How it works

Three steps. Two of them are ours.

  1. 1

    We get on a call and map how you operate

    One conversation, about an hour. We walk through the processes you actually run, what systems they touch, and what those processes throw off as a by-product. Most of the value turns out to be sitting in an archive nobody has opened in years.

    What we need from you: one call with someone who knows the operation.

  2. 2

    We propose a strategy, then build it

    We come back with a written assessment: which assets are sellable, what they are plausibly worth, how they get cleaned and anonymised, and who the buyers are. If you sign it off, we implement the whole thing — collection, de-identification, packaging, the legal work and the buyer relationships.

    What we need from you: one sign-off, and one integration window.

  3. 3

    You get paid, and you do nothing

    From then on it runs without you. We handle collection, anonymisation, compliance documentation, dataset packaging, buyer negotiation and contracts. Revenue arrives on a schedule. There is no team to hire, no process to change, and nothing added to anyone's job description.

    What we need from you: nothing.

03

What actually qualifies

Not “your data” in the abstract. These are the specific assets that a buyer will write a cheque for, and the reason each one commands a price.

Email and message threads

Multi-party correspondence is how people actually negotiate, escalate and decide in writing. A public bid valued corporate email and chat at about a cent a message in a forced sale. A clean, consented, domain-specific archive is worth a multiple of that.

Support tickets and resolutions

A ticket is a problem, an attempt, a correction and a confirmed fix. That closed loop is what makes it trainable rather than merely readable.

Complete case and claim files

Intake to settlement, including the exceptions and the reversals. One insurer's claims history is a corpus no lab can reconstruct from the open web at any price.

Legal matters, filing to close

Reasoning chains with a verified outcome attached. Legal-tech builders need work that was actually completed and actually held up, not synthetic argument.

End-to-end workflow traces

Which control, in what order, with what value, and what the outcome proved. This is what agentic AI cannot buy anywhere else, and labs currently pay expert contractors by the hour to recreate it from scratch.

Call and meeting transcripts

Real spoken negotiation, objection and repair, unscripted and domain-specific. Public speech corpora are read aloud and generic by comparison.

Internal documentation and SOPs

The written rules your people follow. Procedure documents teach a model the policy that your traces then show being applied.

Decision and exception logs

Outcome labels are the scarcest thing in training data. A log of what was approved, declined or overridden, and what happened next, is a labelled dataset you already own and have never sold.

Structured back-office records

Invoices, policies, shipments, applications and their lifecycles. Individually mundane, collectively a map of how a regulated industry actually runs.

If you hold something not on this list, it is still worth the call. The categories that pay best today did not exist as categories eighteen months ago.

04

What is it worth

An indicative range, built from public licensing transactions. Nothing here is an offer, and nothing here is a number we invented.

What are you valuing?

How big is the operation?

A mid-size company 300-800 people (~815K records on file)

A small team A small business A growing company A mid-size company A large company A large enterprise
Where these numbers come from

There is no published price list for AI training data. Every rate in this model is derived from a reported transaction or a published labour rate, listed below. Where no public benchmark exists — workflow traces are the clearest case, because labs commission them rather than license them — we say so instead of inventing a citation. The spread between the low and high figure is wide on purpose. Anything narrower would be false precision.

  • Spirit Airlines data estate, competitive bids August 2026

    Google bid $10M, Mercor $7.5M and micro1 $12.5M for a bundle including 100M corporate emails and 500M Teams messages. That normalises to roughly $0.005-$0.021 per message. It was a distressed bankruptcy sale, so we treat it as a floor rather than a fair-market price.

    Source
  • Reddit content licensing 2024, revisited 2026

    $60M per year from Google, with $203M in total contracted licensing disclosed at IPO. The reference point for bulk conversational text with a resolution attached.

    Source
  • News Corp and OpenAI May 2024

    $250M over five years. The largest disclosed content licence, and the proof that archives trade at scale rather than as curiosities.

    Source
  • HarperCollins and Microsoft November 2024

    $5,000 per title, shared evenly with the author, which works out near $10-$17 per page of edited prose. The practical ceiling for long-form written material.

    Source
  • Tempus, de-identified clinical records Reported 2024-2025

    $200M over three years for anonymised patient and genomic data. The closest public comparable for a regulated, per-case archive sold lawfully.

    Source
  • Agentic trajectories are commissioned, not licensed 2026

    There is no liquid market for pre-existing task traces. What exists is the cost of creating them: expert contractors bill $85-$250 an hour, and a typical multi-step trajectory costs a lab in the low hundreds of dollars. The one public catalogue of pre-existing computer-use traces prices privately on request. Our rate is derived as a fraction of creation cost. It is not an observed price, and we say so.

    Source
  • Photobucket, $0.05 to $1.00 per image April 2024

    One seller, one catalogue, a twenty-fold spread depending on the buyer and the quality of the material. The clearest public evidence that condition, not volume, sets the price.

    Source
  • Session and spoken material, per hour 2025-2026

    $1-$4 per minute for ordinary footage, rising to as much as $1,000 an hour for first-person and task-demonstration material. Used here to bound hour-denominated assets.

    Source
  • Vertical legal and claims text is unpriced in public 2023-2025

    No per-document benchmark exists. Only whole-company comparables do: Clio acquired vLex for $1B and Thomson Reuters acquired Casetext for $650M, then chose to build on Westlaw rather than license it out. Supply is being withheld, which is precisely why it is scarce.

    Source
  • There is no price list 2026

    Every figure in this model derives from a reported transaction or a published labour rate. Nobody publishes a rate card for AI training data. Any site that shows you one confident number is guessing with more confidence than the evidence supports.

    Source
05

First call to first payment

Roughly ninety days, of which about two hours are yours.

  1. Day 0

    First call

    One hour. We map the operation and the data it produces.

  2. Day 1–10

    Written assessment

    What is sellable, what it is worth, how it gets cleaned, who buys it.

  3. Day 10–30

    Legal groundwork

    DPIA, lawful basis, contracts. Nothing is collected before this is done.

  4. Day 30–55

    Pilot extract

    A first packaged dataset from a narrow slice, so both sides can see the real thing.

  5. Day 55–80

    Buyer placement

    We take it to the market, negotiate terms and close.

  6. Day 90

    First payment

    And from here it is passive.

06

The questions you are about to ask

Does our data leave our systems?

The raw data does not. De-identification happens at the point of capture, inside your environment, before anything is transmitted. What we receive has already had the personal data removed. We could not hand over your customers' details to a buyer even if a buyer asked, because we never hold them.

What about our clients' and employees' consent?

The output is anonymised to the point where it is no longer personal data, which puts it outside the scope of the GDPR entirely. That is the standard we work to, and the DPIA documenting it exists before collection starts. Where a consent basis is the cleaner route for a given asset, we build the consent flow rather than assume it.

Could a buyer work out that the data is ours?

Not from the data. Datasets are packaged so that the source organisation is not identifiable from the content, and attribution is a commercial decision you make, not a default. Some clients want to be named because it is good positioning. Most do not.

What does this cost us?

Nothing up front. There is no fee, no licence and no platform charge. We are paid out of what the data earns, on terms agreed with you before anything is collected, so we only do well if you do.

How long until we see revenue?

About ninety days from the first call in a typical case, most of which is legal groundwork and buyer negotiation rather than engineering. A narrow, clean archive can move faster. A complex multi-entity one takes longer.

What if we want to stop?

Collection stops when you say so, and the contract says so in writing. Datasets already licensed remain licensed under their existing terms, which is a constraint of the market rather than of our agreement, and we are explicit about it up front rather than in a footnote.

Who actually buys this?

Frontier labs, and the vertical AI builders behind them — legal-tech, insur-tech, healthcare automation — training agents that have to complete regulated work end to end. They are currently paying expert contractors by the hour to reconstruct, badly, work that your business has already done and filed.

Is our industry too small or too boring for this?

Boring is the point. The material that is hardest to obtain is exactly the material nobody thought to publish. Scarcity, not glamour, sets the price.

Find out what you are sitting on.

One call, about an hour, with someone who knows how the operation runs. You leave it with a written view of which of your assets are sellable and what they are plausibly worth. There is no cost and no obligation attached to that.