← All posts

What Makes a Dataset Machine-Ready? The 12-Question Test

Data Marketplace Team ·

A CSV is a rumor. A machine-ready dataset is testimony.

The rumor says "here are some numbers." It doesn't say what the columns mean, where the rows came from, whether anyone checked them, or whether you're allowed to use them. A human analyst can chase those answers down: email the author, open the PDF, squint at the header. An AI agent can't. It gets one pass at the file and whatever the file says about itself.

So here is the counter-intuitive part. Machine-ready does not mean clean. A perfectly clean table with no context is still a rumor. A slightly messy table that declares its schema, its units, its version, its license and its hash is testimony: every claim it makes can be checked.

This post defines the term and gives you a tool to apply it: the Machine-Ready Test, 12 questions an agent asks a dataset before trusting it.

Why the bar moved

For twenty years, dataset documentation was written for a colleague who could ask. "What's val2?" "Oh, that's the seasonally adjusted one."

That reader is changing. Agents now search for data, compare candidates, sample rows and decide what to use, often mid-task with no person watching. Every gap in the documentation becomes a guess, and guesses compound.

The rule that follows is simple, and it's the most useful sentence in this post:

Write your README for a reader who can't ask follow-up questions.

If a fact about your data lives only in your head, in a Slack thread, or in a PDF linked from a web page nobody's agent will open, it doesn't exist for a machine.

The Machine-Ready Test

Twelve questions, grouped into four concerns. Score each one 0, 1 or 2. The maximum is 24.

Shape: can a machine parse it correctly?

  1. Is there a declared schema?
  2. Are the types explicit?
  3. Is every column described?

Meaning: can a machine interpret it correctly? 4. Are units and grain stated? 5. Are there stable identifiers and join keys? 6. Are delivery formats and file metadata defined?

Trust: can a machine check what it's told? 7. Can I tell where it came from and which version this is? 8. Are there per-file hashes, published before I pay? 9. Is quality measured, not described?

Permission and access: may a machine use it, and can it look first? 10. Is the license machine-readable? 11. Has personal data been scanned, and is the inspected scope recorded? 12. Can I sample or probe it before committing?

The scoring rubric

# Question 0 points 1 point 2 points
1 Declared schema? None; header row only Schema inferred by a tool, not declared Declared, machine-readable schema that matches the files
2 Explicit types? Everything is a string Types described in prose Typed columns, with nullability and formats (e.g. ISO 8601 dates)
3 Every column described? No data dictionary Some columns described Full data dictionary, one entry per column
4 Units and grain? Neither stated One of the two Both: "one row = one country-month; value in percent"
5 Stable identifiers / join keys? Free-text names ("Germany", "DE", "Deutschland") IDs present but home-grown Standard keys (ISO country codes, dates) declared as join keys
6 Formats and metadata? Whatever the seller exported One documented format Typed formats (CSV, Parquet, JSON) with encoding, delimiter and compression stated
7 Source and version? Unknown Source named, no version Source plus a pinned version or snapshot date
8 Per-file hashes before purchase? None Hash available only after delivery Per-file sha256 published before purchase
9 Quality measured? Seller adjectives ("high quality!") One-off statistics, or a score with no components A composite scored from measured components (completeness, schema validity, dictionary coverage, freshness), with the unmeasurable parts declared rather than assumed
10 Machine-readable license? None, or "contact us" Named license in prose or a PDF A licence identifier on the listing plus machine-readable terms
11 PII scan with inspected scope? No scan "Scanned," scope unknown Scan report that records what was actually inspected
12 Sample or probe first? Buy to see A static sample file Seller-permitted sample rows via API, plus a bounded query or aggregate path

What your score means

  • 0–8: Rumor. Someone says there's data here. A machine can't confirm anything about it.
  • 9–16: Hearsay. Some claims are written down, but the important ones (license, version, integrity) can't be checked.
  • 17–21: Affidavit. Documented and mostly verifiable. Usually missing one trust or permission signal.
  • 22–24: Testimony. The dataset answers every question, and each answer can be checked.

Illustrative example: a data.csv on a shared drive, with a header row, one line of README, and a note saying "public data, use freely." It scores roughly 3 out of 24: a named source, one documented format, a partial column list. Everything a machine needs to trust it is missing. It was written for a colleague.

Walking the questions that matter most

A few questions deserve more than a table row.

Units and grain (Q4) cause the quiet failures. Average percentages with basis points, or treat monthly rows as annual, and you get a confident, wrong number. Neither error throws an exception. State both, per column.

Per-file hashes before purchase (Q8) are the difference between "trust me" and "check me." A hash published after delivery shows the file wasn't corrupted in transit. A sha256 published before purchase shows the file you receive is the exact file you evaluated. We cover this in depth in Why AI Agents Need Data Provenance and put it to work in How to Verify a Dataset Before You Buy It.

Measured quality (Q9) is where most listings bluff, and where the honest answer includes an admission. A score with no components is a mood. A score with components tells you which property was weak, and a careful score also tells you how much of itself it could actually measure. A dataset with no declared schema can't be scored for schema validity; the right behavior is to say so, not to quietly average over the gap.

Inspected scope (Q11) is where most "PII-free" claims fall apart. "We scanned it" means little unless you know what was scanned: every file, every column, a sample? A scan that recorded nothing about its scope is not the same as a scan that inspected everything and found nothing. A clean scan differs from an absent one.

Sample or probe (Q12) lets an agent test a hypothesis before spending money. Sample rows show the texture of the data; a bounded query or an aggregate checks its contents without a bulk export. For sensitive data, the aggregate path is often the only path: see How Privacy-Preserving Compute Lets AI Use Data It Shouldn't Download.

An example machine-ready manifest

Here's what 24 out of 24 can look like, written as a single manifest. This is illustrative: a teaching format, not the response shape of any API. It borrows the title of a real Eurostat HICP listing, but the column names, file names, coverage dates and field values are invented for the example. Where a field exists on Data Marketplace, we use its real name (version.files[].sha256, inspectedScope, measuredWeight); the rest are descriptive labels. Hashes are shortened placeholders.

{
  "title": "Eurostat HICP — Monthly Annual Inflation Rate, 32 EU and EFTA Economies, 2015–2025",
  "source": "Eurostat",
  "licence": { "identifier": "CC-BY-4.0", "attribution_required": true },
  "grain": "one row = one economy × one month",
  "coverage": { "geography": "32 EU and EFTA economies", "window": "2015-01 to 2025-12" },
  "update_frequency": "monthly",
  "join_keys": ["geo", "time_period"],
  "data_dictionary": [
    { "column": "geo", "type": "string", "nullable": false,
      "description": "Economy code (ISO 3166-1 alpha-2 style)" },
    { "column": "time_period", "type": "string", "format": "YYYY-MM", "nullable": false,
      "description": "Reference month" },
    { "column": "annual_rate_pct", "type": "decimal", "unit": "percent", "nullable": true,
      "description": "HICP annual rate of change vs same month previous year; null where not published" }
  ],
  "quality": {
    "components": { "completeness": 0.97, "schemaValidity": 1.00,
                    "dictionaryCoverage": 1.00, "freshness": 0.88 },
    "measuredWeight": 1.00
  },
  "personal_data": { "scan": "automated PII detection", "inspectedScope": "all files, all columns" },
  "sampling": { "sample_rows": "seller-permitted", "query": "bounded, ≤100 rows", "compute": "aggregates only" },
  "version": {
    "files": [
      { "name": "hicp_monthly.parquet", "format": "parquet", "sha256": "9f2c…e41a" },
      { "name": "hicp_monthly.csv", "format": "csv", "encoding": "utf-8", "sha256": "4b7d…0c93" }
    ]
  }
}

Read it the way an agent would. It can join this to another country-month table without guessing, it knows attribution is required, it knows a null means "not published," it can see which quality components were actually measured, and it can check after delivery that the bytes match the hashes it saw before paying.

How Data Marketplace attaches each element

We built the listing flow so that a published product scores high on this test by construction. Here is the mapping, question by question.

  • Schema, types, dictionary (Q1–3): Every dataset ships with a machine-readable schema and a data dictionary. describe_product returns the per-resource data dictionary. Sellers set it with set_documentation, which requires a non-empty README, and an empty README is one of the gates that blocks publishing.
  • Units, grain, join keys (Q4–5): These live in the dictionary and README, and the seller writes them. We can require a README; we can't write your units for you. search_datasets ranks by category, geography, licence and join keys, so declared keys also make a product easier to find.
  • Formats (Q6): Delivery is CSV, Parquet or JSON.
  • Source and version (Q7): A purchasable product pins a current version. Publishing is gated: publish_product refuses with listing_policy_unmet and a list of { gate, reason } rejections until the listing has its metadata (title, summary, readme, categories), a non-empty readme, a price, and either temporal coverage or an update frequency. The seller sets coverage, update frequency, the price and the licence identifier with set_pricing.
  • Per-file hashes (Q8): The pinned version's per-file sha256 commitments (version.files[].sha256) are published before purchase. The hash isn't added at the end, either: finalize_upload checks the uploaded file's size and sha256 against the upload ticket before any scan starts. At delivery, get_delivery returns the commitment again as commitmentHash beside download links that live for 300 seconds.
  • Measured quality (Q9): The quality breakdown is a composite of four measured components: completeness at 0.35, schema validity at 0.25, dictionary coverage at 0.20 and freshness at 0.20. When a component can't be measured the remaining weights are renormalized, and measuredWeight reports how much of the score was measurable. If neither completeness nor schema validity can be measured, the score is null, which is the honest answer and not a zero. Freshness gets a 30-day grace window and then decays linearly to zero at 730 days, so a slow-moving series isn't penalized for being a month old while a stale "daily" feed is. describe_product returns the breakdown so agents can rank on measured numbers rather than the seller's description, and minQualityScore lets a search filter weak listings out. For the broader framing of what makes a dataset useful, our earlier post Introducing Predictive Utility Scores sets out the wider model; the four components above are what the shipped breakdown measures today.
  • License (Q10): The listing carries a licence identifier, such as CC-BY-4.0, set with set_pricing, and the product spec carries machine-readable terms, so the license check can happen before purchase. The five policy archetypes (OA-CCBY, NEL, EXL, DRV, FED) are the ODRL layer that evaluates those terms, not a value the seller types in. Background: Why Machine-Readable Licenses Will Replace PDF Contracts.
  • PII and inspected scope (Q11): Every upload runs through automated personal-data detection. Sellers watch it with get_scan_status, which reports per file: scanStatus, piiStatus, piiFlagged, piiReviewApproved, settled and a scan report. You poll until every file is settled, not until it reaches a particular status, because a failed file never reaches one. A flagged file needs an administrator's approval before it can be published; the seller cannot wave it through. On the buyer's side, sampling availability carries inspectedScope, which records what the scan actually inspected and is null when that wasn't recorded. We made the null visible on purpose.
  • Sample and probe (Q12): Sellers choose preview rows with set_sampling_policy. Agents read them free with describe_dataset. On a licensed dataset, query_dataset runs a whitelisted equality-filter query capped at 100 rows ("a probe, not a bulk export"), and compute_dataset returns a privacy-preserving aggregate.

One more design choice: when a seller runs publish_product with gaps, it returns every unmet listing gate at once, not one failure at a time. If you're preparing data to list, How to Sell a Dataset walks through the full flow.

Candidly, the platform guarantees the presence of most of these signals, not their quality. A dictionary can exist and still be vague. That's why the test is scored 0–2, not yes/no.

FAQ

Is "AI-ready data" the same as machine-ready data? They overlap. "AI-ready" usually means cleaned and formatted for training or analysis. Machine-ready is broader: the dataset must also explain itself (schema, dictionary, units), prove itself (version, hashes, measured quality) and declare its permissions (license, PII posture) in a form software can read.

Does machine-ready mean the data is clean? No. A dataset can be clean and still be a rumor if nothing about it can be checked. Machine-ready means the data's claims about itself are explicit and verifiable, including honest statements of where it's incomplete.

What does a measured quality score actually measure? On Data Marketplace, four components: completeness (weighted 0.35), schema validity (0.25), dictionary coverage (0.20) and freshness (0.20). Weights are renormalized when something can't be measured, measuredWeight says how much of the score was measurable, and the score is null when neither completeness nor schema validity could be measured at all.

Does a clean PII scan mean a dataset contains no personal data? Only within the scope the scan inspected. On Data Marketplace, inspectedScope records what was inspected and is null when that wasn't recorded. A clean scan with a recorded scope is evidence; a missing scope is an open question.

Can I run the Machine-Ready Test on a listing before buying? Yes. describe_product returns the metadata, per-resource data dictionary, licence, quality breakdown, sampling availability and version file hashes, and describe_dataset returns seller-permitted sample rows for free. In the web app you can sample 50 rows before you buy.


Score a listing yourself: browse the catalog at app.datamarketplace.io/discover, or connect Claude in one line at datamarketplace.io/claude.

What Makes a Dataset Machine-Ready? The 12-Question Test — Data Marketplace