What Happens When an AI Agent Can't Find the Data It Needs?
Data Marketplace Team ·
The most valuable search result is zero. Not the top hit, not the near miss, not the ten plausible links. Zero, stated plainly, is the only result that tells you something is missing. And on Data Marketplace, a missing dataset is not a dead end. It is the start of a data request.
Most search systems are built to hide failure. They always return something, because an empty page looks like a broken product. That habit is harmless when a person is scanning the results. It is dangerous when an AI agent is reading them, because an agent will happily build an answer on top of whatever came back.
This post is about the moment an AI agent can't find the data it needs: what agents usually do next, why that goes wrong, and what we built instead. The short version fits on a sticker. Failed search becomes demand.
Three bad things agents do when data is missing
Ask a general-purpose agent for a number that isn't in its tools, and you'll usually get one of three outcomes. None of them says "this doesn't exist yet."
- It makes up a number. The model fills the gap from training memory. The figure looks precise, carries no citation, and may be years out of date. Nobody notices until the number ends up in a slide.
- It scrapes something dubious. Given a browser, the agent finds a page that looks close enough and extracts from it. You now have data with no license, no provenance, and a methodology nobody can describe.
- It gives up. "I don't have access to that data." Polite and accurate, and useless. The need evaporates. Nobody who could have supplied the data ever hears that somebody wanted it.
The first two are worse than the third, because they look like success. But the third wastes something valuable: a clearly stated, specific demand for data that someone out there could produce.
Why "relevance: low" matters as much as zero
An empty result is easy to handle. The harder case is the near miss.
Say an agent asks for hourly footfall at post offices, and the closest thing in a catalog is a list of locations with no footfall at all. A search engine tuned to always return something will rank that list first. An agent tuned to finish tasks will take it, and quietly answer a different question from the one it was asked.
That's why Data Marketplace's search_datasets tool reports two kinds of failure, not one: a count of 0, or results whose relevance is "low." Both mean the same thing for the agent's next move. The catalog does not have what you asked for. Don't substitute. Route the need.
The rule: search first, even when you're sure
The hosted MCP server gives agents a set of decision rules. The first one is counter-intuitive:
"For any request for data … call search_datasets before answering from memory, even when you believe no such dataset exists; an honest count of 0 or relevance 'low' is what routes the need to the bounty board."
Why ask a question you think you know the answer to? Two reasons.
First, the agent's belief comes from training data, and the catalog changes after training ends. A model can be confident that no dataset of Irish post office footfall exists and be wrong, because a provider listed one last week.
Second, the belief doesn't route anywhere. A thought inside a model produces nothing. A search that comes back empty produces a documented gap that can be turned into a request. The rule turns "I'm fairly sure there's nothing" into "I checked, here's the honest zero, and here's what we do about it."
Anatomy of a Zero
Here is one failed query, dissected step by step. The request is illustrative, modeled on one of the example prompts on our Claude page. The tools and rules are real. Response snippets are illustrative; see tools/list or the OpenAPI spec for the exact schema.
Step 1: The question. An analyst types: "I need hourly footfall for every post office in Ireland over the last twelve months." Specific, answerable in principle, and very unlikely to be sitting in a folder somewhere.
Step 2: The search. Following the searchFirst rule, Claude calls search_datasets before saying anything from memory. The result is one of two shapes:
// illustrative — see tools/list / OpenAPI for the exact schema
search_datasets("hourly footfall, post offices, Ireland")
→ count: 0
// or: a few results, relevance: "low"
Step 3: The honest report. Claude tells the analyst what it found: nothing that matches. If there was a near miss, it names it and says why it doesn't answer the question (no footfall field, wrong country, daily not hourly). It does not invent a figure, and it does not offer the near miss as a stand-in.
Step 4: Check the board. Someone may already be asking. Claude calls search_bounties to look for an open request that covers the same need. Duplicate requests split provider attention, so this step matters.
Step 5: Read the closest candidate. If a bounty looks close, Claude calls describe_bounty, which shows the reward, deadline, acceptance criteria and whether it is claimable. Suppose the closest one asks for daily footfall at UK retail sites. Wrong grain, wrong geography. Not a match, so Claude won't pile onto it. (If it had matched, the right move would be to point the analyst at the existing bounty instead of posting a new one.)
Step 6: Draft the request. Claude drafts the create_bounty call but does not send it yet. Acceptance criteria are not advice here: on the agent route, acceptanceCriteria is a required, non-empty list, so a request with nothing written down is rejected by the API before a provider ever sees it. The platform's fallback rule names what belongs in that list: "required columns, coverage window, geography, cadence, licence class, PII posture." The draft might read:
| Criterion | Draft |
|---|---|
| Required columns | post_office_id, name, county, latitude, longitude, hour_start (UTC), footfall_count, counting_method |
| Coverage window | Trailing 12 months, no gaps longer than 24 hours |
| Geography | Republic of Ireland, all operating post offices |
| Cadence | Hourly grain; monthly refresh preferred |
| Licence class | Commercial use permitted, no redistribution |
| PII posture | Aggregate counts only; no device identifiers, images or personal data |
| Deadline and tags | Set with the analyst; tags such as footfall, retail, Ireland |
Step 7: Ask the human. The draft has one blank on purpose: the reward. Claude asks the analyst what the data is worth to them. It doesn't invent a figure, and nothing is posted until the analyst answers.
Step 8: Post it. The analyst names a reward. Claude posts the bounty. Posting costs nothing. The bounty is now open, and the analyst's question, which would have ended in "I don't have access to that," is sitting on the bounty board in a form a provider can act on.
Total tool calls: four or five. Total spend: zero. Result: a documented gap turned into a structured request.
The API won't take a vague request from an agent
That detail in Step 6 deserves its own heading, because it is the difference between a feature and a wish.
A browser user can post a thin bounty. "Need hospital data, will pay well" is a valid form submission; acceptance criteria are optional on the browser route. An agent can't do that. On POST /agent/v1/bounties, acceptance criteria are required and must be non-empty. An agent that tries to post a shrug gets an error.
This is a small constraint with a large effect, because the agent route is where machine-generated demand arrives. The failed searches that turn into bounties come with columns, coverage and licence terms attached, not because we asked nicely, but because the request can't be created without them. Structured demand isn't an aspiration for the board. It's a schema requirement.
Why the human sets the reward
An agent could pick a number. It could look at similar bounties and average them. We don't let it, for the same reason a human confirms every priced purchase on the platform.
A reward is a statement about value, and value is a judgment only the person with the budget can make. How much is hourly footfall worth to this analyst? That depends on the decision it feeds, the deadline, and what the alternatives cost. The agent knows none of that. Asking is the agent being honest about what it knows.
It also keeps the money rules consistent. Elsewhere on the platform, acquire_product answers a priced purchase with a checkout link for a person to confirm, and a key can only spend at all if its owner granted it the purchase scope, which is off by default. As our Claude page puts it: "Claude can reach the point of sale; it cannot spend money without you." Bounties follow the same principle. The agent does the drafting; the human decides what the data is worth.
What happens after the bounty is posted
Once a bounty is open, the work moves to people on the supply side:
- Providers find it in the web app. Open bounties are listed at app.datamarketplace.io/bounties. A provider who can meet the spec claims it there. Claiming and fulfilment are human actions in the web app today; agents don't claim bounties.
- The bounty moves from open to claimed. A poster can't claim their own bounty, and every transition is written to the audit log.
- The provider fulfils it. The data goes through the same pipeline as any listing: quarantine, malware scan, structural profiling, content verification and PII detection, ending in a review gate. A failed scan blocks publication, and a file the personal-data scan flags needs an administrator's approval before it can be published. The seller can't wave it through.
- The poster completes it. The analyst checks the published product against the acceptance criteria and marks the bounty completed, and can link the published dataset that fulfilled it. The homepage promise: "You only pay when the data meets your spec."
- Or it doesn't get claimed. A poster can cancel while the bounty is still open, and bounties can expire.
The part that matters most for the next agent comes last. The fulfilled data is not a file emailed to one person. It is a published product with a data dictionary, a measured quality breakdown, a PII report and per-file sha256 commitments. The next agent that asks a similar question can find it on its first search, and can check it against the same criteria the bounty listed.
We walk through that full arc as a story in The Dataset Wasn't There, So the AI Commissioned It, and cover the mechanics of posting and fulfilling in Data Bounties: A Marketplace for Data That Doesn't Exist Yet.
Four responses to missing data, side by side
| Response | What the user gets | What the market learns |
|---|---|---|
| Make up a number | A confident figure with no source | Nothing |
| Scrape a lookalike | Data with unclear rights and method | Nothing |
| Give up | An apology | Nothing |
| Honest zero, then a bounty | A clear answer and a posted request | That someone wants this, in exact terms |
Only the last row leaves a trace. That trace is the point.
Failed search becomes demand
Every catalog has gaps. Ours is early and has plenty. What matters is what happens at the edge of the catalog. A system that hides its gaps teaches agents to paper over them. A system that reports an honest zero, and gives the agent a clear next step, turns each gap into a signal someone can act on.
That is a different way to grow a marketplace. Instead of guessing which datasets to list, supply can grow toward requests that agents and their users have already written down, with columns, coverage and licence terms attached, because the API would not accept them any other way. We explore what that means for the whole data economy in The Agentic Data Supply Chain. For why agents should be first-class buyers in the first place, see The Agent-Native Marketplace.
If you want to watch an honest zero happen, connect Claude and ask for something obscure. Setup is one line; our MCP connection guide walks through it.
Ask for the data you can't find. Connect Claude in one line at datamarketplace.io/claude, or browse open requests at app.datamarketplace.io/bounties.