Repository navigation
Preventing tool poisoning: save signatures of possible tool calls #82
Description
Activity
- addednot go-live blockerThis issue has been reviewed and determined to not be a blocker to go-liveThis issue has been reviewed and determined to not be a blocker to go-live
on May 27, 2025 I had a request on this for VS Code. I didn't dive deep into it but there are a couple scenarios that complicate this.
- Servers that provide a dynamic tools
- Tools that localized (Adds i18n support modelcontextprotocol#115)
- Tools that may change depending on the model (Feature Request: Model-Aware Content Adaptation for MCP modelcontextprotocol#469)
Just letting clients do this may make more sense because the client will know at least (2) and (3) or know when they change. But the ecosystem tooling here is still quite immature.
Reacted by yaoIf we focus on the notion of possible tools (i.e. any tool that could ever surface, beyond just at-initialization), those three problems go away, right?
Yes. It's a bit of a pain because, given both (2) and (3) happen it's a
NxMmultiplier to the number of tools, but it's possible.Reacted by Tadas Antanavicius(though practically only a relatively small number of servers will get localized into a large number of languages)
I think tool representation overall, either in the registry spec or MCP metadata, would be useful. Currently clients need to start servers in order to know what tools, prompts, etc. are available. Because MCP servers can now do nice authentication, this can lead to a bunch of login prompts for the user when they first start interacting with their language model. If servers optionally (there might always be servers with nondeterministic tools) published their starting set of tools/prompts, then we could avoid this.
I think tool representation overall, either in the registry spec or MCP metadata, would be useful. Currently clients need to start servers in order to know what tools, prompts, etc. are available.
Agreed, in particular for tool representation in MCP metadata. Unrelated to tool poisoning, but it might allow clients to statically index tools for client-side tool search.
Hi everyone, I'm glad that @tadasant raised this.
We are facing a challenge that would also be solved by having easy access to the tools provided by each MCP server. In our case, we want to avoid forcing a user to authenticate to a service before we can even confirm it has the tools to solve their problem (via
tools/list).Could the
serverschema within the registry API be augmented to includetools? This would allow a server operator to choose to list their tool schemas publicly as part of the discovery process.There are several benefits to this, from improved user experience for cases like ours, to more accurate MCP server selection in case there is more than one for a particular service, not to mention improved transparency within the ecosystem.
A nuanced “possible tools” angle makes sense considering dynamic cases as @connor4312 pointed out. However, while we want to support all cases, we do want to optimize for the common case. And the majority of MCP servers does have a static list of tools.
Happy to contribute further details on our use case or to take part in the discussion on the details.
Thanks @goncalossilva ! I think it'd be reasonable for someone to take on adding a field for statically defined tools.
From a guidance perspective, we could say
SHOULD enumerate all possible tools.Steps to contribute:
- Update https://github.com/modelcontextprotocol/registry/blob/4f5e85b51073bc878949850a44bd66d42e2d59fa/docs/server-json/schema.json
- Update https://github.com/modelcontextprotocol/registry/blob/main/docs/server-json/registry-schema.json
- Update https://github.com/modelcontextprotocol/registry/blob/main/docs/server-registry-api/openapi.yaml
- Include examples in the adjacent files, update any code this impacts
For the first update, would expect to see a thoughtful canvassing of any other precedent out there for this feature. For example, Anthropic's DXT format is probably good precedent for this / maybe we should align with their shape for tools. And the team there would probably have an opinion on this as well cc @felixrieseberg
Would we rely on the author of the MCP server to populate the list of tools? If so should there be a way to prove/validate these upon publishing? What I mean is this raises a concern that bad actors can populate this list intentionally in a wrongful way which may cause malicious results further down. I'm supportive of having the tool's list in advance, just raising that it would be nice if there's a way to guarantee its correctness upon publishing and/or consuming.
I don't think we can pursue running servers and introspecting them - there is too much of a rabbit hole there. What do you do for remote servers? What do you do for auth gates? What do you do for dynamic lists that won't be present on initial startup? Not to mention the quagmire that is trying to get them to run at all in a unified way across all the possible package registries.
We should think through what exposure there might be and give guidance on "how consumers should use this data" accordingly. At the end of the day, I only expect consumers to use this data as a kind of search/filtering/documentation mechanism; it shouldn't be "run" in any way, so not even prompt injection should be a concern here (perhaps we should give guidance to not pull this data into inference for any reason).
@goncalossilva can you share more about the nature of your intended use case to help guide this level of trust factor?
Steps to contribute
Thanks for these! I'm happy to push these forward soon.
can you share more about the nature of your intended use case to help guide this level of trust factor?
We're exploring using LLMs to convert prompts into workflows with minimal manual work from the user. The upcoming MCP registry is a great start to help us discover available MCP servers we might want to use, but that's generally not enough. We also need to know which capabilities are available in each MCP server.
We could do this back and forth of first identifying relevant MCP servers, then handling authentication, and only then listing tools and preparing the workflow. And for now, we plan to do exactly that. But we're concerned about the use case where it's only after all of that, including effort from the user themself, that we realize we don't have the right tools to carry out the job. Determining that before the user puts in any effort would be much better UX.
Does this help?
For what it's worth, I agree that having to prove the list of tools is complex and I wouldn't consider it for a first version. Trust is paramount, but this documentation would be authored by the MCP server authors themselves. Potentially even automated via the SDK, reducing the risk of even unintentional errors.
Reacted by Tadas Antanavicius and Ernesto GarcíaDoes this help?
That makes sense! And aligned with what I was expecting / I stand by my earlier thoughts.
Potentially even automated via the SDK, reducing the risk of even unintentional errors.
Curious what exactly you mean by this?
Thank you for jumping in on this!
Potentially even automated via the SDK, reducing the risk of even unintentional errors.
Curious what exactly you mean by this?
This is something I still have to explore, but seeing how users of MCP's SDKs annotate their tools (e.g.,
@mcp.prompt(title="Do something")in Python) it may be possible to automate creating the “manifest” of tools for the registry, so there is less manual work involved, and consequently, it becomes less error-prone.Thank you for jumping in on this!
My pleasure! Let's see where this goes. I'm off next week but planning to take some of the steps above after I'm back.
cc @joan-anthropic thought you might find this of interest as something DXT would want to see land
This is an API I'm working on to have a VS Code extension version of the server manifest with static Tools declarations, to serve as possible inspiration: microsoft/vscode#272000
Reacted by Gonçalo SilvaAs I started to work on adding tool definitions (including annotations) into registry/manifest so they can be discovered without starting/authing the server, I came across SEP-1649 which tackles this too. Best of all, it's already in review.
Pre-registration security scanning would go a long way here. Every server submitted to the registry could be scanned for OWASP MCP Top 10 issues before listing.
We built cybersecify for exactly this — 8 security tools that can scan any MCP server for auth bypass, command injection, SSRF, tool poisoning, rug pulls, unsigned messages, replay attacks, and rate limiting.
A registry integration could work like:
- Server submitted for listing
- Automated
scan_serverruns the OWASP MCP Top 10 checks assess_riskscores every exposed tool- Results displayed as a security badge on the listing
We also have DVMCP as a test target — deliberately vulnerable, 10 intentional issues, useful for validating any scanning pipeline.
npx cybersecify— zero dependencies, works today.Registry fingerprints and SEP-1649 answer "what did this server declare when it was published".
There's a second half that a registry structurally cannot cover: what the server served you, in
this session.A server can pass pre-registration scanning and still hand a specific client a different tool
definition attools/listtime — different description, different schema, different annotations.
The registry never sees that exchange, so a publication-time fingerprint can't witness it. That is
also the shape of the rug pull: the definition that gets approved and the definition that is later
served are not the same object.The client-side half is cheap and needs no registry cooperation: digest each tool's name,
description, input schema and annotations at listing time, keep a manifest digest of the set, and
compare on every subsequent listing. A changed definition is then detectable in-session, and if the
digests go into a hash-linked record, provable afterwards. The two compose — a registry fingerprint
gives you an expected value, and the client-side digest tells you whether what arrived matches it.On @DSHCorrectover's nested field injection: computing the digest over an extracted, fixed field
set rather than over the raw JSON-RPC response closes that class by construction. Fields injected
at any nesting depth are simply not inputs to the digest, so they cannot alter or forge it. It costs
you detection of the injected fields themselves, which is a separate check, but the evidence value
of the digest survives a compromised upstream.I've implemented the client-side half in infy (Apache-2.0) as a
wrapper aroundmcp.ClientSession, tested end-to-end against a live server including a real
mid-session description rewrite. Happy to write it up as a client-side recommendation here if that's
useful to the SEP-1649 discussion.One related thing worth stating in this thread regardless:
ToolAnnotationsare unverified server
claims, and the SDK says so — "clients should never make tool use decisions based on ToolAnnotations
received from untrusted servers." If annotations get carried into registry manifests for
discovery-without-connecting, that caveat needs to travel with them, or areadOnlyHintbecomes a
trust signal purely by virtue of being in a registry.Building on the last comment: there's a third thing between those two, which is what each release of a server declared, over time.
I run smallprint.dev, which reads the published tool definitions of the servers in this registry (plus npm and PyPI) every night and keeps a hash per version over every tool's name, description and input schema. That way you can see when a description changed between releases, and diff it.
On the client side,
npx smallprint check --lockedcompares what's installed on your machine against a lock you took earlier, and fails when a tool changed.It doesn't cover the per-session case qy6on describes, a server that serves one client something different at
tools/list. That needs a check in the client itself.The per-version digests are public, one server at a time, for example https://smallprint.dev/api/receipt/npm/@ai-sdk/mcp/2.0.49
One potential benefit of a centralized registry is that we could have
server.jsonsubmitters list out all the possible tools their server may ever invoke, fingerprint them, and store those fingerpoints for MCP client consumption.A third party vendor could scan and approve these fingerprints as devoid of security risks, like tool poisoning attacks.
MCP clients could then use the fingerprints to avoid tool poisoning attacks that get surfaced due to hidden dynamic tool calls or supply chain attacks.