MCP sampling and elicitation — the newest attack surface in the spec
MCP sampling lets servers invoke the client LLM and query users directly, inverting trust. Review controls for that inversion.
The established MCP threat model runs in one direction: the client's agent decides to call a tool, the server executes it and returns a result, and the review questions cluster around what that result contains and what the tool description talked the model into doing. Transport, tool surface, context injection, and the authentication boundary all describe risks that live inside that direction of control — the client initiates, the server responds.
Two primitives in the MCP specification do not fit that direction. Sampling lets a server request that the client's LLM generate a completion on the server's behalf — the server, not the client's agent, decides an inference call should happen, and shapes the prompt that drives it. Elicitation lets a server request structured input directly from the human user, mid-session, outside the orchestrating agent's own reasoning about what it needs to ask. Both invert who is “in control” of a step in the interaction, and a review built entirely around the client-calls-server model has no established place to put either one.
This piece is about what a review has to check specifically for these two primitives, on top of — not instead of — the established MCP threat model.
Why this is not the established MCP model again
Transport security, tool-surface poisoning, context injection through tool results, and the client-versus-user authentication gap are all still live concerns for any MCP deployment that also uses sampling or elicitation — a server capable of requesting a sampled completion is still subject to the same transport and authentication requirements as one that only exposes tools. None of that goes away, and a review that has not covered it is not covered on the fundamentals regardless of whether sampling is in scope.
What is genuinely new is the direction of initiative. Every threat in the established model assumes the server is reactive: it responds to a call the client's agent chose to make. Sampling and elicitation make the server proactive in a specific, protocol-sanctioned way — it can originate a request into the client's own inference capability or directly to the human. A review that only asks “what happens when the agent calls this server” never asks the question that matters for these two primitives: “what happens when the server calls the agent, or the user, instead.”
What sampling is
Sampling is a request type an MCP server can send to the client: “generate a completion, using this prompt, with these constraints.” The client is not required to honor it — the specification frames sampling as something the client mediates, typically with a human-in-the-loop approval step before the request is actually sent to the model — but the request itself originates from the server, and the server supplies the prompt content the client's LLM will process.
The intended use case is legitimate and, in isolation, useful: a server that needs the client's LLM to summarize something, extract structured data from a document the server has access to, or make a judgment call as part of completing a task, without the server needing its own model access or API key. Instead of the server running its own inference (which it may not have budget or credentials for), it borrows the client's.
What elicitation is
Elicitation is a request type a server can send to ask the human user a structured question directly — a form, a set of fields, a yes/no confirmation — and receive the answer back into the server's own processing, rather than the answer being mediated by the orchestrating agent's reasoning about what to ask and why.
The legitimate use case mirrors sampling's: a server midway through a task discovers it needs a piece of information only the user can supply — a missing parameter, a disambiguation between two plausible interpretations, a confirmation before a consequential action — and elicitation lets it ask for that directly instead of returning control to the agent and hoping the agent phrases the right question back to the user.
Why sampling inverts the trust direction
In the standard agentic model, every inference call is initiated by the client's own orchestration logic, and every prompt that reaches the model passes through whatever review, logging, and injection-resistance controls the client has built around its own prompt construction. That is the assumption most of the existing agentic AI security review — goal anchoring, prompt-injection resistance, output classification — is built on: the client controls what gets asked.
Sampling breaks that assumption cleanly. A tool server — potentially one the security review has scoped narrowly, on the theory that it can only do what its tool manifest permits — can trigger an LLM call and shape the prompt that drives it. If that server-originated prompt is not put through the same review the client's own prompts receive, it is an unaudited path to model behavior: content the server controls reaches the model as an authoritative generation request, not as untrusted data the model is merely reading.
Tool calls vs sampling vs elicitation — who initiates, and what's new
| Tool call (established) | Sampling | Elicitation | |
|---|---|---|---|
| Who normally initiates | Client's agent decides to call a tool | Server requests a completion from the client's LLM | Server requests input directly from the human user |
| What is exposed | A scoped function call with typed parameters | A model inference call, with a prompt the server shapes | A form or question surfaced to the user, framed by the server |
| New risk | Covered by the tool-surface and context-injection model | An unaudited path to model behavior, run on the user's quota and context | A social-engineering vector that can look like ordinary agent UX |
A tool call is the server asking the agent to do something. Sampling is the server asking the model to think something. The second is a much larger grant of trust to hand a component the review scoped as “a narrow tool with a fixed manifest.”
A server running its own agenda through your quota
The concrete risk that follows from the trust inversion is straightforward to state: a malicious or compromised MCP server can use sampling to run inference the user never asked for, on the user's own model quota and within the user's own session context, framed as a legitimate part of completing the task the user actually requested. Because the request travels through the sanctioned sampling channel rather than through a scoped tool call, it does not show up in a tool-call audit log the same way a tool invocation would — the review artefact most teams already collect for MCP deployments is blind to it by construction.
This is meaningfully different from a server simply having too broad a tool manifest, which is a scoping problem the existing least-privilege review already catches. Sampling abuse is not about the server doing something with a tool it should not have — it is about the server directing what the client's own model does, without ever calling a tool at all.
Consent and visibility requirements
Both risks reduce to the same underlying question: does the client surface a sampling or elicitation request to the user as something distinct from a normal tool-call result, or does it blend invisibly into the conversation as though it came from the orchestrating agent itself? A client that clearly labels “Server X is requesting a model completion” or “Server X is asking you a question directly” gives the user (and a reviewer inspecting session logs afterward) the information needed to evaluate whether the request is appropriate. A client that renders the result of a sampling call or the prompt of an elicitation request identically to the agent's own output removes that signal entirely.
This is a client-implementation property, not a protocol guarantee — the specification permits a human-in-the-loop approval step for sampling, but does not mandate a specific visual treatment, and elicitation's framing is similarly left to the client. A review of any MCP client library has to check what that specific client actually does, not what the specification permits it to do.
Review checklist
For any system where an MCP client library supports sampling or elicitation, a review is complete on this surface when it can answer each of these with a configuration artefact or a session transcript, not a description of the protocol's general design:
- Are sampling and elicitation enabled by default for every connected server, or opt-in per server with an explicit allowlist?
- Is there a distinct, user-visible indicator when a server-initiated sampling or elicitation event occurs, separate from the agent's own conversational turns?
- Does a sampling request from a server pass through the same prompt-review and injection-resistance controls applied to the client's own prompts, or does it bypass them as a separate code path?
- Is there a rate or scope limit on how much inference a single server can trigger via sampling in a given session or time window?
- For elicitation specifically: is there a policy restricting what categories of information a server is permitted to request directly from the user (credentials, payment details, and similar should not be requestable through this channel at all)?
- Does the audit log capture sampling and elicitation events as their own event type, distinguishable from ordinary tool calls, so a post-incident review is not blind to this channel?
Where a client cannot answer these — where sampling and elicitation are supported but invisible in the logs and indistinguishable from the agent's own behavior in the UI — the honest finding is that the deployment has granted every connected server a channel to the model and the user that the rest of the MCP review never evaluates.
Review the primitives that invert the trust direction, not just the tool surface.
Drel's MCP security review covers sampling and elicitation as their own surface — distinct from the tool manifest and context-injection checks — so a connected server's ability to direct model behavior or solicit user input gets the scrutiny the tool surface already gets.
A note on scope: Drel reviews assessed systems against documented architecture, configuration and intent. It does not ingest live telemetry from production environments. Dispositions reflect the assessed system at the time of review and the re-assessment triggers that govern when the disposition must be revisited.