BlogTechnical

MCP sampling and elicitation — the newest attack surface in the spec

MCP sampling lets servers invoke the client LLM and query users directly, inverting trust. Review controls for that inversion.

Drel11 min read

The established MCP threat model runs in one direction: the client's agent decides to call a tool, the server executes it and returns a result, and the review questions cluster around what that result contains and what the tool description talked the model into doing. Transport, tool surface, context injection, and the authentication boundary all describe risks that live inside that direction of control — the client initiates, the server responds.

Two primitives in the MCP specification do not fit that direction. Sampling lets a server request that the client's LLM generate a completion on the server's behalf — the server, not the client's agent, decides an inference call should happen, and shapes the prompt that drives it. Elicitation lets a server request structured input directly from the human user, mid-session, outside the orchestrating agent's own reasoning about what it needs to ask. Both invert who is “in control” of a step in the interaction, and a review built entirely around the client-calls-server model has no established place to put either one.

This piece is about what a review has to check specifically for these two primitives, on top of — not instead of — the established MCP threat model.

Why this is not the established MCP model again

Transport security, tool-surface poisoning, context injection through tool results, and the client-versus-user authentication gap are all still live concerns for any MCP deployment that also uses sampling or elicitation — a server capable of requesting a sampled completion is still subject to the same transport and authentication requirements as one that only exposes tools. None of that goes away, and a review that has not covered it is not covered on the fundamentals regardless of whether sampling is in scope.

What is genuinely new is the direction of initiative. Every threat in the established model assumes the server is reactive: it responds to a call the client's agent chose to make. Sampling and elicitation make the server proactive in a specific, protocol-sanctioned way — it can originate a request into the client's own inference capability or directly to the human. A review that only asks “what happens when the agent calls this server” never asks the question that matters for these two primitives: “what happens when the server calls the agent, or the user, instead.”

What sampling is

Sampling is a request type an MCP server can send to the client: “generate a completion, using this prompt, with these constraints.” The client is not required to honor it — the specification frames sampling as something the client mediates, typically with a human-in-the-loop approval step before the request is actually sent to the model — but the request itself originates from the server, and the server supplies the prompt content the client's LLM will process.

The intended use case is legitimate and, in isolation, useful: a server that needs the client's LLM to summarize something, extract structured data from a document the server has access to, or make a judgment call as part of completing a task, without the server needing its own model access or API key. Instead of the server running its own inference (which it may not have budget or credentials for), it borrows the client's.

What elicitation is

Elicitation is a request type a server can send to ask the human user a structured question directly — a form, a set of fields, a yes/no confirmation — and receive the answer back into the server's own processing, rather than the answer being mediated by the orchestrating agent's reasoning about what to ask and why.

The legitimate use case mirrors sampling's: a server midway through a task discovers it needs a piece of information only the user can supply — a missing parameter, a disambiguation between two plausible interpretations, a confirmation before a consequential action — and elicitation lets it ask for that directly instead of returning control to the agent and hoping the agent phrases the right question back to the user.

Why sampling inverts the trust direction

In the standard agentic model, every inference call is initiated by the client's own orchestration logic, and every prompt that reaches the model passes through whatever review, logging, and injection-resistance controls the client has built around its own prompt construction. That is the assumption most of the existing agentic AI security review — goal anchoring, prompt-injection resistance, output classification — is built on: the client controls what gets asked.

Sampling breaks that assumption cleanly. A tool server — potentially one the security review has scoped narrowly, on the theory that it can only do what its tool manifest permits — can trigger an LLM call and shape the prompt that drives it. If that server-originated prompt is not put through the same review the client's own prompts receive, it is an unaudited path to model behavior: content the server controls reaches the model as an authoritative generation request, not as untrusted data the model is merely reading.

Tool calls vs sampling vs elicitation — who initiates, and what's new

Tool call (established)SamplingElicitation
Who normally initiatesClient's agent decides to call a toolServer requests a completion from the client's LLMServer requests input directly from the human user
What is exposedA scoped function call with typed parametersA model inference call, with a prompt the server shapesA form or question surfaced to the user, framed by the server
New riskCovered by the tool-surface and context-injection modelAn unaudited path to model behavior, run on the user's quota and contextA social-engineering vector that can look like ordinary agent UX
A tool call is the server asking the agent to do something. Sampling is the server asking the model to think something. The second is a much larger grant of trust to hand a component the review scoped as “a narrow tool with a fixed manifest.”

A server running its own agenda through your quota

The concrete risk that follows from the trust inversion is straightforward to state: a malicious or compromised MCP server can use sampling to run inference the user never asked for, on the user's own model quota and within the user's own session context, framed as a legitimate part of completing the task the user actually requested. Because the request travels through the sanctioned sampling channel rather than through a scoped tool call, it does not show up in a tool-call audit log the same way a tool invocation would — the review artefact most teams already collect for MCP deployments is blind to it by construction.

This is meaningfully different from a server simply having too broad a tool manifest, which is a scoping problem the existing least-privilege review already catches. Sampling abuse is not about the server doing something with a tool it should not have — it is about the server directing what the client's own model does, without ever calling a tool at all.

Elicitation as a social-engineering vector

Elicitation carries a different, more direct risk: a server can solicit sensitive input straight from the user — a credential, a confirmation the user would not have given if they understood who was actually asking, a piece of personal information beyond what the visible task requires — under the appearance of ordinary agent interaction. To the user, an elicitation request and a question the orchestrating agent itself decided to ask can be visually and behaviorally indistinguishable, which is exactly what makes it a usable social-engineering vector rather than an obviously external request.

The failure mode is not hypothetical in shape, even where any specific incident is: a compromised or intentionally malicious third-party MCP server connected for one legitimate purpose (say, a scheduling integration) uses elicitation mid-task to ask the user to “confirm your account password to continue” or “re-enter your payment details to complete this booking,” and the request arrives inside the same conversational surface the user already trusts for the agent's own questions.

Review checklist

For any system where an MCP client library supports sampling or elicitation, a review is complete on this surface when it can answer each of these with a configuration artefact or a session transcript, not a description of the protocol's general design:

  • Are sampling and elicitation enabled by default for every connected server, or opt-in per server with an explicit allowlist?
  • Is there a distinct, user-visible indicator when a server-initiated sampling or elicitation event occurs, separate from the agent's own conversational turns?
  • Does a sampling request from a server pass through the same prompt-review and injection-resistance controls applied to the client's own prompts, or does it bypass them as a separate code path?
  • Is there a rate or scope limit on how much inference a single server can trigger via sampling in a given session or time window?
  • For elicitation specifically: is there a policy restricting what categories of information a server is permitted to request directly from the user (credentials, payment details, and similar should not be requestable through this channel at all)?
  • Does the audit log capture sampling and elicitation events as their own event type, distinguishable from ordinary tool calls, so a post-incident review is not blind to this channel?

Where a client cannot answer these — where sampling and elicitation are supported but invisible in the logs and indistinguishable from the agent's own behavior in the UI — the honest finding is that the deployment has granted every connected server a channel to the model and the user that the rest of the MCP review never evaluates.

Review the primitives that invert the trust direction, not just the tool surface.

Drel's MCP security review covers sampling and elicitation as their own surface — distinct from the tool manifest and context-injection checks — so a connected server's ability to direct model behavior or solicit user input gets the scrutiny the tool surface already gets.

A note on scope: Drel reviews assessed systems against documented architecture, configuration and intent. It does not ingest live telemetry from production environments. Dispositions reflect the assessed system at the time of review and the re-assessment triggers that govern when the disposition must be revisited.

Related hub pages