Teaching an LLM to pull MCP Resources and Prompts on demand (instead of drowning it in context)
Teaching an LLM to pull MCP Resources and Prompts on demand (instead of drowning it in context) How we wired the Model Context Protocolβs βapplication-controlledβ primitives into a model-controlled tool-calling loop β and why that small shift changes everything about context hygiene. TL;DR The Model Context Protocol (MCP) gives a server three ways to expose capability: tools, resources, and prompts. Tools drop straight into an LLMβs function-calling loop. Resources and prompts donβt β theyβre application-controlled, so most integrations just dump every resourceβs content into the system prompt and hope for the best. That approach bloats context, truncates large documents, breaks on binary files, and gives the model zero say in what it actually needs. Our fix: promote resources and prompts into synthetic, auto-approved LLM tools β read_resource(uri) and invoke_prompt(name). The system prompt now carries only a lightweight catalog (URIs + descriptions). The model reads a resource only when it decides it needs one, through the exact same tool-calling machinery it already uses. On-demand, selective, full-fidelity. The three MCP primitives β and the control model nobody talks about MCP defines three server capabilities, but the interesting part is whoβs in control of each: Primitive Who decides when itβs used Natural fit for tool-calling? Tools The model (it calls them) β
Yes β this is what function calling is Resources The application / user β No native hook in the loop Prompts The user (usually a slash-command) β No native hook in the loop Tool calling is model-controlled by design: the LLM emits a tool_use block, you execute it, you feed the result back. Beautiful. Resources and prompts are application-controlled. The specβs mental model is a human clicking βattach this fileβ or β/use this prompt template.β There is no obvious place for them inside an autonomous agentβs reasoning loop. So what do most integrations do? The naive approach (and why it hurts) The path of least resistance is to fetch every resource at startup and paste it into the system prompt: ## Available Resources ### Resource: SUM ABAP Test Matrix URI: sap-btp://sum-abap-v1 Content: <β¦ 11,000 characters of markdown β¦> ### Resource: API Docs URI: sap-btp://api-docs Content: <β¦ more β¦> Four problems show up fast: Context bloat. Every request pays for every resource, whether or not itβs relevant. Ten resources Γ a few thousand tokens each = a system prompt that dwarfs the actual conversation. Truncation. To keep bloat sane, you cap content (content[:2000]) β and now large documents are silently chopped. In our case the SUM ABAP matrix lost its entire product list and output-format section below the 2,000-char line. Binary breaks. A PDF or PNG resource has no meaningful text form. Naive extractors try blob.as_string(), hit a UnicodeDecodeError, and quietly emit β[No content available]β. No agency. The model canβt say βI donβt need any of these right nowβ or βgive me that one, in full.β Itβs force-fed. The insight: make resources and prompts look like tools Hereβs the shift. The LLM already has a clean, well-understood way to ask for something on demand: it calls a tool. So instead of fighting the control model, we translate it. We register two DARA-internal tools that donβt exist on any MCP server β theyβre synthesized client-side: read_resource(uri) β fetches one resourceβs content when the model asks. invoke_prompt(name) β injects a named prompt template into the conversation when the model asks. The system prompt now advertises only a catalog β names, URIs, and descriptions, no content: ## Available Resources The following resources can be read on demand. To read one, call the read_resource tool with its exact URI. Do not assume a resourceβs contents until you have read it. ### Resource: SUM ABAP Test Matrix URI: sap-btp://sum-abap-v1 Description: SUM (Software Update Manager) test matrix specification for ABAP products The model sees what exists, then reaches for exactly what it needs β through the tool loop it already speaks fluently. Implementation walkthrough We built this on top of LangGraph + langchain-mcp-adapters, but the pattern is framework-agnostic. 1. Synthesize the tool with a schema read_resource is a StructuredTool with a one-field schema. Its func is a no-op lambda β we never actually run it as a function; we intercept it in the graph (see step 3). from pydantic import BaseModel, Field from langchain_core.tools import StructuredTool class _ReadResourceArgs(BaseModel): uri: str = Field(description=βThe exact URI of the MCP resource to read, β βe.g. βsap-btp://sum-abap-v1β.β) read_resource_tool = StructuredTool.from_function( func=lambda uri: "", # placeholder β handled in the graph name=βread_resourceβ, description=( βRead the full contents of an MCP resource by its URI. β βCall this when the user asks to read, open, or summarize a resource, β βor when you need a resourceβs contents to answer. β βOnly resources listed under βAvailable Resourcesβ can be read.β ), args_schema=_ReadResourceArgs, ) tools = list(tools) + [read_resource_tool] allowed_tools_without_review.append(βread_resourceβ) # auto-approve, no human gate Two things matter here: The description is the UX. Itβs the only instruction the model gets on when to call this. Write it like a prompt, because it is one. Auto-approval. In an agent with human-in-the-loop review, reading a read-only resource shouldnβt require a click. We add it to the βno reviewβ allow-list so the graph routes straight to execution. 2. Build a content map once, keyed by URI We fetch resources once (the adapter already returns their content as Blob objects) and index them by URI so the on-demand read is an O(1) lookup β no second network round-trip: resources_content = {} for res in (resources or []): uri, name, description, content = _extract_resource_fields(res) # URIs can arrive as pydantic AnyUrl objects β normalize to str so the # plain-string URI the LLM passes actually matches the dict key. (Gotcha!) uri = str(uri) if uri is not None else "" if uri and uri != βunknownβ: resources_content[uri] = {βnameβ: name, βcontentβ: content} 3. Intercept the tool call in the graph When the model emits a read_resource call, we donβt invoke a function β we look up the content and hand it back as a tool message. Because a tool result flows naturally back into the modelβs context, the content lands only when requested: if tool_call[βnameβ] == βread_resourceβ: uri = str(tool_call[βargsβ].get(βuriβ, "")) res = resources_content.get(uri) if res: new_messages.append({ βroleβ: βtoolβ, βnameβ: βread_resourceβ, βcontentβ: res[βcontentβ], # full content, no truncation βtool_call_idβ: tool_call[βidβ], }) else: available = list(resources_content.keys()) new_messages.append({ βroleβ: βtoolβ, βnameβ: βread_resourceβ, βcontentβ: fβResource β{uri}β not found. Available URIs: {available}β, βtool_call_idβ: tool_call[βidβ], }) continue The mirror image for prompts: invoke_prompt injects the templateβs messages as real Human/AI turns, which is exactly how the MCP spec intends prompts to be surfaced. 4. Handle binary honestly (donβt decode bytes as text) langchain-mcp-adapters collapses every resource into a Blob: text lands in .data as a str, binary as raw bytes, with the media type on the .mimetype attribute (not in metadata β a common trip-up). So we branch on the mime type instead of blindly calling .as_string(): def _is_text_mime(mime: str) -> bool: mime = (mime or "").lower().split(β;β)[0].strip() return (mime.startswith(βtext/β) or mime.endswith((β+jsonβ, β+xmlβ, β+yamlβ)) or mime in {βapplication/jsonβ, βapplication/xmlβ, βapplication/yamlβ}) # inside the extractor, for a Blob: data = getattr(resource, βdataβ, None) if isinstance(data, bytes): if _is_text_mime(mime): return uri, name, description, data.decode(βutf-8β) # Binary (PDF, PNG, β¦) β an honest descriptor, NOT raw bytes/base64 return uri, name, description, ( fβ[Binary resource: {mime or βapplication/octet-streamβ}, β fβ{_human_size(len(data))}. This is not text and cannot be inlined; β fβopen it with a client that handles its media type.]β ) This is aligned with the MCP spec itself: binary payloads belong in typed media content blocks or are referenced by URI β never stuffed into a text field. A model reading [Binary resource: application/pdf, 240.0 KB] knows exactly what itβs looking at and can decide what to do, instead of choking on garbage or getting a misleading βno content.β The flow, end to end βββββββββββββββββββ β User message β ββββββββββ¬βββββββββ β βΌ ββββββββββββββββββββββββββββββββββββββββββββ β System prompt = resource CATALOG only β β (URIs + descriptions, NO content) β ββββββββββ¬ββββββββββββββββββββββββββββββββββ β βΌ βββββββββββββββββ β LLM decides β ββββ¬ββββββββββ¬βββ β β needs a β β doesnβt need one resourceβ ββββββββββββββββΊ Answer directly βΌ ββββββββββββββββββββββββββββ β tool_use: read_resource β β (uri) β ββββββββββββββ¬ββββββββββββββ βΌ ββββββββββββββββββββββββββββββββ β Graph intercepts the call β β (auto-approved, no HITL gate) β ββββββββββββββ¬ββββββββββββββββββ βΌ ββββββββββββββββββββββββββββββββ β Look up URI in content map β βββββββββ¬ββββββββββββββββ¬βββββββ β text β binary βΌ βΌ ββββββββββββββββββ ββββββββββββββββββββββββ β Full content β β Descriptor: β β as tool messageβ β mime type + size β βββββββββ¬βββββββββ βββββββββββββ¬βββββββββββ β β βββββββββββββ¬ββββββββββββ βΌ (back to LLM βββΊ answer) Only the βneeds a resourceβ branch ever pays the content cost β and it pays the full cost, untruncated, for just that resource. Why this is nice Token-efficient. The system prompt holds a catalog (tens of tokens per resource), not a library. Content enters context only on read. Selective. The model β or the user, phrasing a request β decides what to load. Ten irrelevant resources cost almost nothing. Full fidelity. No truncation cap needed, because youβre no longer defending against ten simultaneous dumps. The one resource you asked for arrives whole. Binary-safe. Mime-driven handling means PDFs and images degrade to a clear descriptor instead of a crash or a lie. Spec-aligned. Application-controlled primitives stay application-controlled β we just expose an affordance for the model to request them, rather than pre-deciding on its behalf. Uniform mental model. Resources, prompts, and real tools all flow through one loop. No special-case rendering paths, no bespoke context-stuffing logic. Gotchas worth stealing AnyUrl vs str. MCP resource URIs often surface as pydantic AnyUrl objects. If your content map is keyed by AnyUrl and the LLM passes a plain string, .get() silently misses. Normalize to str on both sides. Mime lives on .mimetype, not metadata. For langchain-mcp-adapters Blobs, the media type is an attribute; metadata only carries the uri. Read the right field or every binary looks like application/octet-stream. The tool description is the routing logic. Thereβs no separate policy telling the model when to read a resource β only the toolβs description. Invest in it. Auto-approve read-only reads. If your agent has human review gates, forcing a click to read a read-only resource kills the UX. Allow-list it. A no-op func is fine. The synthesized tool never executes as a function; the graph intercepts it. The lambda is just there to satisfy the schema. Where it goes next The binary branch is the single hook for richer handling: extract PDF text server-side, or emit an ImageContent block to a vision-capable model for images. Because everything already funnels through one read_resource path, adding a modality is a localized change β not a re-architecture. The bigger takeaway: when a protocol primitive doesnβt fit your execution model, donβt force the model to swallow it up front. Give the model an affordance to ask, and let the loop it already understands do the rest. Built on the Model Context Protocol (2025-03-26), LangGraph, and langchain-mcp-adapters. The pattern is framework-agnostic β anywhere you have tool-calling and MCP, you can do this.