Skip to content

docs: add limits of tool search document - #15

Open
kurtisvg wants to merge 3 commits into
mainfrom
docs/limits-of-tool-search
Open

kurtisvg wants to merge 3 commits into
mainfrom
docs/limits-of-tool-search

Conversation

@kurtisvg

Copy link
Copy Markdown

Analysis of why Tool Search is not sufficient progressive discovery (from @SamMorrowDrums, @helloeve, and myself).

- Add comprehensive analysis of native tool calling mechanisms and limitations of tool search
- Compare server, proxy, client, and model-provided tool search approaches
- Detail tool blindness, lookup latency, prompt caching impact, and lessons from Skills
cliffhall
cliffhall previously approved these changes Sep 7, 2026

@cliffhall cliffhall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM! 👍

@LucaButBoring

LucaButBoring commented Sep 22, 2026 •

Copy link
Copy Markdown

I think this is a great review of the problem space, but the suggested solution space seems inconsistent with those problems - as I read it, there are a few existing ways to do tool search and they're all flawed because tool search itself is hard to do well:

  • Make the client deal with it: Tool search isn't easy to build well - it ideally needs to be done without busting the prompt cache, which means having an immutable tool list.
  • Make the server deal with it via wrapping its tools in action-executors: Creates indirection and results in proliferation of search tools. I would also argue that tool search not being easy to build well applies here, too.
  • Use a proxy search server: Becomes an SPOF for auth and availability, and also creates a level of indirection.
  • Make the inference provider deal with it: Great for prompt caching, but there's no consistency between or within providers.
  • Tool search in general also introduces additional model turns and tool-blindness.

The suggested solution space inherits several of the same problems, however:

  • Skill-driven discovery: Requires either busting the prompt cache or defining a wrapper tool (and hence introducing indirection).
  • Appendable tool definitions: Requires inference provider support, which mostly isn't present today except in OpenAI's Responses API, and will likely be inconsistent between/within providers in the long run.

If we're looking to carve out a tractable solution space, I think we need to narrow the problem space as well - we can effectively distill a few interrelated constraints inherent to progressive disclosure from this document, irrespective of if we use tool search or some other mechanism:

  1. Conversations are append-only: Basically no inference providers support declaring a tool mid-conversation (without first declaring it in the top-level tool list) as of today. This is an inference provider problem which we cannot solve ourselves, but we should try to shape our proposed solution to be compatible with mid-conversation tool declarations in case that ever does become widely-supported. Mid-conversation tool declarations notwithstanding, whatever solution we introduce needs to be cache-friendly; ergo, it must somehow be compatible with an append-only conversation format.
  2. Agent turns are expensive: Updating the top-level tool list busts the prompt cache, and agents incur quadratic costs over the course of a conversation, plus cache writes are much more expensive than cache reads. The underlying economics don't favor busting the prompt cache outside of compaction (there's a bunch of other nuances and exceptions to that but it's a good rule of thumb).
  3. Action-executor tools introduce a level of indirection: An easy way to keep the prompt cache happy is to route all of your tools through a single action-executor that takes the upstream tool name as an argument. However, this means throwing away a certain degree of "model-friendliness" given that models are largely trained on native tools (I actually think the validity of this point is debatable but we can accept it for now). If the client itself defines the action-executor tool that it's wrapping other servers in, this also means that the server itself does not know how it's being exposed to the LLM and can't effectively optimize for any particular paradigm. We discussed variants as a possible solution to this previously, which may or may not be worth mentioning in the document.
  4. Any level of indirection worsens tool blindness: Native tool definitions carry the full tool and argument descriptions with them, and force them into the context so the model is basically unable to ignore them. If you put a layer between the model and the tool you want to invoke, you lose this property and need to make up for it some other way (perhaps via skills).

Presenting a strong solution to progressive disclosure in MCP will likely mean accepting these constraints and working within them, rather than discovering some way to work around them entirely. This will mean inheriting at least some of the problems tool search has today while hopefully mitigating others.

I think the way the document currently frames the solutions doesn't clearly emphasize the underlying constraints imposed by inference economics (it does mention them, it just doesn't center them as themes), so it's unclear how any of the suggested solutions might actually be better or worse than tool search in the end. The information is great, it just feels directionless when zoomed-out because the throughline is a bit weak.


Anyways my meta-suggestion here is honestly to rework the exploration of the solution space significantly (everything from "We can learn a lot from Skills" onwards; everything before that is mostly fine) and put something roughly along the lines of the underlying constraints I have above in between the problem and solution sections - not to say that what I have above is necessarily the precise set of constraints to discuss, but it's something to start with at least.

@kurtisvg

Copy link
Copy Markdown
Author

Thanks for the thorough review, @LucaButBoring.

As we've discussed on Discord, OpenAI supports additional_tools, and Claude now also supports inline tool definitions in beta. Previously undeclared tools can therefore become natively callable mid-conversation. That’s distinct from extending a provider-hosted search catalog, whose support remains API-specific.

I’ve tightened the doc to address the broader feedback:

  • Clarified the native-calling, caching, and provider-support tradeoffs.
  • Separated implementation constraints from discovery limitations.
  • Made explicit that grouping can improve discovery while still inheriting activation, registration, and discoverability constraints.

The Skills example now illustrates a possible grouping mechanism for MCP primitives, with its remaining limitations called out.

Copilot AI lite review requested due to automatic review settings September 28, 2026 14:57
Copilot stopped reviewing on behalf of kurtisvg due to an error September 28, 2026 15:18

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Copilot review overview

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Adds a new documentation page explaining why tool search/progressive discovery is difficult to integrate cleanly with native tool calling and prompt caching across MCP and major model providers.

Changes:

  • Document tradeoffs of server/proxy/client/provider-hosted tool search approaches and their impact on native tool calling.
  • Add a provider support matrix plus discussion of mid-conversation tool registration.
  • Describe practical limitations (tool blindness, repeated searches) and relate them to Skills/CLI-style grouping.
File Description
docs/​limits-of-tool-search.md New doc detailing limitations and patterns for tool search and progressive tool discovery with MCP and provider tool-calling APIs.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +204 to +207
| API or platform | Deferred tool search | Scope and limitations |
| :---- | :---- | :---- |
| OpenAI Responses API | Native | GPT-5.4 and later; hosted and client-executed search. ([OpenAI](https://developers.openai.com/api/docs/guides/tools-tool-search)) |
| Azure OpenAI Responses API | Native | GPT-5.4 and later on a supported deployment; hosted and client-executed search. ([Microsoft](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/tool-search)) |
`get_build_logs` never becomes a native tool. Its schema is data returned by
`search_tools`, while the actual native call is the weakly typed `execute_tool`.
The model provider can validate only the generic executor's schema, not the
arguments of the discovered tool.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants