Local Orchestrator Uses Narrow Tool Set for Small Models
October 5, 2026 · 4 min read
Also published in עברית, العربية, Nederlands
The orchestrator runs on a single RTX 3090 24Gb with Qwen3.6 27B and handles more than ninety tools. It selects a tool at each step based on the request. The question is whether to show all tools or only those that match the request. The orchestrator can work with both large cloud models and small models that fit on a consumer GPU. The minimal working setup uses RTX 3090 24Gb plus Qwen3.6 27B. The goal is to run corporate tasks on this minimal setup with acceptable quality and speed. Usually the orchestrator displays the full list of tools and the model chooses which one to call. Hermes, OpenClaw and Ouroboros follow this pattern. Some agents provide a catalog mode that only shows names and short descriptions, allowing the model to load the needed tool itself. These agents are designed for frontier models such as Claude and GPT‑5 that can hold a hundred tools without loss. The author’s task is different: they need to run a 27B model in a closed loop on a single RTX 3090. For such a model a narrow set performs better. Before the main call a router examines the request text, picks a section such as “datasets”, “mail”, “presentations”, and then a sub‑section. The model receives descriptions of nearly ninety tools, but only about twenty of them. It does not know about the remaining tools at this stage. Experiments in August measured how often the model chose the correct tool. With a narrow set of forty‑three correct choices out of forty‑four, the full set achieved only thirty‑nine correct out of forty‑four. The full context was 2.7 times larger and the run took one‑third longer. The author wonders whether a more precise tool description could raise accuracy to one hundred percent, but the current results show a clear gap.
Four failures all belonged to a family of tools with similar names. One was renamed from “combine datasets” to “manage datasets”, another from “draw chart” and “rename columns” to “analyze dataset”, and a third from “extract news sites” to “raw page load”. The weak model does not make random mistakes; it confuses similar candidates. A narrow set helps not because it is shorter but because it contains fewer mutually similar options. The hypothesis that a strong model needs a wide set was tested in July with an eight‑query exploratory run comparing a local model to GPT‑4.1. The local model answered seven of eight correctly on the narrow set and six of eight on the full set, while GPT‑4.1 answered seven of eight correctly on the narrow set and all eight on the full set. This suggests that a weak model benefits from a narrow set, whereas a strong model may be limited by it. The author therefore added a routing mode that can be configured separately for each model. Although eight queries are not proof, they provide a useful benchmark. In September the author ran the entire corpus on two cloud models, GigaChat‑2‑Max and gpt‑oss 120B, using all four selection modes available in the product. The corpus had grown to ninety‑one tools. Both cloud models performed better with the narrow set, achieving higher accuracy, smaller context size and lower latency. The hypothesis that strong models need a wide set was not confirmed. When analyzing the failures, half of them were not actual selection errors. In the narrow set, GigaChat missed three clarifying questions such as “which project key?” and “in which slot is the dataset?”. If such clarifying questions are counted as successful answers, the success rate becomes forty‑four out of forty‑four.
The corpus, however, only counts a clarification as a success in one of the forty‑nine cases, because it treats the decision to answer a clarification as a separate behavior rather than a model property. With different tool descriptions the result might improve. The downside of a narrow set is a different kind of failure: when the router finds no matching section the system assumes the question is about knowledge and fabricates an answer without calling any tool. This works for factual questions like “what is OKR”, but fails for interactive queries such as “list the devices on the network”. In that case the model pretends to scan the network, writes a plan and never invokes a tool. The correct response would be to admit that the network scan tool is unavailable and suggest an alternative. To be able to produce such a response the model must first have a way to answer without a tool, otherwise it will simply pretend to work. The author concludes with a practical lesson: when benchmarking, make sure the measurement follows the same flow as the intended product, otherwise you are measuring something else. The routing mode must be attached to the model, not to the system, so switching models also switches the routing configuration. For the three models measured on the full corpus the narrow set was the best configuration. The rule that strong models need a wide set remains a hypothesis that can only be tested on frontier models with a full run, not on a handful of queries. Your own experience may differ, and the author invites discussion.