Tool Design Is Prompt Design
Most teams spend days tuning a system prompt and about twenty minutes naming their tools. That ratio is backwards. The model sees your tool names, descriptions, and schemas as text in its context, sitting directly alongside the prompt you agonized over. They are prompt. They just live in a different file.
Names are the first disambiguation
Give a model get_orders and get_orders_recent and it will choose wrong some fraction of the time, because the names don't encode the distinction well enough to decide from.
Name tools by what separates them from their neighbors, not by what they do in isolation. get_all_orders_for_customer and get_last_n_orders_for_customer are longer and unambiguous. Verbosity in a tool name costs a few tokens once. Ambiguity costs a wrong call every time the situation is close.
A common version of this: an agent asked for a customer's latest order calls get_orders, gets the full history, and summarizes the wrong one. Nothing in the names told it which tool returns the recent slice. Renaming the pair and adding a "do not use for" line to each description fixes it more reliably than another paragraph in the system prompt.
This is the tool-selection failure mode from Evals for Production Agents. Test it directly.
Split when the arguments diverge
The tempting move is one flexible tool with a mode parameter. Fewer tools, less context, less to maintain.
It works until the arguments start disagreeing. When mode: "search" needs a query string and mode: "fetch" needs an id, and half the parameters are irrelevant depending on the mode, the model has to reason about which fields apply before it can reason about the task. That is where malformed calls come from.
The rule: split into separate tools when the required arguments diverge. Keep them together when the arguments share a shape and only behavior varies slightly.
Descriptions should say when, and what
Most tool descriptions describe the function. The model already has a decent guess from the name and schema. What it doesn't know is when this tool is the right choice and when it isn't.
Write for the question the model is asking:
That last sentence does more work than the rest of the description. Telling a model what a tool isn't for is one of the cheapest accuracy improvements available.
Declare safety in the same place. Mark each tool with whether it's read-only, whether it's idempotent, and whether it requires approval. The orchestration layer uses those flags to decide what an agent may do unattended, and they belong beside the tool definition. How those flags get used at runtime: Durable Agents.
Schemas: flat, explicit, required
The rules that make structured output reliable make tool arguments reliable, for the same reasons. Flat beats nested, because deeply nested argument objects get filled incorrectly. Enums beat free strings, with an explicit member for unknown rather than an optional field. Required beats optional: some models fill optional fields with a plausible value instead of leaving them out; test yours.
// Before: the model must infer which fields matter from the mode{ name: "orders", description: "Look up orders.", input_schema: { type: "object", properties: { mode: { type: "string" }, // "search" | "fetch" | "summary"? customer: { type: "object" }, // nested, shape unclear query: { type: "string" }, // required for search only id: { type: "string" }, // required for fetch only options: { type: "object" }, // anything goes }, },} // After: one job, flat arguments, nothing ambiguous{ name: "get_recent_orders_for_customer", description: "Returns the 5 most recent orders for one customer. Use when the user asks " + "about recent activity or a specific recent order. Do not use for history " + "over time or aggregate reporting; use get_order_summary_by_period instead.", input_schema: { type: "object", properties: { customer_id: { type: "string", description: "Format: CUS-XXXXX" }, status: { type: "string", enum: ["any", "open", "shipped", "cancelled"], // "any", not optional }, }, required: ["customer_id", "status"], },}The second version is longer in tokens and shorter in failure modes. That is almost always the right trade.
Related reasoning in Structured Output That Holds in Production.
Error messages are prompt too
If you change one thing after reading this, change your error messages.
When a tool fails, whatever it returns goes straight into the model's context and becomes the basis for its next decision. So Error: 500 teaches the model nothing, and it will usually respond by making the identical call again.
Write tool errors as instructions:
No customer found with id "cust_8813".Customer ids have the format CUS-XXXXX.If you have an email address instead, use find_customer_by_email.The model can act on that. It will reformat the id or switch tools, which is what you would want a competent teammate to do. Every tool error is a chance to steer the next call. Most teams waste it on a status code.
Every tool costs context before it does anything
Every tool's name, description, and schema is sent on every call, whether or not the model uses it. A few dozen detailed tools can run to thousands of tokens before the user says anything. That leaves less room for the work and makes selection harder.
Curate per client and per task. An agent doesn't need every tool your organization has. It needs the ones relevant to its job.
Returns count too. A tool returning a full API payload hands the model a wall of JSON, most of it irrelevant and all of it consuming context. Return the fields that inform the next step. If the full object is occasionally needed, expose a second tool that fetches it by id. Truncate long lists explicitly, with a count, so the model knows more exists.
The protocol-level version of this problem: MCP in Production.
Parallel calls change how you write tools
Models now issue several tool calls in a single turn when the calls are independent. That can cut a multi-tool turn from several round trips to one, and it only holds if your tools are independent: no shared mutable state, no ordering assumption buried in a description.
If two tools must run in sequence, say so in the description of the second one. Nothing else will stop the model from firing both at once.
When an agent behaves strangely, read the tool catalog before the system prompt.
Related
Working on something like this?
Start a Conversation