Tools are the bridge between LLMs and the outside world. The LLM decides which tool to call (and with what arguments) based on the tool's description. Good tool design = reliable agent behavior.
The @tool decorator
The simplest way to define a tool:
from langchain_core.tools import tool
@tool
def get_weather(city: str) -> str:
'''Get the current weather for a city.
Args:
city: The city name, e.g., 'Mumbai' or 'New York'.
Returns:
A string describing current weather.
'''
# Real implementation would call weather API
return f"It's 28°C and sunny in {city}."
@tool
def search_products(query: str, category: str = None, limit: int = 5) -> list[dict]:
'''Search the product catalog by keyword.
Args:
query: Search keyword, e.g., 'running shoes'.
category: Optional category filter ('shoes', 'apparel', 'accessories').
limit: Max results to return.
Returns:
List of products with name, price, sku.
'''
return [{"name": "...", "price": 49.99, "sku": "..."}, ...]
Two critical pieces:
1. The docstring IS the description
The LLM reads it to decide WHEN to call the tool. Make it clear:
- WHAT the tool does.
- WHEN to use it.
- WHAT arguments it needs.
- WHAT it returns.
A vague docstring = unreliable tool calling.
2. Type hints define the schema
city: str becomes {"city": {"type": "string"}} in the tool schema. limit: int = 5 becomes optional with default. Use Pydantic models for complex inputs.
Binding tools to an LLM
from langchain_anthropic import ChatAnthropic
llm = ChatAnthropic(model="claude-sonnet-4-6")
tools = [get_weather, search_products]
llm_with_tools = llm.bind_tools(tools)
response = llm_with_tools.invoke("What's the weather in Mumbai?")
# response.tool_calls = [{"name": "get_weather", "args": {"city": "Mumbai"}}]
The LLM looks at the user message, examines the bound tools, decides whether to call one. If yes, it returns a structured tool call instead of text.
Executing the tool call
def call_tool(tool_call):
tool_map = {t.name: t for t in tools}
tool = tool_map[tool_call["name"]]
return tool.invoke(tool_call["args"])
response = llm_with_tools.invoke("What's the weather in Mumbai?")
if response.tool_calls:
for tc in response.tool_calls:
result = call_tool(tc)
print(f"{tc['name']}({tc['args']}) → {result}")
For automated tool execution, use ToolNode (in LangGraph) or build a custom loop. We'll cover the production pattern in Module 2.
Tool design principles
Principle 1 — One tool, one job
# Bad — too many responsibilities
@tool
def manage_order(action: str, order_id: str, **kwargs) -> dict:
'''Get info, refund, cancel, or update an order.'''
if action == "get": ...
# Good — clear, separate tools
@tool
def get_order_status(order_id: str) -> dict:
'''Get current status of an order.'''
@tool
def initiate_refund(order_id: str, reason: str) -> dict:
'''Initiate a refund for an eligible order.'''
@tool
def cancel_order(order_id: str) -> dict:
'''Cancel a pending order.'''
The LLM picks tools by description. Specific tools with clear descriptions = reliable selection. Polymorphic tools = the LLM struggles to fill arguments correctly.
Principle 2 — Verbose descriptions
# Bad
@tool
def get_order(order_id: str) -> dict:
'''Get order.'''
# Good
@tool
def get_order_status(order_id: str) -> dict:
'''Look up the current status, tracking info, carrier, and estimated delivery
of a customer's order. Use this when the customer asks about their order's
location, shipping, delivery, or current state. Requires the order_id from
the customer.
Args:
order_id: The order ID number, e.g., "12345" or "ORD-2026-00045".
Returns:
dict with: status (string), tracking_number, carrier, estimated_delivery_date.
'''
The LLM reads this. Verbose, specific descriptions improve tool selection.
Principle 3 — Bounded and safe by default
# Bad — too much power
@tool
def run_sql(query: str) -> list:
'''Execute SQL query on the database.'''
return db.execute(query) # The LLM can DROP TABLE.
# Good — bounded
@tool
def get_customer_orders(customer_id: str, limit: int = 10) -> list:
'''Get the most recent orders for a specific customer.'''
return db.execute(
"SELECT * FROM orders WHERE customer_id = ? ORDER BY created_at DESC LIMIT ?",
customer_id, limit
)
Don't give the LLM raw database access. Parametrize tools for specific operations.
Principle 4 — Errors return informative messages
@tool
def get_order_status(order_id: str) -> dict:
'''Get order status.'''
try:
order = db.get_order(order_id)
if not order:
return {"error": f"No order found with ID {order_id}. Verify the customer's order number."}
return {"status": order.status, ...}
except Exception as e:
return {"error": f"Database temporarily unavailable: {e}"}
Tools that throw exceptions break the agent. Tools that return error dicts let the LLM recover (e.g., apologize, ask for correct order ID, escalate).
Principle 5 — Pydantic for complex inputs
from pydantic import BaseModel, Field
class TicketCreate(BaseModel):
customer_id: str
category: str = Field(description="One of: complaint, feature_request, billing, technical")
summary: str = Field(description="One-paragraph summary")
priority: str = Field(description="One of: low, medium, high, urgent")
@tool(args_schema=TicketCreate)
def create_support_ticket(**ticket: TicketCreate) -> dict:
'''Create a customer support ticket.'''
# ...
Pydantic models give richer validation and clearer schemas to the LLM.
Tools that take time / cost money
For tools that:
- Take >1 second.
- Cost money per call.
- Have side effects.
Add guard rails:
@tool
def initiate_refund(order_id: str, reason: str) -> dict:
'''Initiate a refund. WARNING: this triggers a real financial transaction.'''
# Check eligibility first
if not is_eligible(order_id):
return {"error": "Not eligible"}
# Require confirmation flag for high-value
order = get_order(order_id)
if order.amount > 1000 and not config.get("auto_approve_high_value"):
return {"requires_approval": True, "amount": order.amount}
return refund_service.create(order_id, reason)
The agent gets a structured response indicating approval needed → can ask the user or escalate, instead of just calling the tool blindly.
Common tool design mistakes
- Vague descriptions. "Gets data" — useless for tool selection.
- Polymorphic tools. One tool with many actions confuses the LLM. Split into specific tools.
- Raw API access. Don't let the LLM run arbitrary SQL or shell commands.
- Tools that crash. Return error dicts instead of throwing.
- Too many tools at once. 10+ tools confuses the model. Group with toolkits, only expose relevant ones per use case.
- No examples in description. Sometimes a one-liner example clarifies argument format.
Toolkits
For collections of related tools, use Toolkits (or just lists):
# A "support agent" toolkit
support_tools = [
get_order_status,
check_refund_eligibility,
initiate_refund,
create_support_ticket,
search_knowledge_base,
]
llm_with_tools = llm.bind_tools(support_tools)
For different agent personas, bind different tool sets. The "support agent" doesn't need the "admin" toolkit.
Takeaway
Tools = LLM's hands in the world. Design with verbose, specific descriptions. One tool, one job. Pydantic for complex inputs. Return error dicts, not exceptions. Bind a focused tool set per agent — don't drown the model in options. The next module covers how agents loop over tools to accomplish goals.
Production Deep Dive: Pydantic v2 Schemas & Argument Validation
When defining tools with @tool, adhere to these three production rules:
- Always provide an explicit docstring: The LLM inspects the function docstring to understand when to choose the tool. Vague docstrings cause tool hallucination and incorrect selection.
- Use Pydantic v2
Field(description=...): Every argument must carry a natural language description, type annotation, and sensible default if optional. - Handle tool exceptions gracefully: Never let an unhandled exception crash the agent loop. Catch database connection drops or API 429 rate limits, and return a clean error string in the
ToolMessageso the model can explain the problem or try an alternative tool.