This lesson on Tool-Calling / Single-Agent Loop Lite is hands-on and example-driven. You will implement a lightweight, framework-free single-agent loop that binds tool-calling models directly to your retriever. You will enforce strict step caps, token cost budgets, and structured execution logging while validating grounded citation checks before returning final answers.
What You'll Be Able To Do
- Construct a native Python tool-calling loop using direct LLM API tool definitions without external orchestrator dependencies.
- Implement an iterative retrieve-draft-cite-check workflow to verify groundedness before terminating execution.
- Enforce deterministic exit conditions using hard iteration limits and running token spend caps.
- Instrument structured JSON logging across agent iterations, tool arguments, and retriever payloads.
Detailed Concept Walkthrough
1. Single-Agent Loop Architecture
A single-agent loop is a deterministic while-loop that presents tools to an LLM, invokes requested functions, and feeds results back until the model generates a final response or hits a stopping boundary.
- Mechanism: The application initializes a message history list and enters a bounded while-loop where the model evaluates available tools against user intent.
- Execution Flow: If the model returns a
tool_callspayload, the runtime parses tool names and arguments, executes the local retriever function, appends the tool result message to the history, and re-invokes the model. - Best Practice: Keep the tool registry lean by exposing only necessary retriever methods with strict, JSON-schema-validated argument types to minimize routing errors.
def run_agent_loop(query: str, retriever_fn, max_steps: int = 3) -> str:
messages = [{"role": "user", "content": query}]
for step in range(max_steps):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
tools=TOOL_DEFINITIONS,
tool_choice="auto"
)
msg = response.choices[0].message
messages.append(msg)
if not msg.tool_calls:
return msg.content
for call in msg.tool_calls:
result = retriever_fn(**json.loads(call.function.arguments))
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
return "Execution exceeded maximum allowable step limit."
Key Takeaway: Direct while-loops provide full execution transparency and debuggability without the hidden state overhead of third-party agent frameworks.
2. Retrieve-Draft-Cite-Check Verification
Rather than trusting the model to quote facts directly, the agent executes an explicit verification step that cross-references drafted claims against raw retrieved chunks before returning an answer.
- Mechanism: When the model drafts a candidate answer based on tool outputs, a lightweight verification pass evaluates sentence-level claims against the retrieved context IDs.
- Execution Flow: If ungrounded assertions or hallucinated citations are detected, the loop rejects the candidate draft, injects a correction prompt into the message history, and requests a revised response.
- Under the Hood: Verifying citations via programmatic string matching or a low-latency secondary model check prevents fabricated facts from propagating to downstream consumers.
def verify_citations(draft_text: str, retrieved_chunks: list[dict]) -> bool:
# Ensure every cited source identifier matches an actual retrieved document id
cited_ids = set(re.findall(r"\[doc:(\d+)\]", draft_text))
valid_ids = {str(c["id"]) for c in retrieved_chunks}
if not cited_ids or not cited_ids.issubset(valid_ids):
return False
return True
Key Takeaway: Never emit retrieved answers without validating that cited source keys strictly exist within the tool payload.
3. Cost Caps and Safety Guardrails
Autonomous agent loops require explicit operational guardrails to prevent infinite tool-calling recursion, excessive API charges, and run-away context sizes.
- Mechanism: A state tracker accumulates prompt tokens, completion tokens, and dollar expenditures on every cycle through the loop.
- Under the Hood: If cumulative token usage exceeds a hard threshold (e.g., 8,000 tokens) or execution time exceeds a latency cap, the loop breaks immediately and routes to a fallback handler.
- Syntax Rule: Combine explicit loop counters (
for step in range(max_steps)) with dynamic token tracking to guarantee termination under any network or model failure mode.
class BudgetGuard:
def __init__(self, max_tokens: int = 10000, max_cost_usd: float = 0.05):
self.max_tokens = max_tokens
self.max_cost_usd = max_cost_usd
self.total_tokens = 0
self.total_cost = 0.0
def track_usage(self, usage_obj, cost_per_1k: float = 0.00015):
self.total_tokens += usage_obj.total_tokens
self.total_cost += (usage_obj.total_tokens / 1000.0) * cost_per_1k
if self.total_tokens > self.max_tokens or self.total_cost > self.max_cost_usd:
raise BudgetExceededException("Cost safety ceiling reached.")
Key Takeaway: Hard step limits and token budget guards are non-negotiable primitives for production-grade agent loops.
4. Structured Telemetry and Audit Logging
Every decision, tool query, raw retriever payload, and model token count must be logged in a machine-readable schema for offline evaluation and latency profiling.
- Mechanism: The agent wraps every model request and tool invocation with a structured logger that outputs JSON events to stdout or an observability backend.
- Under the Hood: Capturing tool call latency alongside retrieved chunk relevance enables golden-suite regression testing against Module 2 evaluation benchmarks.
- Best Practice: Never log sensitive user PII in tool arguments, but always log query hashes, step indexes, tool execution durations, and final termination reasons.
import structlog
logger = structlog.get_logger()
def log_tool_event(step: int, tool_name: str, args: dict, latency_ms: float):
logger.info(
"agent_tool_execution",
step=step,
tool=tool_name,
arguments=args,
duration_ms=round(latency_ms, 2)
)
Key Takeaway: Structured telemetry converts opaque multi-turn agent execution paths into deterministic, debuggable trace events.
Topics Covered in Tool-Calling / Single-Agent Loop Lite
- Agent Loop Lite Overview (0:00 - 1:15) — Examines the minimal single-agent while-loop architecture without external framework dependencies.
- Tool Schema and Invocation (1:15 - 2:45) — Demonstrates declaring retriever tool schemas and handling tool call requests from model completions.
- Retrieve-Draft-Cite Loop (2:45 - 4:10) — Walks through the iterative cycle of retrieving documents, drafting answers, and validating inline citations.
- Budget Guards & Cost Caps (4:10 - 5:35) — Implements strict execution step counters and cumulative token spend guardrails.
- Structured Audit Telemetry (5:35 - 7:00) — Instruments structured JSON telemetry logging across agent turns to feed evaluation suites.
GenAI + RAG Agents (DS Lite) Cheat Sheet
-
tools = [{"type": "function", "function": {...}}]— Defines callable retriever JSON schema for LLMtools=[{"type":"function","function":{"name":"search","parameters":{"type":"object","properties":{"q":{"type":"string"}}}}}] -
msg.tool_calls— Checks if model requested external tool executionsif response.choices[0].message.tool_calls: handle_tools(response.choices[0].message.tool_calls) -
role: 'tool'— Appends tool execution result back to conversationmessages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(chunks)}) -
for step in range(MAX_STEPS):— Enforces deterministic loop termination via fixed counterfor step in range(3): ... else: raise TimeoutError("Exceeded step limit") -
structlog.get_logger().info()— Emits structured key-value audit event logslogger.info("agent_step", step=step, tokens=resp.usage.total_tokens)
Comparison Table
| Architecture Pattern | State Management | Primary Trade-off |
|---|---|---|
| Direct RAG | Single-turn prompt array | Fastest; cannot re-query missing info |
| Single-Agent Loop Lite | Explicit native Python loop | Full control; lightweight zero-dependency logic |
| Heavy Framework Agent | Graph state machine abstraction | High features; complex debugging overhead |
Common Pitfalls
- Mistake: Allowing unbound while loops without step limits. Avoid: Always iterate over a fixed range ceiling such as range(max_steps).
- Mistake: Appending tool responses without matching tool_call_id keys. Avoid: Strictly pass tool_call_id in the tool message dictionary.
- Mistake: Failing to validate generated citations against returned doc IDs. Avoid: Run an assertion check matching cited IDs to context chunks.
- Mistake: Silently dropping tool execution exceptions inside loop bodies. Avoid: Catch retriever errors and pass the failure message into context.
FAQs
- Why build a native while-loop instead of using LangGraph or CrewAI? A native loop eliminates framework abstractions, giving you complete visibility into state mutations, cost controls, and debugging traces.
- What happens if the retriever returns empty or irrelevant chunks? The agent passes the empty context back to the model, prompting it to either reformulate the search query or report missing information.
- Where in the loop should the citation check be executed? Execute the citation validation immediately after the model returns a final text response without tool calls, before returning to the caller.
- How does this single-agent loop connect to the M2 golden test suite? Every completed loop execution emits structured traces that evaluate precision, recall, and hallucination metrics against your golden eval dataset.