Agents look magical in demos and fragile in production. The difference is almost always the plumbing around the model, not the model itself.
Treat every tool as an API contract: strict input schemas, typed outputs and clear error messages the model can read and recover from.
Keep context small and intentional. Summarise finished steps, drop raw tool output once it is used, and store long-term state outside the prompt.
Finally, put hard limits everywhere: max steps, max cost, timeouts and a human fallback. Reliable agents are bounded agents.
