7 min read
How the KV cache shapes AI agent design
Prompt caching only works when each request begins with the same tokens as the last. The design rules that follow for AI agents, and how to check them.
7 min read
Prompt caching only works when each request begins with the same tokens as the last. The design rules that follow for AI agents, and how to check them.
5 min read
An agent re-sends its whole conversation every turn. What that does to a bill, what a real project's numbers showed, and the five habits that changed it.
8 min read
AI coding tools are context engines. How the harness works, the practices that improve what it produces, and what to settle before a team adopts them.
6 min read
The mental model, the five levers that actually work, and the failure modes to watch for when you put an LLM into your daily workflow.