The problem
AI agents draft content for several brands. People review, approve and publish it. To save one campaign, the agent had to make nine separate tool calls: list, create, fill, submit, and repeat for every draft.
The model was doing bookkeeping that code should do. Every extra call cost a full round trip of context.
What I changed
- One call for one intent. “Create this campaign” is now a single call.
- Code handles the exact work. IDs, versions and retries no longer depend on the model.
- People keep the final say. Agents draft. People approve and publish.
How I checked, and what I found
I ran the same 24 tasks on the same data, before and after the rebuild. The database decided pass or fail, not a model’s opinion. In total I ran 595 controlled runs.
- DeepSeek V4.1 Flash: 134.6k to 36.6k tokens per task (73% less). Success 96.6% to 100%.
- Haiku 4.5: 158.0k to 53.4k tokens (66% less). Success 87.5% to 98.6%.
- Sonnet 5: 172.1k to 67.9k tokens (61% less). Success 87.5% to 88.9%.
No run took an off-limits action. In a blind test, an AI judge slightly preferred the new writing, but not by a statistically significant margin.
The catch
Sonnet’s success on create tasks fell from 23 of 24 to 19 of 24. In each failed run, it stopped to ask for a missing fact, as the brand guidance tells it to. That is a product decision, and I do not count it as a bug in the rebuild.
Limits
Token counts include cached tokens and host overhead, so less context does not mean the same cut in cost. The test library was small. An AI model judged the writing quality. These are first-party measurements, not an independent audit.