Nine calls became two

I rebuilt the tools an AI agent uses to draft marketing campaigns. Then I ran a controlled benchmark to see whether it helped.

RoleTechnical lead · design, build, evaluation
TimelineSeptember 2026
Models testedDeepSeek V4.1 Flash · Haiku 4.5 · Sonnet 5
61 to 73%less context sent to the model per task (median, three models)
5→2median tool calls per task for Haiku and Sonnet
98.6%Haiku task success, up from 87.5%

The problem

AI agents draft content for several brands. People review, approve and publish it. To save one campaign, the agent had to make nine separate tool calls: list, create, fill, submit, and repeat for every draft.

The model was doing bookkeeping that code should do. Every extra call cost a full round trip of context.

What I changed

  • One call for one intent. “Create this campaign” is now a single call.
  • Code handles the exact work. IDs, versions and retries no longer depend on the model.
  • People keep the final say. Agents draft. People approve and publish.

How I checked, and what I found

I ran the same 24 tasks on the same data, before and after the rebuild. The database decided pass or fail, not a model’s opinion. In total I ran 595 controlled runs.

  • DeepSeek V4.1 Flash: 134.6k to 36.6k tokens per task (73% less). Success 96.6% to 100%.
  • Haiku 4.5: 158.0k to 53.4k tokens (66% less). Success 87.5% to 98.6%.
  • Sonnet 5: 172.1k to 67.9k tokens (61% less). Success 87.5% to 88.9%.

No run took an off-limits action. In a blind test, an AI judge slightly preferred the new writing, but not by a statistically significant margin.

The catch

Sonnet’s success on create tasks fell from 23 of 24 to 19 of 24. In each failed run, it stopped to ask for a missing fact, as the brand guidance tells it to. That is a product decision, and I do not count it as a bug in the rebuild.

Limits

Token counts include cached tokens and host overhead, so less context does not mean the same cut in cost. The test library was small. An AI model judged the writing quality. These are first-party measurements, not an independent audit.

Read the other case study: a promo video with no camera →

Work with me

I can run the same kind of study on your agent, MCP server or AI workflow, then help fix what it finds.