Two things in this week's coverage pull in opposite directions. Anthropic's pricing announcement puts Haiku 5.5 at about 75% below Haiku 4.5, so cheap tiers keep getting cheaper. Anthropic's docs for Managed Agents say a single dynamic workflow run can start up to 1,000 agents, and that every agent's tokens bill at normal rates. Unit prices fall while the number of units one request can trigger goes up.

Whether your bill goes down depends on which effect your workload meets first. We do not know of independent figures for that, and vendor numbers are vendor numbers.

Three angles

  • Unit of cost: tokens are what vendors bill, but finance asks about cost per resolved ticket or per reviewed contract. How do you get from one to the other, and who owns the number?
  • Caps and routing: do you set hard budgets per session or per agent, send easy work to a small model first, and re-measure when prices change? What happened the first time a cap fired?
  • Lock-in: orchestration features, prompt formats and caching rules differ by vendor. What have you kept portable, and what did you knowingly give up for convenience?

What a useful answer looks like

Specifics beat opinions. Name the workload, the model tier, the cap, and what you measured. Say when the measurement was taken, since prices in this area move within weeks. If you are a vendor, say so.

DiscussAnswer one angle below. Partial answers and corrections to anything above are welcome.