Topics

Routing

1 article and 3 community signals on Routing for people running AI in production, written and curated by Wren with every source linked.

Updated Oct 10, 2026

Articles

What practitioners are saying

All signals →
  • Hacker News1d ago

    Why isn't the industry freaking out about DeepSeek 4.1 Flash?

    The author reports running all-day coding sessions on DeepSeek 4.1 Flash for under a dollar and argues Chinese labs can undercut frontier pricing the way generic drug makers do. The top replies push back: several commenters say Claude Opus still wins clearly on hard coding tasks, and the gap is worth paying for.

    Why it matters If your routing already sends easy work to a cheap tier, this is the thread to read before you pick which cheap tier.

    Discussion on Hacker News →
  • Hacker News1d ago

    Claude Haiku 5.5: what practitioners say after a day with it

    Commenters report near-perfect accuracy on structured classification in production and call it noticeably smarter than GPT-6 Luna, while several argue the 100K-token price cutoff is too low for long agent runs.

    Why it matters The cheapest tier is only cheap under 100K tokens; check your prompt sizes before you route to it.

    Discussion on Hacker News →
  • Spotify engineering1d ago

    Portal by Spotify cut my Claude Code token usage by 90%

    The author describes routing bulk file reading and boilerplate generation from Claude Code to a cheaper worker model (Gemini 2.5 Flash) via two declarative agents, reporting roughly 90% mean token savings on a Java monorepo in four scenarios; they also list what fails: delegated edits, reasoning, and 10-30 second round trips.

    Why it matters Gives a candid routing recipe with the limits stated, and the savings figure is a single author's test rather than a fleet-wide measurement.