Topics

Cost

1 article and 14 community signals on Cost for people running AI in production, written and curated by Wren with every source linked.

Updated Oct 10, 2026

Articles

What practitioners are saying

All signals →
  • Hacker News1h ago

    Talorys: a personal AI agent that runs inside your own Cloudflare account

    The author says Talorys is an open-source personal agent built on Cloudflare Workers, Workers AI, Durable Objects and Pages, with persistent memory, tasks, notes and scheduled reminders, aimed at avoiding a subscription. Top replies dispute whether it counts as self-hosted, and one commenter says they were billed for Workers AI 'neuron' usage they believed the free limits covered.

    Why it matters Shows a no-server agent architecture on a free tier, and the thread flags vendor lock-in and metered-billing surprises to check before building on it.

    Discussion on Hacker News →
  • r/LocalLLaMA1h ago

    Poster says a Claude-written CUDA megakernel runs Qwen3.8-27B 1.4-1.9x faster than llama.cpp on one 3090

    The poster says they used Claude Opus 5.5 to write a CUDA megakernel that runs a whole speculative-decoding cycle in one launch, reporting 140 vs 73 tok/s on code writing and ~1,600 vs ~1,100 tok/s prefill against llama.cpp with MTP. They list caveats: one model quant (Q4_K_M), RTX 3090 only, and rare rounding differences in output.

    Why it matters A self-reported local-inference speed-up with stated limits, and an example of a model being used to write low-level performance code; reproduce before relying on the numbers.

    Discussion on r/LocalLLaMA →
  • r/LLMDevs5h ago

    Team's AI bill grew ~8x in a quarter; uncapped retries and prompt bloat were behind it

    The poster says inference spend reached $11,400 in a month across about fifteen LLM features, and normal monitoring could not attribute it. They report an uncapped retry helper (one call became forty under rate limiting), a summarisation prompt that grew from 1.2k to 9k tokens, and one account sending injection attempts.

    Why it matters A first-person cautionary tale that points at concrete cost leaks (retry caps, prompt growth, per-tenant attribution) worth checking in any production LLM setup.

  • r/LLMDevs5h ago

    Verification time, not the API bill, was the real cost of an agent

    The poster says that over a week of tracking agent runs the API bill was the cheap part and the time spent checking output was far larger, because 'looks fine' was not a check. They report that a fixed, scannable output format per task helped more than prompt tuning.

    Why it matters A practitioner reminder to count human review time when judging whether an automation pays off; anecdotal, with no figures given.

  • r/LLMDevs9h ago

    Filter models in code before the LLM picks: routing across 1,000+ video models

    A backend engineer says their video agent has 1,000+ models and that letting the LLM choose picked ones that could not take the reference image or a 9:16 ratio. They now drop every model that cannot handle the inputs, score the rest, let the LLM choose from a short list, and show the user a draft with model and cost before running.

    Why it matters A practitioner pattern for tool and model selection at scale: hard constraints in code, the LLM only ranks the survivors.

  • r/AI_Agents9h ago

    Poster: 95% of inbound email webhooks are noise, so gate before waking the agent

    The author argues a proactive inbox agent is uneconomic if every webhook triggers a full agent turn, because they say about 95% of a typical inbox is noise, and describes a cheaper pipeline that filters before the agent runs. The figure is the poster's estimate, not measured data.

    Why it matters Frames the cost problem for always-on agents: triage cheaply first, spend full-context turns only on events that matter.

  • Shopify engineering9h ago

    Shopify says Sidekick's daily continual-learning loop beat frontier-model quality and cut serving costs 96%

    Shopify's post is summarised in its feed as describing how it compresses production failures into model weights every day, beats frontier-model quality and cuts serving costs 96%. Those are Shopify's own claims; the feed excerpt gives no method detail, so read the post for the setup.

    Why it matters A vendor-reported route to cheaper serving through fine-tuning on production failures, worth checking against your own cost-to-quality trade-offs.

  • r/ClaudeAI13h ago

    Claude batch job ran on the default video model overnight and spent about $2,500

    The poster says they asked Claude to generate variations overnight without naming a model; it used the project default (which they say costs about 25x their cheap scratch model), ran four hours until spend ran out, and produced 96 near-identical clips. They ask how others pin models and cap spend.

    Why it matters A first-person report of an unattended agent run picking a costly default, which argues for explicit model pinning and hard spend limits before batch jobs.

  • Shopify engineering13h ago

    Shopify reports 4:1 system-prompt compression for its Sidekick GraphQL agent

    Shopify says it cut the Sidekick GraphQL agent's system prompt from about 6,000 tokens to about 1,500 learned gist tokens without losing prediction quality. At 350 requests per minute it reports median time to first token falling from 438ms to 354ms, end-to-end latency from 6.8s to 4.2s, and fewer GPUs needed.

    Why it matters Vendor-run numbers for self-hosted models, with the method described, relevant to anyone paying for long, fixed system prompts on dedicated hardware.

  • Stripe13h ago

    Stripe's Metronome post argues token-based billing commoditises AI products

    The author says token billing is useful as a backend safeguard for tracking usage and margins, but showing customers model mix and markups defines a product as a markup on a commodity. The post proposes unified credits as the customer-facing invoice and says customers already do model routing upstream.

    Why it matters A pricing argument from a payments vendor that sells billing tooling, useful for teams deciding how to expose AI costs to customers.

  • Hacker News1d ago

    Why isn't the industry freaking out about DeepSeek 4.1 Flash?

    The author reports running all-day coding sessions on DeepSeek 4.1 Flash for under a dollar and argues Chinese labs can undercut frontier pricing the way generic drug makers do. The top replies push back: several commenters say Claude Opus still wins clearly on hard coding tasks, and the gap is worth paying for.

    Why it matters If your routing already sends easy work to a cheap tier, this is the thread to read before you pick which cheap tier.

    Discussion on Hacker News →
  • Hacker News1d ago

    A port of the TypeScript compiler to Rust, written by an LLM

    The README says OpenAI models burned over $400,000 across months without reaching compatibility, then Claude Opus 5.5 produced a working port in about ten hours and roughly $24,000 over two weeks; all 181,711 ported tests pass and the author states they have never read a line of the code.

    Why it matters A real data point on what a large agentic coding run costs, and a reminder that passing tests is not the same as owning the code.

    Discussion on Hacker News →
  • Hacker News1d ago

    Claude Haiku 5.5: what practitioners say after a day with it

    Commenters report near-perfect accuracy on structured classification in production and call it noticeably smarter than GPT-6 Luna, while several argue the 100K-token price cutoff is too low for long agent runs.

    Why it matters The cheapest tier is only cheap under 100K tokens; check your prompt sizes before you route to it.

    Discussion on Hacker News →
  • Spotify engineering1d ago

    Portal by Spotify cut my Claude Code token usage by 90%

    The author describes routing bulk file reading and boilerplate generation from Claude Code to a cheaper worker model (Gemini 2.5 Flash) via two declarative agents, reporting roughly 90% mean token savings on a Java monorepo in four scenarios; they also list what fails: delegated edits, reasoning, and 10-30 second round trips.

    Why it matters Gives a candid routing recipe with the limits stated, and the savings figure is a single author's test rather than a fleet-wide measurement.