r/LocalLLaMA1h ago
The poster says they used Claude Opus 5.5 to write a CUDA megakernel that runs a whole speculative-decoding cycle in one launch, reporting 140 vs 73 tok/s on code writing and ~1,600 vs ~1,100 tok/s prefill against llama.cpp with MTP. They list caveats: one model quant (Q4_K_M), RTX 3090 only, and rare rounding differences in output.
Why it matters A self-reported local-inference speed-up with stated limits, and an example of a model being used to write low-level performance code; reproduce before relying on the numbers.
Discussion on r/LocalLLaMA →Hacker News112 points · 211 comments1d ago
The README says OpenAI models burned over $400,000 across months without reaching compatibility, then Claude Opus 5.5 produced a working port in about ten hours and roughly $24,000 over two weeks; all 181,711 ported tests pass and the author states they have never read a line of the code.
Why it matters A real data point on what a large agentic coding run costs, and a reminder that passing tests is not the same as owning the code.
Discussion on Hacker News →Hacker News493 points · 340 comments1d ago
Vals AI reports that Claude Opus 5.5 agents ran quantum-mechanical simulations and surfaced one new compound and one 1999 material as spintronics candidates. The post is careful about caveats: the new compound may be hard to synthesise and the two simulation methods disagree on the old one.
Why it matters A rare agentic-research write-up that states its own limits; useful as a template for how to report agent results internally.
Discussion on Hacker News →Hacker News1043 points · 485 comments1d ago
Commenters report near-perfect accuracy on structured classification in production and call it noticeably smarter than GPT-6 Luna, while several argue the 100K-token price cutoff is too low for long agent runs.
Why it matters The cheapest tier is only cheap under 100K tokens; check your prompt sizes before you route to it.
Discussion on Hacker News →Hamel Husain1d ago
Hamel Husain finds Anthropic's build_eval and hill-climb plugin good at discovering issues other auto-eval approaches miss, especially for handoffs and voice agents, but says it pushes users to write evals before looking at data and bundles too many checks into one evaluator.
Why it matters The most-cited evals practitioner on the web just told you what to do differently with the vendor's tool; read it before you adopt it.
Anthropic customer story1d ago
Anthropic and Zendesk say a five-person team took a custom agent builder from proof of concept to early access in four months on Claude Sonnet through Amazon Bedrock, reached one million agent executions in seven weeks, and that customers saw up to 80% lower handle times and up to 10% higher automated resolution.
Why it matters Vendor-published, so treat the percentages as claims, but the team size and timeline are the useful part for your own plan.
Anthropic customer story1d ago
Anthropic and Pictet say the Swiss bank rolled Claude Code and Cowork out to 700 people, ran 25 workshops for over 500 staff, kept data resident in the EU and Switzerland, and cut a 50-directive gap analysis from two weeks to hours.
Why it matters A regulated-bank rollout with training numbers and residency details is a better template than most enterprise AI announcements.
Spotify engineering1d ago
The author describes routing bulk file reading and boilerplate generation from Claude Code to a cheaper worker model (Gemini 2.5 Flash) via two declarative agents, reporting roughly 90% mean token savings on a Java monorepo in four scenarios; they also list what fails: delegated edits, reasoning, and 10-30 second round trips.
Why it matters Gives a candid routing recipe with the limits stated, and the savings figure is a single author's test rather than a fleet-wide measurement.