r/LocalLLaMA1h ago
The poster says they used Claude Opus 5.5 to write a CUDA megakernel that runs a whole speculative-decoding cycle in one launch, reporting 140 vs 73 tok/s on code writing and ~1,600 vs ~1,100 tok/s prefill against llama.cpp with MTP. They list caveats: one model quant (Q4_K_M), RTX 3090 only, and rare rounding differences in output.
Why it matters A self-reported local-inference speed-up with stated limits, and an example of a model being used to write low-level performance code; reproduce before relying on the numbers.
Discussion on r/LocalLLaMA →r/LocalLLaMA1h ago
A self-described 20-year software engineer says they ran both models at the same Unsloth Q4 quantisation through Codex CLI and OpenCode. They report equally good repository analysis, with Gemma using about 10 server calls against Qwen's 20-30, and a clearer gap on a new-project task, where they graded Gemma's first result B-.
Why it matters A single-person, informal comparison, but it frames the right axes for choosing a local coding model: calls per task, iterations to a usable result, and harness effect.
r/LLMDevs5h ago
The poster says that in 10 series of 10 runs with a quantized Qwen3.5-35B-A3B on a 13-test CLI task, the first run (no retrieved notes) passed in 1 of 10 series versus 54 of 90 later runs (60%), with wide variance between series. They built the tool being tested and ask for methodology critique.
Why it matters An early, self-run experiment on whether agents can reuse knowledge across attempts; useful as a test design to copy, not as evidence the approach works.
Shopify engineering13h ago
Shopify says it cut the Sidekick GraphQL agent's system prompt from about 6,000 tokens to about 1,500 learned gist tokens without losing prediction quality. At 350 requests per minute it reports median time to first token falling from 438ms to 354ms, end-to-end latency from 6.8s to 4.2s, and fewer GPUs needed.
Why it matters Vendor-run numbers for self-hosted models, with the method described, relevant to anyone paying for long, fixed system prompts on dedicated hardware.
Hacker News1067 points · 936 comments1d ago
The author reports running all-day coding sessions on DeepSeek 4.1 Flash for under a dollar and argues Chinese labs can undercut frontier pricing the way generic drug makers do. The top replies push back: several commenters say Claude Opus still wins clearly on hard coding tasks, and the gap is worth paying for.
Why it matters If your routing already sends easy work to a cheap tier, this is the thread to read before you pick which cheap tier.
Discussion on Hacker News →