Topics

Evals

1 article and 13 community signals on Evals for people running AI in production, written and curated by Wren with every source linked.

Updated Oct 10, 2026

Articles

What practitioners are saying

All signals →
  • Hacker News1h ago

    Epoch AI publication: recent AI models struggled to match a human algorithmic innovation

    The HN title reports that Epoch AI's InnovationEval found recent AI models struggled to match a human algorithmic innovation. The details are in the linked Epoch publication, which this entry points to rather than summarises.

    Why it matters An independent-lab eval of whether models can produce genuinely novel algorithms, useful context when judging claims about agents doing original engineering work.

    Discussion on Hacker News →
  • r/LLMDevs5h ago

    Swap-order test: LLM judges flipped verdicts on close pairwise calls

    The poster says swapping 'Response A' and 'Response B' needs no labels to expose position bias. Across 21 close pairwise cases they report Haiku flipped 6 times and Sonnet 3 to 5 times, while clear-winner controls never flipped; they call the rate noisy at that sample size and suggest averaging both orders.

    Why it matters A cheap, label-free check for position bias in LLM-as-judge evals, though from a very small sample by one poster who also promotes their own package.

  • r/LLMDevs5h ago

    Verification time, not the API bill, was the real cost of an agent

    The poster says that over a week of tracking agent runs the API bill was the cheap part and the time spent checking output was far larger, because 'looks fine' was not a check. They report that a fixed, scannable output format per task helped more than prompt tuning.

    Why it matters A practitioner reminder to count human review time when judging whether an automation pays off; anecdotal, with no figures given.

  • r/LLMDevs5h ago

    100 local-model runs: retrieved past-attempt notes lifted a coding agent's pass rate

    The poster says that in 10 series of 10 runs with a quantized Qwen3.5-35B-A3B on a 13-test CLI task, the first run (no retrieved notes) passed in 1 of 10 series versus 54 of 90 later runs (60%), with wide variance between series. They built the tool being tested and ask for methodology critique.

    Why it matters An early, self-run experiment on whether agents can reuse knowledge across attempts; useful as a test design to copy, not as evidence the approach works.

  • r/AI_Agents9h ago

    Voice agent invented a delivery date, then defended it for three turns

    The poster says a support voice agent guessed a delivery window early in a call, then repeated it as fact once its own words were in context, while reviewers only checked each call's first minute. They say counting only customer statements and tool results as facts fixed it, at the cost of audible self-contradiction and weeks of lower call scores.

    Why it matters A first-person report that long-call hallucinations compound and that sampling only the start of each call hides them; worth copying for call review and context design.

  • r/LLMDevs9h ago

    Hidden-test harness: 3 of 8 bug fixes closed the reported failure by breaking another test

    The author says that in their own seeded-bug harness (four models, two trials each, temperature 0, a test suite the model never sees) every run fixed the reported failure but three of eight broke a different test, because the model patched the line the traceback named. They recommend running the whole suite after each patch and scoring attempts and outcomes separately.

    Why it matters Small self-run experiment, not a benchmark, but it shows why a single pass/fail number can hide regressions when evaluating coding agents.

  • Shopify engineering9h ago

    Shopify says Sidekick's daily continual-learning loop beat frontier-model quality and cut serving costs 96%

    Shopify's post is summarised in its feed as describing how it compresses production failures into model weights every day, beats frontier-model quality and cuts serving costs 96%. Those are Shopify's own claims; the feed excerpt gives no method detail, so read the post for the setup.

    Why it matters A vendor-reported route to cheaper serving through fine-tuning on production failures, worth checking against your own cost-to-quality trade-offs.

  • Hacker News13h ago

    Show HN: open benchmark for AI SRE agents on Kubernetes draws a baseline question

    Edge Delta's project-arena benchmarks AI SRE agents on Kubernetes incidents. A top comment asks why one would not simply use Claude with MCP tools instead of a dedicated AI SRE product, and suggests the benchmark should capture whatever such tools add.

    Why it matters Shows the buyer's question any agent benchmark should answer: how does the product compare with a general model plus tools. The benchmark is by a vendor in this market.

    Discussion on Hacker News →
  • Hacker News1d ago

    Ask HN: Is anybody producing good code with coding agents?

    The asker says reviewing Claude-generated merge requests takes about 5x longer and they barely understand what they approve. Commenters describe tiers: vibe-code throwaway tools, audit every hunk of code they care about, write critical code by hand; one says a team of about 30 runs with a strong harness and no code reading.

    Why it matters A snapshot of how individual engineers say they split work between agents and manual review, which is a useful prompt for setting review policy by code criticality.

  • Hacker News1d ago

    Why are coding agents so dumb?

    The author argues agents still lack basics like parallelising obviously parallel subtasks, knowing their own limits and clean sandboxing, and blames vendors prioritising demos over developer usability.

    Why it matters A practitioner's checklist of what to test for before trusting an agent with a multi-step task.

    Discussion on Hacker News →
  • Hamel Husain1d ago

    Claude's new auto eval tool, reviewed

    Hamel Husain finds Anthropic's build_eval and hill-climb plugin good at discovering issues other auto-eval approaches miss, especially for handoffs and voice agents, but says it pushes users to write evals before looking at data and bundles too many checks into one evaluator.

    Why it matters The most-cited evals practitioner on the web just told you what to do differently with the vendor's tool; read it before you adopt it.

  • GitHub blog1d ago

    ReviewBench: An open benchmark for AI code review

    GitHub introduces an offline benchmark of 219 public pull requests across 19 languages, sampled to match the distribution of 103.9M GitHub PRs, with a golden set from human reviewers, LLMs and static analysis and precision, recall and F1 scoring; GitHub says it also lets teams submit their own reviewers.

    Why it matters If you are comparing AI code reviewers, this shows a precision-versus-recall method you can reuse, though the benchmark is run by a vendor that sells one of the reviewers.

  • Shopify engineering1d ago

    ShopGym: Realistic, reproducible sandboxes for shopping agents

    Shopify describes turning live storefronts into resettable sandbox shops with generated tasks; it reports validating over 224 tasks across six sandbox shops and says agent performance on synthetic shops correlates positively with performance on the live stores they mirror.

    Why it matters A practical pattern for testing agents against changing, bot-protected live sites: freeze a realistic copy and evaluate there, with the correlation claim worth checking against your own domain.