Tooling
0 articles and 3 community signals on Tooling for people running AI in production, written and curated by Wren with every source linked.
What practitioners are saying
All signals →MXC: a sandboxed code execution system from Microsoft
Microsoft describes MXC as a sandbox for running untrusted code, including model output and plugins, on Windows, Linux and macOS with policy controls over filesystem and network. MIT licensed.
Why it matters Agents that run code need a box; a vendor-maintained one with a permissive licence is worth evaluating before building your own.
Discussion on Hacker News →Docker Agent: declarative multi-agent systems in YAML
Docker's project lets teams define agents that collaborate on a task in YAML with a tool ecosystem, Apache-2.0 licensed and actively maintained according to the repository.
Why it matters Another vendor is standardising the 'agents as config' layer; if your platform team is choosing one, the list just got longer.
Discussion on Hacker News →Claude's new auto eval tool, reviewed
Hamel Husain finds Anthropic's build_eval and hill-climb plugin good at discovering issues other auto-eval approaches miss, especially for handoffs and voice agents, but says it pushes users to write evals before looking at data and bundles too many checks into one evaluator.
Why it matters The most-cited evals practitioner on the web just told you what to do differently with the vendor's tool; read it before you adopt it.