Coding-Agent Benchmarking with SWE-Bench Regression Gates
Agent quality is a moving target — a model update can shift pass rates overnight. This LangGraph workflow builds a golden task suite, runs agents in parallel Map-Reduce style, scores pass rates and cost per task, diffs against a pinned baseline, and ends in a hard go/no-go promotion gate.