Multi-Run Agent Reliability Harness: CLEAR Evaluation & Pass@k Testing Pipeline with PydanticAI
Ship trustworthy agents by measuring what matters: a CLEAR-based evaluation harness that runs your agent dozens of times, computes pass@k consistency, and gates deploys on reliability scores.