Standards for building agents, better
-
Updated
Jun 3, 2026 - TypeScript
Standards for building agents, better
Agentic testing for agentic codebases
The definitive benchmark for AI agents on OpenClaw. 45 tasks across 4 tiers. Powered by MyClaw.ai
AI agents can generate code, but they still struggle to understand what they build. Reticle gives them runtime perception of web applications.
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments and CI gates.
Ship agents you can audit.
Agent Verifier is a coding agent skill that verifies code against organizational policies, code quality patterns, security requirements, and framework best practices — before code ships. Works with Claude Code, Cursor, Windsurf, and 30+ agents.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
The Regression Testing Framework for AI Agents. Replay · Evaluate · Assert · Catch Regressions — in CI. Like Jest for your AI layer.
infrastructure chaos to test the resilience of ai agents
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
Turn failed AI agent runs into replayable regression tests. Catch regressions before you ship.
Typed Kotlin DSL framework for AI agent systems.
GitHub template for agent-testable SaaS apps. Next.js 16 + shadcn/ui + Neon Postgres + agent-browser e2e testing via accessibility tree.
Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you swap models or edit prompts.
Deterministic runtime for agent evaluation
pytest plugin for deterministic testing of AI agents. Assert agent actions, not vibes.
Behavior-regression testing for LLM agents. 4-class attribution, 6-field FAIL schema, $-cost gating, flaky detection. Bash + jq. Works with opencode today, runner-pluggable.
Record-and-replay for agent decision graphs: reproduce a prod agent failure as a committed regression test, and re-run your fix without live LLM calls.
A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.
Add a description, image, and links to the agent-testing topic page so that developers can more easily learn about it.
To associate your repository with the agent-testing topic, visit your repo's landing page and select "manage topics."