Change a prompt, a model, or one agent step, and nothing tells you whether it broke the last ten things that worked. So you ship it and watch. axonpush replays the production failures you have already had against every change, and fails the build when one comes back. Evals score the model call. axonpush replays the whole request: the tool that returned a 402, the retrieval that came back empty, the query that timed out. A CLI and a GitHub Action. Exit 1, a JUnit file, a PR comment.
Axonpush is a developer tool that replays production agent failures against code changes to prevent faulty merges in CI environments. It provides analytics on model performance and integrates with GitHub through a CLI and GitHub Action.
Scored deterministically. Only candidates that fire a story trigger are sent to a model, so this one has no written angle.
Gaps in our data, not findings about the product. Their weight is redistributed across the 6 we did measure.
A source that found nothing is a measurement. A source that has not run is a gap. Neither means the launch lacks the thing.