Separated capture and scoring
Runner and scorer queues separate endpoint IO from the evaluation logic. Raw payload storage and per-dimension rationales make individual results inspectable.
Evaluation platform prototype
Make an agent’s failure modes inspectable.
An adversarial evaluation platform with a FastAPI control plane, Celery workers, and a Next.js results interface. It captures agent responses and decomposes their behavior into explicit scoring dimensions with per-prompt rationales.
A single leaderboard number hides the reason an agent failed. Separating raw capture, scoring, and aggregation creates a trail from a surprising result back to the response that produced it.
Runner and scorer queues separate endpoint IO from the evaluation logic. Raw payload storage and per-dimension rationales make individual results inspectable.
Weight snapshots travel with a run. Missing dimensions are excluded and remaining weights are normalized, rather than silently counting missing values as zero.
The repository contains a public leaderboard, run results and trace views, paginated per-prompt results, NDJSON export, and completion webhooks.
Follow a prompt through capture, independent scorers, and a weighted result. Recovery is a separate scorer, not wired into the single-turn worker.
Adjust an illustrative resistance score and endpoint measurements. Tool misuse and grounding are held fixed; recovery is absent, following the inspected single-turn worker.
Illustrative inputs, using source scoring bands. No agent endpoint is called and no real agent is being ranked.
Latency is measured outside the provider at the runner boundary. It includes HTTP work rather than claiming to isolate model inference time.
A missing dimension is represented as None and removed from the active denominator. The single-turn worker deliberately leaves recovery unscored.
The grounding implementation uses a small fixed fact corpus and word overlap; this is an inspectable prototype, not general-purpose factual verification.
Implementation details, examples, and project documentation.
Inspected task dispatch, HTTP capture, raw payload storage, score persistence, and finalization.
Inspected exclusion of missing dimensions and active-weight normalization.
Inspected temporal exclusions, fixed facts, and partial credit logic.
Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.