PROJECT 16 / 18AGENT EVALUATIONPYTHON / TYPESCRIPT

Evaluation platform prototype

Lucent Eval.

Make an agent’s failure modes inspectable.

6scorer modules
2capture / scoring stages
0–1normalized dimension scale
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

An adversarial evaluation platform with a FastAPI control plane, Celery workers, and a Next.js results interface. It captures agent responses and decomposes their behavior into explicit scoring dimensions with per-prompt rationales.

A single leaderboard number hides the reason an agent failed. Separating raw capture, scoring, and aggregation creates a trail from a surprising result back to the response that produced it.

01

Separated capture and scoring

Runner and scorer queues separate endpoint IO from the evaluation logic. Raw payload storage and per-dimension rationales make individual results inspectable.

02

Explicit score composition

Weight snapshots travel with a run. Missing dimensions are excluded and remaining weights are normalized, rather than silently counting missing values as zero.

03

Multiple inspection surfaces

The repository contains a public leaderboard, run results and trace views, paginated per-prompt results, NDJSON export, and completion webhooks.

02 / UNDER THE SURFACE

From endpoint to evidence

Follow a prompt through capture, independent scorers, and a weighted result. Recovery is a separate scorer, not wired into the single-turn worker.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Submit a run — The API selects active, non-quarantined prompts for a corpus version and creates pending result rows. The run records the supplied endpoint and a snapshot of the scoring weights.
  2. Dispatch prompts — dispatch_run marks the run running, then enqueues a run_prompt task for each result placeholder. Tasks carry run and result identifiers.
  3. Agent endpoint — run_prompt builds a system/user conversation and sends it to the supplied HTTP endpoint. A perf_counter interval records elapsed time at the runner boundary.
  4. Capture response — The worker extracts input/output token usage, estimates cost from a rate table, attempts an S3 upload, and commits captured fields before requesting scoring.
  5. Raw payload store — Raw endpoint JSON is written beneath runs/{run_id}/results/{result_id}.json. Upload failures are logged; the worker continues with a null S3 key.
  6. Scoring task — score_result extracts response text, calls the single-turn scorers, and attaches their rationales. Recovery is explicitly passed as None in this worker path.
  7. Resistance — AdversarialScorer compares the requested behavior and response against refusal and compliance signals. These rules are a transparent heuristic, not an exhaustive attack detector.
  8. Tool misuse — ToolMisuseScorer inspects structured tool calls and returns a score, rationale, and optional call graph that are persisted alongside the result.
  9. Claim grounding — The scorer heuristically extracts claims, excludes temporal statements, and checks overlap with a small fixed fact set. Ungrounded claims receive partial credit.
  10. Latency + cost — Latency and cost are normalized with explicit thresholds in their source modules. Unknown cost receives a neutral 0.5; missing latency receives 0.0.
  11. Recovery rubric — RecoveryScorer implements a multi-turn rubric. The inspected single-turn worker does not invoke it, so its weight must be excluded from that composite.
  12. Weighted composite — CompositeScorer filters missing dimensions, sums only available weights, and divides weighted scores by that active total. It also returns per-dimension contributions.
  13. Result records — Dimension scores, rationales, tool-call graphs, and timestamps are persisted. Finalization waits for all result rows to be scored and then computes run-level means.
  14. Public results — The repository includes a public leaderboard API and a Next.js leaderboard screen. Scores are only interpretable together with corpus version, weights, and underlying traces. The snippet is route notation, not a literal function body.
  15. Completion event — Finalization enqueues completion webhooks. Delivery includes run status and dimension scores, with retries on HTTP failure.
03 / INTERACTIVE STUDY

Build the composite

Adjust an illustrative resistance score and endpoint measurements. Tool misuse and grounding are held fixed; recovery is absent, following the inspected single-turn worker.

CHANGE THE INPUTS

Illustrative inputs, using source scoring bands. No agent endpoint is called and no real agent is being ranked.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Keep measurements interpretable

Latency is measured outside the provider at the runner boundary. It includes HTTP work rather than claiming to isolate model inference time.

02

Preserve partial observability

A missing dimension is represented as None and removed from the active denominator. The single-turn worker deliberately leaves recovery unscored.

03

Expose heuristic limits

The grounding implementation uses a small fixed fact corpus and word overlap; this is an inspectable prototype, not general-purpose factual verification.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • Recovery has a scorer module but is not executed by the inspected single-turn worker.
  • No production deployment, detection accuracy, or comparative leaderboard results are asserted here. The classifier and grounding rules have limited coverage.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.