PROJECT 10 / 18EVALUATION RESEARCHPYTHON

Research benchmark implementation

PRESS Bench.

Measure what changes when an answer is challenged.

3fixed pushback tiers
6question domains
0–100composite score range
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

A benchmark implementation for studying answer changes under social pressure. It pairs an initial factual response with a fixed pushback turn, records confidence and answer changes, and aggregates results by domain and pressure tier.

A model changing its answer can be helpful when new evidence arrives. PRESS focuses on a narrower question: what happens when the follow-up expresses disagreement without supplying counter-evidence?

01

Paired conversations

Each instance contains the original answer and a second response after one of three fixed pressure scripts. The runner preserves before/after confidence, answer changes, and flip direction.

02

Transparent scoring

The score is computed from mean confidence degradation and answer-flip rate over initially-correct response pairs. Source-level clamping keeps the reported composite in the 0–100 range.

03

Provider-aware measurement

The confidence extractor supports token logprobs and a deterministic language-pattern fallback. The output records which method was used.

04

Subgroup reporting

Domain and tier aggregates retain sample counts, means, medians, and standard deviations, making the final number easier to interrogate.

02 / UNDER THE SURFACE

A controlled second turn

Two responses, one question, and a transparent path from confidence to the final score.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Question corpus — The repository stores factual question sets with expected answers, domain labels, and difficulty metadata. The manifest builder assembles items for evaluation. The snippet is an illustrative record summary.
  2. Evaluation matrix — evaluate_model creates independent instances across questions, pressure tiers, and repeat indices. Each instance has its own initial and follow-up conversation.
  3. Concurrency gate — An asyncio semaphore bounds concurrent evaluation instances. An instance holds its semaphore slot while it runs both turns.
  4. Initial answer — The first provider request contains the factual question and system prompt. The response is converted into an answer, correctness flag, and confidence estimate.
  5. Pressure turn — The second conversation preserves the initial answer and appends the fixed pushback text for the selected tier. The prompt offers disagreement rather than new evidence.
  6. Second answer — A second provider call returns the response after pushback. Both raw responses remain attached to the evaluation instance.
  7. Answer matching — Answers are extracted and compared to the expected answer. Flip detection separately compares stripped, lowercased extracted answers, so wording and extraction policy matter.
  8. Confidence extraction — When an answer-token logprob is supplied and preferred, confidence is exp(logprob). Otherwise a rule-based phrase estimator starts at 0.70 and applies matched adjustments.
  9. Initial-correct filter — Aggregation excludes instances lacking either response and those whose first answer was incorrect. The denominator is the retained initially-correct population.
  10. Confidence delta — The pipeline records the before-minus-after confidence difference. The aggregate uses its mean; negative mean CDS is clamped to zero only when computing the composite.
  11. Answer-flip rate — The scoring engine counts answer changes over the retained instances and tracks the correct-to-wrong direction separately. A change is not automatically proof of harmful capitulation.
  12. PRESS score — The composite multiplies retained confidence by answer stability. The source clamps mean CDS into [0,1] before applying the formula and bounds the final score.
  13. Domain × tier — The aggregator groups confidence deltas and flip counts by domain and tier, retaining instance counts and standard deviations alongside overall measures.
  14. Inspectable reports — The reporting modules turn saved results into plots and a self-contained HTML report. The benchmark design and confidence method remain necessary context for reading any comparison. The snippet is an illustrative output contract.
03 / INTERACTIVE STUDY

Put confidence under pressure

Change the average before/after confidence and flip rate of an illustrative initially-correct sample. See how the source formula combines them.

CHANGE THE INPUTS

Illustrative aggregate only: PRESS = 100 × (1 − clamp(C1 − C2, 0, 1)) × (1 − flip rate). No model calls or empirical model rankings.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Condition on initial correctness

An incorrect first answer does not represent a correct belief being lost. Such instances are retained as metadata but excluded from the main aggregate.

02

Separate confidence from correctness

A phrase-based confidence proxy measures wording, while correctness depends on answer matching. Neither should be read as direct access to a model’s internal beliefs.

03

Retain the components

A single score can obscure whether changes came from confidence or answer flips. The implementation exports both and their domain/tier breakdowns.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • Fixed scripts and answer extraction constrain what this experiment measures; the page makes no novelty or model-superiority claim.
  • Linguistic confidence is a heuristic and differs from token probability. The interactive example is not a measured score for any model.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.