EnterTheLoopentertheloop
For Clinicians
Sovereign AIBlogAboutContact
EnterTheLoopentertheloopClinicians powering AI alignment, training & safety.

Verified against

GMCNMCGPhCHCPC

Follow

Register→
© 2026 EnterTheLoop Ltd  ·  Built in Britain
PrivacyTermsCookies
EnterTheLoopentertheloop

Clinicians powering AI alignment, training & safety.

PrivacyTermsCookies
© 2026 EnterTheLoop Ltd · Built in Britain

Hold a report? It lives in your dashboard.Sign in

Calibration · Reliability reports

The Reliability Report.

A Reliability Report records how a clinician performed across the eight task types, with uncertainty on every figure rather than a bare number. The calibration that issued them closed in September 2026. Reports already issued still stand, and buyers can still verify them.

The daily loopThe methodology

At a glance

Status
No longer issued
Reports issued
Still valid
Format
PDF and JSON
Calibration now
Daily loop
Why it matters

The report is the record.

A Reliability Report is a procurement-grade record of how one clinician performed: a reliability score with confidence intervals on every metric, in a form a buyer can check rather than take on trust. The reports issued under the full calibration are unchanged by its closure.

Coverage

All eight task types.

Evaluation work spans more than rating answers. The calibration exercised the full range, so a report records performance across the work buyers actually need. The daily loop mixes the same eight kinds of task.

01

Rating

Score a single AI response against a quality scale.

02

Comparison

Choose the better of two responses, and say why.

03

Ranking

Order several responses from best to worst.

04

Rubric

Assess against structured, dimensional criteria.

05

Correction

Fix an unsafe or incorrect response.

06

Annotation

Label spans and structure in clinical text.

07

Justification

Write the clinical reasoning behind a judgement.

08

Red-team

Probe for failure modes across the safety taxonomy.

The statistics

Measured, not asserted.

Every metric in a report carries its uncertainty rather than standing as a bare number, which is why a report keeps meaning what it says long after it was issued.

Per-metric confidence intervals

Beta-Binomial intervals for proportions, bootstrap intervals for continuous metrics, so uncertainty is quantified on every score.

Lower confidence bounds

The report read the lower bound rather than the point estimate, so a score reflects worst-case plausible performance.

Per-category coverage

Performance broken down across task types and the safety taxonomy, recorded in the report.

If you hold one

A report already issued still stands.

A Reliability Report issued under the full calibration is yours, as a PDF for people and JSON for systems. It stays on your profile, it is not withdrawn and not reissued, and a buyer you share it with can still verify it.

Sign in to open your report →How it was scored
Calibration now

A short daily loop.

Calibration runs as a set of five cases, most days, on cases written for your own profession. A set takes about four minutes. There is no screening step in front of it and nothing to unlock.

How the daily loop works →
WHAT CHANGED

The two-phase calibration, Quick Screen then Full Calibration, was retired in September 2026. Calibration now runs as a short daily loop. A signed Reliability Report issued under the old pathway still stands, and buyers can still verify it.

Get calibrated.

Registration is free. Calibration runs as a short daily loop: five cases, most days, on cases written for your profession.

Register freeBack to training
EnterTheLoopentertheloopClinicians powering AI alignment, training & safety.

Verified against

GMCNMCGPhCHCPC

Follow

Register→
© 2026 EnterTheLoop Ltd  ·  Built in Britain
PrivacyTermsCookies
EnterTheLoopentertheloop

Clinicians powering AI alignment, training & safety.

PrivacyTermsCookies
© 2026 EnterTheLoop Ltd · Built in Britain