URRSF 3.1_Confidential

URRSF 3.1 — Offline Rule-Based Grading Rubric (9 domains)
RATSe → RATSe-Radiology → RRGF → Layer 5 · Report Governability Assessment · Manual Instrument

URRSF 3.1 Grading Rubric — Rule-Based, Offline Edition

A fully transparent scoring instrument for grading radiology reports by hand. Every weight, anchor, gate, and line of arithmetic is printed on this page and computed in front of you. No model, no network, no hidden multiplier.

Works offline · single file All weights disclosed Nine graded domains EAL C — human-judged Two-tier gates active Educational & research use

Teacher's guide & rationale

Read me first · open to print

This instrument grades a radiology report the way a careful attending would: it separates quality you can trade off from failures you cannot. Everything it does is visible on this page — the weights, the rules, the arithmetic. It is an educational and formative aid, not a certified competency exam.

How to run a grading session (8 steps)
  1. Enter the context and report (§1). Paste the de-identified report, fill in the clinical indication, and set the artifact origin — several checks depend on knowing what the study was ordered to answer and whether AI was involved. Never enter patient names, MRNs, or dates.
  2. Read the anchors (§2). The 1–5 scale is behaviourally anchored: 3 is the acceptable minimum, not a polite average. Calibrate to those descriptions, not to your mood.
  3. Pick the scenario (§3). The scenario changes the domain weights, switches individual criteria on or off, and decides which governance flags are Tier 1 rather than Tier 2. It is a reporting mode, not a specialty: choose it by what the study is for. Surveillance & Interval Comparison assumes a non-acute study read against a prior — it de-emphasises urgency and does not promote the urgent-communication gate, so it must not be used for an acute presentation, including an acute presentation in an oncology patient. The full table is shown; nothing is hidden.
  4. Run the offline analysis (§6). The rule engine pre-fills the machine-checkable criteria and flags gates, each with the exact sentence that triggered it. It is a first pass, not a verdict.
  5. Review every engine suggestion. The engine is a linter, not an oracle. Confirm or override each score and flag. Items tagged engine unsure — confirm especially need your eye.
  6. Score the diagnostic criteria yourself. Diagnostic Accuracy is marked awaits human — the engine cannot see the images and refuses to guess. You must score these before the session can be finalised.
  7. Check the gates (§5). Confirm any non-compensatory gate. A Tier-1 gate fails the report outright; a Tier-2 gate caps it. Use the quoted evidence to justify each.
  8. Sign off and export. Switch to the Results & Summary tab. The sign-off box only unlocks once every awaiting criterion is scored. Then Print/PDF or export the record, which carries the evidence quotes and the knowledge-base version.
The core idea: compensatory domains vs. non-compensatory gates

Part A — graded domains (compensatory). Nine domains scored 1–5 and weighted into a 0–100 profile. Strength in one can offset weakness in another, the way overall report quality really does trade off.

Part B — governance gates (non-compensatory). Some failures cannot be bought back by excellence elsewhere. A report that names the wrong patient, or sends a BI-RADS 4 lesion to routine screening, is not “82% good” — it is a failed report. Gates model that. This is the single most important design decision in the tool: a weighted average alone is unsafe, because it lets a beautifully-written report average away a catastrophic error.

Two tiers. Tier 1 = never-events → verdict locks to FAIL (score 0). Tier 2 = major deficiencies → score capped at 59 (a failing band) but not zeroed. This restores proportionality: a laterality error and a missing follow-up sentence are both serious, but not equally catastrophic.

The verdict is a word, not just a number. The headline is ACCEPTABLE / BELOW STANDARD / CONDITIONAL / FAIL, with a per-domain profile beneath it. A single 0–100 score invites misuse as a ranking and implies a precision the instrument does not have; the profile shows where a report is strong or weak, which is what actually helps a trainee.

What changed in 3.1 — the two restored domains

Earlier versions scored seven domains. Two were missing, and both matter enough to be graded rather than left to the gates alone.

Machine & Human Accountability (9%). Previously handled only by AI-governance gates, which are binary: either undisclosed AI is a never-event or it is nothing. But accountability has degrees — a report can name its author yet carry no timestamp, or disclose AI involvement without documenting who verified it. Those are gradable deficiencies, not never-events. Making it a domain also means a report is measured on provenance even when nothing has gone wrong enough to gate.

Patient-Centered Access (4%). Previously absent from the instrument entirely — neither a domain nor a gate. With open-notes access, patients read these reports directly and often before their clinician explains them. A report that is impeccable for a referrer and incomprehensible to the person it is about has a real deficiency the old rubric could not express. It carries the smallest weight of the nine, which is the honest reflection of its priority relative to getting the diagnosis right.

Note The remaining seven weights were rescaled proportionally so the total stays 100 and the relative ranking among them is unchanged.

Why the engine abstains — and what “EAL C” means

The honest boundary. The offline engine is a rule-based text linter with a clinical knowledge base — not an AI that judges diagnoses. It reliably checks things that live in the words: laterality mismatches, internal contradictions, phantom organs, hedging, measurements, provenance markers, spelling. It is negation-aware, so “no pneumothorax” is read as absent. What it cannot do is judge whether a diagnosis is correct, because that needs the images. So it abstains on Diagnostic Accuracy and hands those criteria to you.

Evidence Assurance Level (EAL). Every domain here is EAL C: a judgment made and owned by a human evaluator, not machine-verified. The label is a promise about who is accountable for the score — you are. A handful of criteria carry a rule-verifiable badge because the engine checks them deterministically; those approach EAL B.

Artifact ≠ event. The tool grades the report artifact. It cannot confirm that an urgent finding was actually phoned to the care team, only that the report says so. Keep that distinction when you grade communication and accountability.

Fairness, limitations, and honest caveats

Writing style is not a defect. Spelling and grammar are shown as advisory only and never lower a score. A different-but-valid register — including non-native-English phrasing or a non-US template house style — is not a quality problem, and the tool deliberately does not penalise it.

The engine can miss things. Rule-based checks catch the phrasings they anticipate. An unusually-worded error can slip past. A clean engine result means “nothing the rules recognised,” not “a good report.” Your review is what makes the grade valid.

The weights are consensus, not validated. The Balanced column is literature-informed; the three scenario columns are reasoned adaptations of it. None has been through a Delphi process, an inter-rater reliability study, or correlation with patient outcomes. Treat the number as formative.

Calibrate your graders. Before grading a real stack, have two evaluators score the same 2–3 anchor reports independently and reconcile. The knowledge-base version stamped on every export tells you which instrument produced which score.

Not a diagnostic device. This is a teaching and quality-reflection aid. It does not replace radiologist judgment, departmental QA, or formal assessment.

1 · Grading session & report

Faculty enters everything here first
Fill in the context and paste the report, set the scenario below, then run the offline analyzer. It reads the words on the page — it cannot see the images.

2 · Behavioural anchors

Apply per criterion
N/ACriterion does not apply to this study. Removed from the denominator — it neither helps nor harms.
1Unsatisfactory. Major deficiency; immediate correction required before the report could be relied on.
2Significant deficiency that reduces the report's clinical usefulness.
3Acceptable minimum standard. Safe, but no better than required.
4Good practice; only minor opportunities for improvement remain.
5Exemplary. A model report suitable for teaching.

3 · Scenario weights — disclosed

The Balanced column is the literature-informed reference set; the three scenario columns are reasoned adaptations of it, not independently validated. Weights only shape the compensatory composite — they can never rescue a gated report. Domains marked N/A in full are removed from both numerator and denominator, so weights renormalise automatically.

4 · Part A — Graded domains (compensatory)

Nine domains · score every criterion

5 · Part B — Governance gates (non-compensatory)

Check any that apply
Tier 1 · Never events — verdict locks to FAIL (0)

Rule: if any Tier-1 gate is checked, the reported score is 0 and the verdict is FAIL — Safety Gate, regardless of Part A. The underlying Part-A quality score is preserved below the verdict for remediation planning only.

Tier 2 · Major deficiencies — score capped at 59

Rule: if any Tier-2 gate is checked (and no Tier-1 gate), the composite cannot exceed 59 — the ceiling of the “Unsatisfactory / Major Deficiencies” band. Part-A excellence cannot buy these back above a failing band, but the report is not treated as a never event.

6 · Run offline analysis

No model · no network · rule-based

This runs a set of deterministic text rules (a linter with a clinical knowledge base — not an AI that judges diagnoses) over the report you entered: laterality, contradictions, phantom organs, hedging, measurements, provenance markers, patient-facing language, and spelling. It is negation-aware and tells you when it cannot reliably segment a report. It pre-fills the scores and gates — each with the exact sentence that triggered it — and hands the diagnostic criteria back to you, marked awaits human.

7 · Engine assessment

Rule-based text checks · auditable

Scope: a rule-based text linter, not an AI that judges diagnoses. It checks the report artifact for internal consistency, completeness, provenance and clarity. It cannot verify that a finding was actually communicated to the care team, nor whether a diagnosis is correct — both need information outside the text.

What the report did well

    What must be corrected

      Mentor's reflection

      8 · Written assessment

      Generated from your selections · editable

      This turns everything you selected above — every domain score, the gates, the maturity level, and the RGI — into a narrative you can edit and paste into your feedback. Composed deterministically from your choices (no model, no network); it narrates what you graded, it does not re-grade.

      9 · Evaluator's summative comment

      Required for gated verdicts

      Grounding: domain criteria draw on the Quality of Report Scale (Yang et al., SAGE Open Medicine 2014), ACR communication practice parameters and actionable-findings categories, and the communication-error taxonomies of Siewert et al. (AJR 2016). Framework: URRSF 3.1 / RRGF, Dr. Sharad Maheshwari — BeResponsibleAI / IRHAI. This manual edition is Evidence Assurance Level C throughout: every judgment is made and owned by the human evaluator. For educational and research use; not a certified competency examination. Do not enter patient identifiers anywhere on this sheet.

      Comments