Shape & Signal

Responsible AI evaluation

Paused

Municipal AI Pilot Lab

A Responsible AI Pilot Evaluation Case Study — Pilot 001: Resident Correspondence Triage

Disclosure

This is an independent portfolio case study, not commissioned by or affiliated with any real municipality. The City of Alderbrook and all Pilot 001 correspondence are synthetic and fictional.

Context

This was a self-initiated research project — Shape & Signal's own investigation into both an AI classification approach and a responsible-AI evaluation methodology, run under realistic constraints rather than on behalf of a client or real municipal government.

Problem

Can AI assist municipal correspondence classification while preserving employee oversight?

Resident correspondence triage is repetitive but consequential — misrouting or missing an urgent message carries real costs — making it a reasonable test case for whether an AI system can recommend without deciding.

Signal

The trigger wasn't a specific municipality's request. It was a question about responsible-AI evaluation itself: whether a rigorous methodology — define the problem, restrict the AI's authority, preregister evaluation criteria, preserve human oversight, measure failures honestly, assess governance readiness, and reach an evidence-based decision — could be demonstrated end-to-end on a realistic task, rather than just described in the abstract.

Hypothesis

An AI classifier, constrained to recommend rather than decide, could reach operationally useful accuracy on correspondence triage while a deterministic human-review layer catches the cases where it shouldn't be trusted alone.

What we built

A classifier run over 200 synthetic resident messages, 3 independent trials each (600 classifications total), schema-validated against a fixed taxonomy, feeding a deterministic human-review router, an evaluation engine, and a 10-area governance assessment.

Constraints

The AI only recommends — category, department, risk, and urgency — and never decides, replies, or acts. Ground truth never touches the model layer except at evaluation time.

Governance findings are restricted to four non-claim words: demonstrated_in_prototype, requires_review, gap_before_production, not_assessed — explicitly excluding words like “compliant,” “safe,” or “certified.”

How it was evaluated

Dataset, ground truth, safety gates, and targets were checksummed and pre-registered before the run. A single controlled run followed, with deterministic review routing, evaluation against the frozen answer key, a 10-area governance assessment, and a final synthesis/decision layer — explicitly flagged as post-hoc, not preregistered.

What happened

Both pre-registered safety gates were satisfied in every trial. Two operational targets — unnecessary-review rate and cost per 1,000 classifications — were missed; both were pre-registered as non-decisive rather than blocking findings.

98.3%Target met

Operationally-acceptable accuracy (pooled)

exact match 92.8%

Target: ≥90%

100%

High-confidence accuracy

1.7%Target met

False-routing rate

2.83% under a stricter sensitivity definition

Target: ≤5%

100%Target met

Critical human-review recall

all 3 trials — safety gate

0Target met

Unsafe auto-routes

all 3 trials — safety gate

98.2% → 100%

Model-only vs. workflow-review recall

18.2%Target missed

Model-layer unnecessary-review rate

Target: ≤15%

$8.14Target missed

Cost per 1,000 classifications

pre-registered as non-decisive

Target: ≤$5.00

3 of 10

Governance areas demonstrated in this prototype

5 of 10

Governance areas requiring municipal review

What we learned

Accuracy and safety-gate performance alone are not sufficient evidence for a deployment decision. Cybersecurity was assessed as a gap before production, and records management was not assessed at all.

A rigorous governance vocabulary — refusing words like “compliant” — turned out to matter as much as the classification numbers themselves for making the result trustworthy.

Decision / current status

“Proceed to a bounded, controlled next-stage evaluation — not to production.”

“Proceed” here means proceeding to the next controlled testing stage under the listed conditions — it is not authorization for production deployment, autonomous routing, or use with real resident data.

What this does not authorize: production deployment, autonomous routing, real resident communication, or use of identifiable resident data without additional privacy, security, and records review.

Have a problem worth running through this same process?