Responsible AI evaluation
PausedMunicipal AI Pilot Lab
A Responsible AI Pilot Evaluation Case Study — Pilot 001: Resident Correspondence Triage
Disclosure
This is an independent portfolio case study, not commissioned by or affiliated with any real municipality. The City of Alderbrook and all Pilot 001 correspondence are synthetic and fictional.
Context
This was a self-initiated research project — Shape & Signal's own investigation into both an AI classification approach and a responsible-AI evaluation methodology, run under realistic constraints rather than on behalf of a client or real municipal government.
Problem
Can AI assist municipal correspondence classification while preserving employee oversight?
Resident correspondence triage is repetitive but consequential — misrouting or missing an urgent message carries real costs — making it a reasonable test case for whether an AI system can recommend without deciding.
Signal
The trigger wasn't a specific municipality's request. It was a question about responsible-AI evaluation itself: whether a rigorous methodology — define the problem, restrict the AI's authority, preregister evaluation criteria, preserve human oversight, measure failures honestly, assess governance readiness, and reach an evidence-based decision — could be demonstrated end-to-end on a realistic task, rather than just described in the abstract.
Hypothesis
An AI classifier, constrained to recommend rather than decide, could reach operationally useful accuracy on correspondence triage while a deterministic human-review layer catches the cases where it shouldn't be trusted alone.
What we built
A classifier run over 200 synthetic resident messages, 3 independent trials each (600 classifications total), schema-validated against a fixed taxonomy, feeding a deterministic human-review router, an evaluation engine, and a 10-area governance assessment.
Constraints
The AI only recommends — category, department, risk, and urgency — and never decides, replies, or acts. Ground truth never touches the model layer except at evaluation time.
Governance findings are restricted to four non-claim words: demonstrated_in_prototype, requires_review, gap_before_production, not_assessed — explicitly excluding words like “compliant,” “safe,” or “certified.”
How it was evaluated
Dataset, ground truth, safety gates, and targets were checksummed and pre-registered before the run. A single controlled run followed, with deterministic review routing, evaluation against the frozen answer key, a 10-area governance assessment, and a final synthesis/decision layer — explicitly flagged as post-hoc, not preregistered.
What happened
Both pre-registered safety gates were satisfied in every trial. Two operational targets — unnecessary-review rate and cost per 1,000 classifications — were missed; both were pre-registered as non-decisive rather than blocking findings.
Operationally-acceptable accuracy (pooled)
exact match 92.8%
Target: ≥90%
High-confidence accuracy
False-routing rate
2.83% under a stricter sensitivity definition
Target: ≤5%
Critical human-review recall
all 3 trials — safety gate
Unsafe auto-routes
all 3 trials — safety gate
Model-only vs. workflow-review recall
Model-layer unnecessary-review rate
Target: ≤15%
Cost per 1,000 classifications
pre-registered as non-decisive
Target: ≤$5.00
Governance areas demonstrated in this prototype
Governance areas requiring municipal review
What we learned
Accuracy and safety-gate performance alone are not sufficient evidence for a deployment decision. Cybersecurity was assessed as a gap before production, and records management was not assessed at all.
A rigorous governance vocabulary — refusing words like “compliant” — turned out to matter as much as the classification numbers themselves for making the result trustworthy.
Decision / current status
“Proceed to a bounded, controlled next-stage evaluation — not to production.”
“Proceed” here means proceeding to the next controlled testing stage under the listed conditions — it is not authorization for production deployment, autonomous routing, or use with real resident data.
What this does not authorize: production deployment, autonomous routing, real resident communication, or use of identifiable resident data without additional privacy, security, and records review.