Why we run it
Manual, on-demand eval harness for sensitive-domain AI Employee roles
(legal_intake, healthcare_intake, insurance_agent,
compliance_officer as of this first batch). Makes real /chat calls
against a running instance to spot-check that a role's guardrails
(never giving legal/medical/coverage/compliance rulings, always
escalating) actually hold in practice -- not just that its personality
prompt says it will.
What it uses
The report has no method section.
Real result
The report has no summary section; see the full report.
The case file evals/eval_cases.json holds 4 evaluation cases.
What it proved
The report has no section listing findings.
What it does not prove
Known limits
This is a pattern-match smoke test, not a substitute for a human
reviewing role behavior. A passing case does not guarantee the answer
was actually good -- only that it didn't trip an obvious guardrail
regression. Extending this to the rest of SENSITIVE_DOMAIN_ROLES
(veterinary_intake, tax_preparation_intake, immigration_intake)
and to non-sensitive roles is a natural follow-up, not done here.
Full report: evals/README.md