RagLeap
Tests

AI Employee role evaluations

Why we run it

Manual, on-demand eval harness for sensitive-domain AI Employee roles (legal_intake, healthcare_intake, insurance_agent, compliance_officer as of this first batch). Makes real /chat calls against a running instance to spot-check that a role's guardrails (never giving legal/medical/coverage/compliance rulings, always escalating) actually hold in practice -- not just that its personality prompt says it will.

What it uses

The report has no method section.

Real result

The report has no summary section; see the full report.

The case file evals/eval_cases.json holds 4 evaluation cases.

What it proved

The report has no section listing findings.

What it does not prove

Known limits

This is a pattern-match smoke test, not a substitute for a human reviewing role behavior. A passing case does not guarantee the answer was actually good -- only that it didn't trip an obvious guardrail regression. Extending this to the rest of SENSITIVE_DOMAIN_ROLES (veterinary_intake, tax_preparation_intake, immigration_intake) and to non-sensitive roles is a natural follow-up, not done here.

Full report: evals/README.md

Compiled by script from the public repository, its CI logs and the GitHub API. To correct something, open an issue.

All tests