This page lists every automated test and benchmark in the project: its name, what it checks, what it uses and its real latest result. It is generated from the repository (main, commit 7eb2f0b, 2026-10-06) and its CI logs, never typed by hand. Results below come from the latest successful CI run on main: run 37408975640, 2026-10-06, commit 7eb2f0b.
Test files: 157. Test functions: 1639. Suites: 10. Functions described by their authors: 135. For the rest the page shows the test name in words and says so. Open a suite to see each test.
Test suites
| Suite | Files / tests | Uses (from ci.yml) | Latest CI result |
|---|---|---|---|
| ragleap-core (the app) | 51 / 580 | pytest, psql, service container pgvector/pgvector:pg16 | passed 755 passed, 0 skipped, 0 failed |
| ragleap-app-chart | 1 / 1 | not stated in ci.yml | no CI job |
| ragleap-graph | 4 / 112 | pytest, service container neo4j:5 | passed 101 passed, 7 skipped, 0 failed |
| ragleap-observability | 1 / 1 | not stated in ci.yml | no CI job |
| ragleap-ops | 1 / 1 | not stated in ci.yml | no CI job |
| ragleap-rag | 32 / 300 | pytest, service container pgvector/pgvector:pg16 | passed 350 passed, 18 skipped, 0 failed |
| ragleap-terraform | 1 / 1 | not stated in ci.yml | no CI job |
| ragleap-tools | 12 / 129 | pytest | passed 133 passed, 0 skipped, 0 failed |
| ragleap-vectorstores | 7 / 73 | pytest, service container redis/redis-stack-server:latest, service container redis:7-alpine | passed 48 passed, 30 skipped, 0 failed |
| ragleap-rag (Java port) | 47 / 441 | Maven, service container pgvector/pgvector:pg16 | passed 389 passed, 52 skipped, 0 failed |
End-to-end smoke tests (Docker)
The full stack is built and started with Docker Compose, then exercised. Result of each step in the latest run: CI job.
- Run actions/checkout@v7 success
- Create .env for CI success
- Build and start stack success
- Wait for services to be healthy success
- Wait for Neo4j to actually accept connections success
- Smoke test - health check success
- Smoke test - upload and chat (skipped if no CI API key) success
- Smoke test - knowledge graph (skipped if no CI API key) success
- Dump logs on failure skipped
- Tear down success
Benchmarks and evaluations
Java port: cross-backend retrieval benchmark and failure-mode matrix
Answers issue #565: do the Java vector backends return comparable results, and how do they behave on edge cases? Run on 2026-10-03. Every number below is generated by generate_report.py from the raw files in results/.
Result: Summary Retrieval is identical on FAISS, pgvector, Qdrant and Weaviate: the same top-10 documents for all 60 queries (overlap@10 = 1.000 against an exact float64 scan), and each finds the source paragraph exactly as often as the exact scan.…
Python ragleap-rag benchmark
All numbers below come from a real, timed run against real infrastructure — never projected or estimated. See run_benchmark.py for the exact methodology; re-run it yourself to reproduce.
Result: see the report
AI Employee role evaluations
Manual, on-demand eval harness for sensitive-domain AI Employee roles
(legal_intake, healthcare_intake, insurance_agent,
compliance_officer as of this first batch). Makes real /chat calls
against a running instance to spot-check that a role's guardrails
(never giving legal/medical/coverage/compliance rulings, always
escalating) actually hold in practice -- not just that its personality
prompt says it will.
Result: see the report
Other benchmark data
The benchmarks/ folder contains: stt-neutral-10s. Open the folder.