RagLeap
Tests

Tests

This page lists every automated test and benchmark in the project: its name, what it checks, what it uses and its real latest result. It is generated from the repository (main, commit 7eb2f0b, 2026-10-06) and its CI logs, never typed by hand. Results below come from the latest successful CI run on main: run 37408975640, 2026-10-06, commit 7eb2f0b.

Test files: 157. Test functions: 1639. Suites: 10. Functions described by their authors: 135. For the rest the page shows the test name in words and says so. Open a suite to see each test.

Test suites

SuiteFiles / testsUses (from ci.yml)Latest CI result
ragleap-core (the app)51 / 580pytest, psql, service container pgvector/pgvector:pg16passed 755 passed, 0 skipped, 0 failed
ragleap-app-chart1 / 1not stated in ci.ymlno CI job
ragleap-graph4 / 112pytest, service container neo4j:5passed 101 passed, 7 skipped, 0 failed
ragleap-observability1 / 1not stated in ci.ymlno CI job
ragleap-ops1 / 1not stated in ci.ymlno CI job
ragleap-rag32 / 300pytest, service container pgvector/pgvector:pg16passed 350 passed, 18 skipped, 0 failed
ragleap-terraform1 / 1not stated in ci.ymlno CI job
ragleap-tools12 / 129pytestpassed 133 passed, 0 skipped, 0 failed
ragleap-vectorstores7 / 73pytest, service container redis/redis-stack-server:latest, service container redis:7-alpinepassed 48 passed, 30 skipped, 0 failed
ragleap-rag (Java port)47 / 441Maven, service container pgvector/pgvector:pg16passed 389 passed, 52 skipped, 0 failed

End-to-end smoke tests (Docker)

The full stack is built and started with Docker Compose, then exercised. Result of each step in the latest run: CI job.

  • Run actions/checkout@v7 success
  • Create .env for CI success
  • Build and start stack success
  • Wait for services to be healthy success
  • Wait for Neo4j to actually accept connections success
  • Smoke test - health check success
  • Smoke test - upload and chat (skipped if no CI API key) success
  • Smoke test - knowledge graph (skipped if no CI API key) success
  • Dump logs on failure skipped
  • Tear down success

Benchmarks and evaluations

Java port: cross-backend retrieval benchmark and failure-mode matrix

Answers issue #565: do the Java vector backends return comparable results, and how do they behave on edge cases? Run on 2026-10-03. Every number below is generated by generate_report.py from the raw files in results/.

Result: Summary Retrieval is identical on FAISS, pgvector, Qdrant and Weaviate: the same top-10 documents for all 60 queries (overlap@10 = 1.000 against an exact float64 scan), and each finds the source paragraph exactly as often as the exact scan.…

Python ragleap-rag benchmark

All numbers below come from a real, timed run against real infrastructure — never projected or estimated. See run_benchmark.py for the exact methodology; re-run it yourself to reproduce.

Result: see the report

AI Employee role evaluations

Manual, on-demand eval harness for sensitive-domain AI Employee roles (legal_intake, healthcare_intake, insurance_agent, compliance_officer as of this first batch). Makes real /chat calls against a running instance to spot-check that a role's guardrails (never giving legal/medical/coverage/compliance rulings, always escalating) actually hold in practice -- not just that its personality prompt says it will.

Result: see the report

Other benchmark data

The benchmarks/ folder contains: stt-neutral-10s. Open the folder.

Compiled by script from the public repository, its CI logs and the GitHub API. To correct something, open an issue.

All tests