Why we run it
All numbers below come from a real, timed run against real infrastructure — never projected or estimated. See run_benchmark.py for the exact methodology; re-run it yourself to reproduce.
What it uses
The report has no method section.
Real result
The report has no summary section; see the full report.
What it proved
The report has no section listing findings.
What it does not prove
Notes & Caveats
- Gemini free tier has two separate caps: 5 requests/minute and 20 generation requests/day. The daily cap, not the per-minute one, is the binding constraint for this run — it was exhausted partway through
gemini_pgvector's queries, beforegemini_faisscould run any. Numbers below reflect that real constraint, not Gemini's underlying latency in isolation. gemini_pgvectorquery numbers are a partial sample (n=8 of 15 requested) — the remaining 7 failed with429 RESOURCE_EXHAUSTEDonce the daily quota ran out. The P50/P95/mean above are computed only from the 8 that succeeded.gemini_faisshas zero successful queries this run — not because FAISS failed, but because the shared daily Gemini quota was already spent bygemini_pgvector's queries beforegemini_faissran. Ingestion succeeded fully (40/40 docs) for this config; only the query-latency/cost numbers are pending. A follow-up run (next UTC day, or with a paid Gemini tier) is needed to fill this in — see Phase 6 tracking.- Grok/xAI was excluded entirely this run — the configured key fails authentication (confirmed independent of ragleap-rag via a raw API call). Will be added once a working key is available.
- Ollama's latency (P50 ~67s, P95 ~85s per query) is real, not a bug — this reflects
qwen2.5:0.5brunning full retrieval + generation on this VPS's CPU with no GPU acceleration. It's the genuine tradeoff of a fully local, zero-cost setup on modest hardware, not a flaw in the pipeline.
Full report: packages/ragleap-rag/benchmark/BENCHMARKS.md