Why we run it
Answers issue #565: do the Java vector backends return comparable results, and how do they behave on edge cases? Run on 2026-10-03. Every number below is generated by generate_report.py from the raw files in results/.
What it uses
Method
- Data: SQuAD v1.1 dev set (CC BY-SA 4.0). Every 7th paragraph is a document (296 documents); the first question of 60 of them is a query whose gold document is the paragraph it was written from.
- Vectors: embedded once with nomic-embed-text (Ollama), 768 dimensions, then the same vectors are stored in every backend, so embedding noise cannot affect the comparison.
- Ground truth: an exact float64 cosine scan computed in the test itself; no backend is used as its own reference.
- Metrics: overlap@10 (share of the exact top-10 returned), top-1 agreement, gold hit@k (does the backend find the source paragraph), the score range, and the deviation of each returned score from (cosine + 1) / 2.
- Environment: Java
openjdk version "21.0.12.1" 2026-08-18; Qdrant 1.19.1; Weaviate 1.27.0; pgvector 0.8.6 on Postgres; 1 CPU, 3.8 GB RAM.
Reproduce
python3 java/ragleap-rag/benchmark/prepare_dataset.py # downloads SQuAD, embeds with Ollama, writes /tmp/bench/corpus.json cd java/ragleap-rag RAGLEAP_BENCH_DIR=/tmp/bench RAGLEAP_BENCH_BACKENDS=faiss,pgvector,qdrant,weaviate mvn test -Dtest=CrossBackendBenchmarkTest RAGLEAP_BENCH_DIR=/tmp/bench RAGLEAP_BENCH_BACKENDS=faiss,pgvector,qdrant,weaviate mvn test -Dtest=FailureModeMatrixTest
Needs Qdrant on :6333, Weaviate on :8081 and a Postgres with pgvector and an empty database ragleap_java_bench on :5433 (see the test sources for the connection settings).
Environment it ran in
- date: 2026-10-03
- java: openjdk version "21.0.12.1" 2026-08-18
- qdrant: 1.19.1
- weaviate: 1.27.0
- pgvector: 0.8.6
- ollama: 0.32.5
- embeddingModel: nomic-embed-text (Ollama), 768 dimensions
- cpus: 1
- ramGb: 3.8
Real result
Summary
- Retrieval is identical on FAISS, pgvector, Qdrant and Weaviate: the same top-10 documents for all 60 queries (overlap@10 = 1.000 against an exact float64 scan), and each finds the source paragraph exactly as often as the exact scan.
- The benchmark found two score-scale bugs (FAISS and Weaviate returned raw cosine in [-1, 1] while pgvector and Qdrant return [0, 1]); both are fixed (#574, #576).
- The failure-mode matrix (19 cases x 4 backends) found three more bugs in Weaviate, FAISS and Qdrant, all fixed (#584, #585). After the fixes, 15 cases behave the same on all four backends and 4 differ in documented ways.
- Not covered: Milvus and Pinecone (no infrastructure to run them live; they are tested against HTTP stubs only), and approximate-search recall at scale (see Limits).
Retrieval results (after the fixes)
Exact float64 scan: gold hit@1 / @5 / @10 = 0.883 / 0.933 / 0.983.
| Backend | overlap@10 | top-1 agreement | gold hit@1 / @5 / @10 | score range | deviation from (cosine+1)/2, mean / max | median latency |
|---|---|---|---|---|---|---|
| FAISS | 1.000 | 1.000 | 0.883 / 0.933 / 0.983 | 0.744 to 0.949 | 0.00003 / 0.00005 | 1.0 ms |
| pgvector | 1.000 | 1.000 | 0.883 / 0.933 / 0.983 | 0.744 to 0.949 | 0.00003 / 0.00006 | 4.2 ms |
| Qdrant | 1.000 | 1.000 | 0.883 / 0.933 / 0.983 | 0.744 to 0.949 | 0.00003 / 0.00005 | 7.9 ms |
| Weaviate | 1.000 | 1.000 | 0.883 / 0.933 / 0.983 | 0.744 to 0.949 | 0.00003 / 0.00005 | 33.9 ms |
Latency is indicative only: one shared CPU with heavy steal time, a 296-document corpus, and run-to-run variation (Weaviate's p95 varied several-fold between runs).
What it proved
Bugs found and fixed
| Finding | Fix |
|---|---|
| Weaviate returned raw cosine in [-1, 1] | #574 |
| FAISS returned raw cosine in [-1, 1] | #576 |
| FAISS, Qdrant and Weaviate accepted or mishandled a wrong-dimension vector on insert (FAISS stored it silently) | #584 |
| Weaviate raised a GraphQL error when searching an empty collection or filtering on a key no chunk had yet | #585 |
| Qdrant and Weaviate left a local row behind when the server write failed, so the chunk could not be retried | #585 |
The Python FaissBackend and PineconeBackend still return raw cosine scores (issue #577).
What it does not prove
Limits
- 296 documents is small: Qdrant and Weaviate use HNSW indexes, but at this size they behave like an exact scan. This measures integration correctness, score comparability and retrieval quality, not approximate-search recall at scale.
- 60 queries from one dataset and one embedding model.
- One known gap: if a server applies a write but the response is lost, retrying the same chunk can still fail on Weaviate.
- Milvus and Pinecone are not benchmarked.
Proof: issues and pull requests
Every item this report references, read live from GitHub.
- #565 Cross-backend retrieval benchmark and failure-mode matrix for the Java port — closed 2026-10-04 · issue by @antonyrag · opened 2026-10-03
- #577 Python FaissBackend and PineconeBackend return raw cosine scores in [-1, 1], unlike the other backends — still open · issue by @antonyrag · opened 2026-10-03
- #574 fix(java-port): WeaviateBackend scores on the 0-1 scale, found by the cross-backend benchmark — merged 2026-10-03 · pull request by @antonyrag · opened 2026-10-03
- #576 fix(java-port): FaissBackend scores on the 0-1 scale, found by the cross-backend benchmark — merged 2026-10-03 · pull request by @antonyrag · opened 2026-10-03
- #584 fix(java-port): insertChunk rejects wrong-dimension vectors before storing, found by the failure-mode matrix — merged 2026-10-03 · pull request by @antonyrag · opened 2026-10-03
- #585 fix(java-port): Weaviate empty-collection and unknown-filter-key errors, and orphan local rows on failed server writes (Qdrant, Weaviate) — merged 2026-10-03 · pull request by @antonyrag · opened 2026-10-03
Full report: java/ragleap-rag/benchmark/RESULTS.md