Published to Maven Central: io.github.antonyrag:ragleap-rag:0.11.0
New: the document parser (DocumentParser): 27 of the Python package's 28 formats, each checked against the real Python parser's output on 79 generated fixture files. Parquet is deliberately not ported (Java Parquet readers need a very heavy Hadoop dependency tree). Legacy .xls needs the optional Apache POI dependency.
Fixed (found by the new cross-backend benchmark and failure-mode matrix):
- FAISS and Weaviate scores are now on the same [0, 1] scale as pgvector and Qdrant. They previously returned raw cosine in [-1, 1], so re-check any score thresholds you use with those two backends (#574, #576).
insertChunkrejects a wrong-dimension vector on FAISS, Qdrant and Weaviate, before anything is stored (#584).- Weaviate returns no results instead of an error when searching an empty collection or filtering on a key no chunk has yet (#585).
- Qdrant and Weaviate no longer leave a local row behind when a server write fails, so the chunk can be retried (#585).
Dependencies: PDFBox 3.0.8 and jsoup 1.23.2. Apache POI 5.5.1 is optional.
Verification: 422 tests, 0 failing in a full local run with Postgres, Qdrant, Weaviate and Ollama (tests that need live services skip themselves in CI). The benchmark covers pgvector, FAISS, Qdrant and Weaviate (identical retrieval: overlap@10 = 1.000 on 296 documents and 60 queries); Milvus and Pinecone are not live-verified. The raw data and the generator script are in the repository.
Not yet ported: OCR, web, audio and video ingestion. The Maven Central description, which still described the 0.5.0 feature set, now matches the library.