Published to Maven Central: io.github.antonyrag:ragleap-rag:0.12.0
New: ingestion. IngestionService ingests text, files, images, audio and video: it sanitizes, warns on prompt-injection phrases (a warning, not a block), applies input guardrails, chunks, embeds and stores the chunks, and fires the on_ingest hooks. If storing fails part-way, the document it created is deleted before the error propagates (the Python package leaves it half-ingested). Underneath it:
- OCR through the Tesseract binary: identical to the Python output on 26 generated images.
- Audio extraction from video through the ffmpeg binary: MP3 bytes identical to Python on 13 clips.
- Transcription with OpenAI Whisper, Deepgram or a custom function. Not live-verified: no paid accounts were available. It is stub-tested against the documented request shapes, the same label as Milvus and Pinecone.
Behaviour change: archives (.zip, .docx, .xlsx, .pptx, .odt, .ods, .odp, .epub) with more than 1,000 members or more than 100 MB declared uncompressed are now rejected, matching Python 0.13.0 (adjustable with DocumentParser.setZipLimits). This closes the known issue listed on v0.11.0.
Verification: 497 tests, 0 failing in a full local run with Postgres, Qdrant, Weaviate, Ollama, Tesseract and ffmpeg; tests that need live services or those binaries skip themselves in CI. The ingestion service is tested against FAISS and a stub backend, not end to end against pgvector, Qdrant or Weaviate. Requires the tesseract and ffmpeg binaries for OCR and video.
Not yet ported: URL and batch ingestion, and the high-level RagLeap entry point (ask, retrieval, streaming, evaluation).