RagLeap
Guide

Getting Started

Install

pip install ragleap-rag
# or
uv add ragleap-rag

You'll need a PostgreSQL database with the pgvector extension enabled, and an API key for at least one supported provider (Gemini, Anthropic, OpenAI, or any OpenAI-compatible endpoint including local Ollama).

Your first ingest + ask

from ragleap import RagLeap, ProviderConfig, EmbeddingConfig

rag = RagLeap(
    database_url="postgresql://user:pass@localhost:5432/mydb",
    primary=ProviderConfig(provider="gemini", model="gemini-3.6-flash", api_key="..."),
    embedder=EmbeddingConfig(
        provider="gemini", model="models/gemini-embedding-001",
        dimensions=3072, api_key="...",
    ),
)
rag.init_schema()

result = rag.ingest_text(filename="handbook.pdf", text="...your document text...")
answer = rag.ask("What's our PTO policy?")

print(answer["answer"])
print(answer["sources"])

model= is always required for both primary and embedder — there are no hardcoded model defaults (see the breaking change in v0.9.0 and why: no provider's default model stays safe to hardcode forever).

Ingesting other content types

rag.ingest("report.pdf", raw_bytes)                          # any of 28 formats
rag.ingest_url("https://example.com/article")                 # clean text extraction
rag.ingest_image("scan.png", raw_bytes, mode="ocr")            # OCR or vision captioning
rag.ingest_audio("call.mp3", raw_bytes)                        # pluggable transcription
rag.ingest_video("webinar.mp4", raw_bytes)                     # audio extraction + transcription

Running locally with zero API cost (Ollama)

rag = RagLeap(
    database_url="postgresql://user:pass@localhost:5432/mydb",
    primary=ProviderConfig(provider="custom", model="qwen2.5:0.5b",
                            api_key="ollama", base_url="http://localhost:11434/v1"),
    embedder=EmbeddingConfig(provider="custom", model="nomic-embed-text",
                              dimensions=768, api_key="ollama",
                              base_url="http://localhost:11434/v1"),
)

Fully local, no API key cost — see Benchmarks for real measured latency on this setup (it's meaningfully slower than a hosted API on CPU-only hardware, which is the honest tradeoff of running free and local).

Next steps

This page mirrors the GitHub wiki page. The wiki is the source of truth and may be newer.

Back to the blog