RagLeap
Tests

Tests: ragleap-rag

Test files: 32. Test functions: 300. Described by their authors (docstring or @DisplayName): 51. Where there is no description, the line shows the test name in words and says so.

Latest result: ragleap-rag-tests passed 350 passed, 18 skipped, 0 failed CI job

test_async_and_batch.py — 7 tests 7 passed, 0 failed, 0 skipped

Tests for async ingest/ask methods and ingest_batch() concurrent mixed-type ingestion with partial-success semantics.

  • test_aingest_text_works passed
    Aingest text works (no description in the source; shown from the name)
  • test_aask_works passed
    Aask works (no description in the source; shown from the name)
  • test_aask_stream_yields_pieces passed
    Aask stream yields pieces (no description in the source; shown from the name)
  • test_ingest_batch_all_succeed passed
    Ingest batch all succeed (no description in the source; shown from the name)
  • test_ingest_batch_partial_failure_does_not_block_others passed
    Ingest batch partial failure does not block others (no description in the source; shown from the name)
  • test_ingest_batch_preserves_input_order passed
    Ingest batch preserves input order (no description in the source; shown from the name)
  • test_aask_stream_respects_rerank_and_metadata_filter passed
    Aask stream respects rerank and metadata filter (no description in the source; shown from the name)
test_cache.py — 8 tests 8 passed, 0 failed, 0 skipped

Tests for the in-memory query embedding cache — hit/miss tracking, LRU eviction, disabled-cache behavior, and integration through ask().

  • test_cache_miss_then_hit passed
    Cache miss then hit (no description in the source; shown from the name)
  • test_cache_key_includes_model passed
    Cache key includes model (no description in the source; shown from the name)
  • test_cache_stats_tracks_hits_and_misses passed
    Cache stats tracks hits and misses (no description in the source; shown from the name)
  • test_cache_evicts_least_recently_used_when_full passed
    Cache evicts least recently used when full (no description in the source; shown from the name)
  • test_cache_clear_resets_everything passed
    Cache clear resets everything (no description in the source; shown from the name)
  • test_rag_cache_disabled_returns_zeroed_stats passed
    Rag cache disabled returns zeroed stats (no description in the source; shown from the name)
  • test_rag_ask_repeated_query_is_a_cache_hit passed
    Rag ask repeated query is a cache hit (no description in the source; shown from the name)
  • test_rag_cache_backend_redis_requires_redis_url passed
    Rag cache backend redis requires redis url (no description in the source; shown from the name)
test_chunker.py — 7 tests 7 passed, 0 failed, 0 skipped

Tests for TextChunker — chunk windowing/overlap, and token_count accuracy (real tiktoken counts when available, honest word-count fallback with token_count_is_exact=False when tiktoken can't load).

  • test_chunk_text_basic_windowing passed
    Chunk text basic windowing (no description in the source; shown from the name)
  • test_chunk_text_empty_input_returns_empty_list passed
    Chunk text empty input returns empty list (no description in the source; shown from the name)
  • test_chunk_text_raises_on_invalid_overlap passed
    Chunk text raises on invalid overlap (no description in the source; shown from the name)
  • test_chunk_text_reports_token_count_is_exact_flag passed
    Chunk text reports token count is exact flag (no description in the source; shown from the name)
  • test_chunk_text_token_count_matches_tiktoken_when_exact passed
    Chunk text token count matches tiktoken when exact (no description in the source; shown from the name)
  • test_chunk_text_fallback_word_count_when_tiktoken_unavailable passed
    Forces the fallback path regardless of the real environment, so
  • test_chunk_text_fallback_on_encoding_load_failure passed
    tiktoken installed but its encoding can't load (e.g. no network
test_ci_extras.py — 1 tests 1 passed, 0 failed, 0 skipped

Fails in CI (CI=true) if an optional extra the suite depends on is missing, so tests that need it cannot silently become skips. Locally these are skipped.

  • test_extra_is_installed_in_ci passed
    Extra is installed in ci (no description in the source; shown from the name)
test_cost.py — 17 tests 17 passed, 0 failed, 0 skipped

Tests for cost tracking - per-call USD cost computation from real provider-reported token usage, and budget-triggered fallback to a cheaper/local provider on subsequent calls.

  • test_compute_cost_known_provider_model passed
    Compute cost known provider model (no description in the source; shown from the name)
  • test_compute_cost_unknown_provider_returns_none passed
    Compute cost unknown provider returns none (no description in the source; shown from the name)
  • test_compute_cost_unknown_model_returns_none passed
    Compute cost unknown model returns none (no description in the source; shown from the name)
  • test_compute_cost_no_usage_returns_none passed
    Compute cost no usage returns none (no description in the source; shown from the name)
  • test_compute_cost_ollama_wildcard_is_zero passed
    Compute cost ollama wildcard is zero (no description in the source; shown from the name)
  • test_cost_tracker_pricing_table_override_wins passed
    Cost tracker pricing table override wins (no description in the source; shown from the name)
  • test_cost_tracker_override_does_not_remove_other_seed_models passed
    Overriding one model for a provider shouldn't wipe out the other
  • test_cost_tracker_accumulates_across_calls passed
    Cost tracker accumulates across calls (no description in the source; shown from the name)
  • test_cost_tracker_no_budget_set_is_never_over_budget passed
    Cost tracker no budget set is never over budget (no description in the source; shown from the name)
  • test_cost_tracker_over_budget_after_threshold_crossed passed
    Cost tracker over budget after threshold crossed (no description in the source; shown from the name)
  • test_ask_result_has_cost_field_with_known_pricing passed
    Ask result has cost field with known pricing (no description in the source; shown from the name)
  • test_ask_result_cost_unavailable_for_unpriced_model passed
    A model not in the seed pricing table - cost must honestly report
  • test_ask_pricing_table_override_applies_through_rag_constructor passed
    Ask pricing table override applies through rag constructor (no description in the source; shown from the name)
  • test_cumulative_cost_grows_across_multiple_ask_calls passed
    Cumulative cost grows across multiple ask calls (no description in the source; shown from the name)
  • test_budget_fallback_triggers_switch_to_fallback_provider passed
    First call stays on primary (nothing spent yet, so not over budget).
  • test_no_budget_fallback_configured_never_switches_provider passed
    budget_usd_per_month set but no budget_fallback provider -> stays
  • test_ask_stream_cost_is_always_unavailable passed
    Streaming has no token usage data (see generate_answer_stream's
test_embedding.py — 29 tests 25 passed, 0 failed, 4 skipped

Tests for embedding provider expansion (v0.7.0), updated for v0.9.0's removal of all hardcoded model/dimension defaults - every provider now requires explicit model= and dimensions= (constructor arg or env var), matching the precedent "together" always set. Split into two groups: config-resolution tests (no network calls, no keys needed beyond fakes) and real live tests against a genuinely running local Ollama instance, skipped automatically if Ollama isn't reachable so CI never fails on its absence.

  • test_mistral_requires_explicit_model_and_dimensions passed
    Mistral requires explicit model and dimensions (no description in the source; shown from the name)
  • test_mistral_works_with_explicit_model_and_dimensions passed
    Mistral works with explicit model and dimensions (no description in the source; shown from the name)
  • test_together_requires_explicit_model_and_dimensions passed
    Together requires explicit model and dimensions (no description in the source; shown from the name)
  • test_together_works_with_explicit_model_and_dimensions passed
    Together works with explicit model and dimensions (no description in the source; shown from the name)
  • test_cohere_requires_explicit_model_and_dimensions passed
    Cohere requires explicit model and dimensions (no description in the source; shown from the name)
  • test_cohere_works_with_explicit_model_and_dimensions passed
    Cohere works with explicit model and dimensions (no description in the source; shown from the name)
  • test_cohere_defaults_base_url_to_first_party_when_unset passed
    Regression test for antonyrag/ragleap-core#360's CO1 finding:
  • test_cohere_honors_explicit_base_url passed
    The actual bug fix: a caller-supplied base_url must now be used
  • test_voyage_requires_explicit_model_and_dimensions passed
    Voyage requires explicit model and dimensions (no description in the source; shown from the name)
  • test_voyage_works_with_explicit_model_and_dimensions passed
    Voyage works with explicit model and dimensions (no description in the source; shown from the name)
  • test_voyage_defaults_base_url_to_first_party_when_unset passed
    Same regression coverage as cohere's, for V1's finding.
  • test_voyage_honors_explicit_base_url passed
    Voyage honors explicit base url (no description in the source; shown from the name)
  • test_missing_api_key_raises_for_mistral passed
    Missing api key raises for mistral (no description in the source; shown from the name)
  • test_missing_api_key_raises_for_cohere passed
    Missing api key raises for cohere (no description in the source; shown from the name)
  • test_missing_api_key_raises_for_voyage passed
    Missing api key raises for voyage (no description in the source; shown from the name)
  • test_missing_model_raises_even_with_valid_api_key passed
    The core of the v0.9.0 change - having an API key isn't enough,
  • test_missing_dimensions_raises_even_with_model_specified passed
    Missing dimensions raises even with model specified (no description in the source; shown from the name)
  • test_ollama_needs_no_api_key_but_still_needs_model_and_dimensions passed
    Mirrors generation.py's existing ollama exemption from requiring
  • test_env_var_fallback_for_mistral passed
    Env var fallback for mistral (no description in the source; shown from the name)
  • test_unknown_provider_raises passed
    Unknown provider raises (no description in the source; shown from the name)
  • test_explicit_dimensions_always_required_no_override_concept passed
    There's no "default to override" anymore - dimensions must always
  • test_custom_provider_requires_base_url passed
    Custom provider requires base url (no description in the source; shown from the name)
  • test_custom_provider_works_with_all_fields_explicit passed
    Proves the escape hatch this session was about: any OpenAI-
  • test_custom_provider_env_var_fallback passed
    Custom provider env var fallback (no description in the source; shown from the name)
  • test_custom_provider_dispatches_to_openai_compatible_path passed
    Confirms EmbeddingService actually routes provider="custom" through
  • test_ollama_embed_text_returns_real_vector skipped
    Ollama embed text returns real vector (no description in the source; shown from the name)
  • test_ollama_embed_batch_returns_real_vectors skipped
    Ollama embed batch returns real vectors (no description in the source; shown from the name)
  • test_ollama_embed_text_empty_string_returns_none skipped
    Ollama embed text empty string returns none (no description in the source; shown from the name)
  • test_ollama_full_rag_integration_with_faiss skipped
    Real end-to-end proof: Ollama embeds -> stored in a real FAISS
test_evaluation.py — 8 tests 8 passed, 0 failed, 0 skipped

Tests for rag.evaluate() - deterministic retrieval hit-rate, keyword coverage, and citation groundedness checks. Not testing real semantic quality (the fake embedder/generator have none) - testing that the scoring logic itself is correct, using real Postgres full-text search for retrieval and the fake generator's real prompt-echo behavior to create genuinely checkable keyword overlap.

  • test_evaluate_requires_at_least_one_case passed
    Evaluate requires at least one case (no description in the source; shown from the name)
  • test_evaluate_retrieval_hit_rate_all_hits passed
    Evaluate retrieval hit rate all hits (no description in the source; shown from the name)
  • test_evaluate_retrieval_hit_rate_partial_miss passed
    Evaluate retrieval hit rate partial miss (no description in the source; shown from the name)
  • test_evaluate_keyword_coverage_uses_real_answer_content passed
    The fake generator echoes the prompt's tail, which includes the
  • test_evaluate_keyword_coverage_partial passed
    Evaluate keyword coverage partial (no description in the source; shown from the name)
  • test_evaluate_groundedness_when_keyword_in_cited_chunk passed
    Ingest text containing 'bananas' so it ends up in the cited
  • test_evaluate_returns_none_for_metrics_with_no_applicable_cases passed
    Evaluate returns none for metrics with no applicable cases (no description in the source; shown from the name)
  • test_evaluate_passes_through_ask_kwargs passed
    Evaluate passes through ask kwargs (no description in the source; shown from the name)
test_faiss_backend.py — 10 tests 10 passed, 0 failed, 0 skipped

Real tests for FAISSBackend - genuine FAISS index + SQLite sidecar, no mocking of the backend itself. Skipped automatically if the [faiss] extra isn't installed, so it never blocks CI runs without it.

  • test_faiss_ingest_and_ask_roundtrip passed
    Faiss ingest and ask roundtrip (no description in the source; shown from the name)
  • test_faiss_dense_search_finds_correct_document passed
    Dense-only search with a non-semantic fake embedder can only
  • test_faiss_backend_does_not_support_sparse passed
    Faiss backend does not support sparse (no description in the source; shown from the name)
  • test_faiss_hybrid_mode_gracefully_degrades_to_dense passed
    hybrid=True should not crash against a backend with no sparse
  • test_faiss_list_documents passed
    Faiss list documents (no description in the source; shown from the name)
  • test_faiss_delete_document_removes_it passed
    Faiss delete document removes it (no description in the source; shown from the name)
  • test_faiss_delete_unknown_document_returns_false passed
    Faiss delete unknown document returns false (no description in the source; shown from the name)
  • test_faiss_metadata_filter_post_filters_correctly passed
    Faiss metadata filter post filters correctly (no description in the source; shown from the name)
  • test_faiss_update_document_preserves_filename passed
    Faiss update document preserves filename (no description in the source; shown from the name)
  • test_faiss_persistence_across_backend_instances passed
    The real point of persist_directory= - data survives a fresh
test_guardrails.py — 8 tests 8 passed, 0 failed, 0 skipped

Tests for input_guardrails/output_guardrails - user-supplied validation callbacks that extend (not replace) sanitization and injection-risk detection.

  • test_input_guardrail_can_modify_text passed
    Input guardrail can modify text (no description in the source; shown from the name)
  • test_input_guardrail_violation_aborts_ingestion_with_nothing_stored passed
    Input guardrail violation aborts ingestion with nothing stored (no description in the source; shown from the name)
  • test_input_guardrails_run_in_order passed
    Input guardrails run in order (no description in the source; shown from the name)
  • test_ask_output_guardrail_passes_through_when_no_violation passed
    Ask output guardrail passes through when no violation (no description in the source; shown from the name)
  • test_ask_output_guardrail_blocks_and_replaces_answer passed
    Ask output guardrail blocks and replaces answer (no description in the source; shown from the name)
  • test_ask_without_output_guardrails_has_no_blocked_key passed
    Ask without output guardrails has no blocked key (no description in the source; shown from the name)
  • test_ask_stream_guardrail_violation_logs_warning_but_still_yields passed
    Ask stream guardrail violation logs warning but still yields (no description in the source; shown from the name)
  • test_ask_stream_output_guardrail_passes_through_when_no_violation passed
    Ask stream output guardrail passes through when no violation (no description in the source; shown from the name)
test_ingestion.py — 9 tests 9 passed, 0 failed, 0 skipped

Tests for RagLeap.ingest_text() — chunking, storage, sanitization, injection-risk warnings, and error handling.

  • test_ingest_text_returns_document_id_and_chunk_count passed
    Ingest text returns document id and chunk count (no description in the source; shown from the name)
  • test_ingest_text_empty_raises_value_error passed
    Ingest text empty raises value error (no description in the source; shown from the name)
  • test_ingest_text_whitespace_only_raises passed
    Ingest text whitespace only raises (no description in the source; shown from the name)
  • test_ingest_text_sanitizes_control_chars_by_default passed
    Ingest text sanitizes control chars by default (no description in the source; shown from the name)
  • test_ingest_text_sanitize_false_preserves_raw_text passed
    Ingest text sanitize false preserves raw text (no description in the source; shown from the name)
  • test_ingest_text_logs_warning_on_injection_risk passed
    Ingest text logs warning on injection risk (no description in the source; shown from the name)
  • test_ingest_text_stores_metadata passed
    Ingest text stores metadata (no description in the source; shown from the name)
  • test_ingest_stores_metadata passed
    Regression test for the real gap ingest() had: it silently
  • test_ingest_without_metadata_still_works passed
    Confirms the default (no metadata=) path is unchanged - the
test_lifecycle.py — 10 tests 10 passed, 0 failed, 0 skipped

Tests for document lifecycle: list_documents, delete_document, update_document — including the known metadata-loss limitation.

  • test_list_documents_returns_ingested_docs passed
    List documents returns ingested docs (no description in the source; shown from the name)
  • test_list_documents_includes_chunk_count passed
    List documents includes chunk count (no description in the source; shown from the name)
  • test_list_documents_ordered_most_recent_first passed
    List documents ordered most recent first (no description in the source; shown from the name)
  • test_delete_document_removes_it_and_returns_true passed
    Delete document removes it and returns true (no description in the source; shown from the name)
  • test_delete_document_unknown_id_returns_false passed
    Delete document unknown id returns false (no description in the source; shown from the name)
  • test_delete_document_cascades_to_chunks passed
    Delete document cascades to chunks (no description in the source; shown from the name)
  • test_update_document_creates_new_document_id passed
    Update document creates new document id (no description in the source; shown from the name)
  • test_update_document_preserves_filename_if_not_given passed
    Update document preserves filename if not given (no description in the source; shown from the name)
  • test_update_document_renames_when_filename_given passed
    Update document renames when filename given (no description in the source; shown from the name)
  • test_update_document_known_limitation_metadata_is_lost passed
    Documents a REAL, currently-open limitation (see README/Part 5
test_memory.py — 7 tests 7 passed, 0 failed, 0 skipped

Tests for persistent conversation memory (session-scoped, Postgres-backed).

  • test_get_history_empty_for_new_session passed
    Get history empty for new session (no description in the source; shown from the name)
  • test_ask_with_session_id_stores_history passed
    Ask with session id stores history (no description in the source; shown from the name)
  • test_ask_without_session_id_stores_nothing passed
    Ask without session id stores nothing (no description in the source; shown from the name)
  • test_multiple_turns_accumulate_in_order passed
    Multiple turns accumulate in order (no description in the source; shown from the name)
  • test_sessions_are_isolated passed
    Sessions are isolated (no description in the source; shown from the name)
  • test_clear_session_removes_all_messages passed
    Clear session removes all messages (no description in the source; shown from the name)
  • test_history_injected_into_prompt_via_generator passed
    The fake generator echoes the tail of the prompt it received -
test_metadata_filtering.py — 4 tests 4 passed, 0 failed, 0 skipped

Tests for metadata_filter on ask() — JSONB containment filtering, the multi-tenant isolation mechanism.

  • test_metadata_filter_restricts_results_to_matching_tenant passed
    Metadata filter restricts results to matching tenant (no description in the source; shown from the name)
  • test_metadata_filter_hybrid_mode_also_respects_filter passed
    Metadata filter hybrid mode also respects filter (no description in the source; shown from the name)
  • test_no_metadata_filter_returns_from_any_tenant passed
    No metadata filter returns from any tenant (no description in the source; shown from the name)
  • test_metadata_filter_no_match_returns_empty_sources passed
    Metadata filter no match returns empty sources (no description in the source; shown from the name)
test_milvus_backend.py — 16 tests 16 passed, 0 failed, 0 skipped

Tests for MilvusBackend - NOT live-verified against a real Milvus/ Zilliz Cloud instance (same honest caveat as the other new vector backends this session). Mocks the actual pymilvus MilvusClient entirely; every method signature used was verified against the actual installed pymilvus==3.0.1 package's real source during development (including that id_type="string" is a genuinely supported primary key type, not assumed from documentation alone).

  • test_requires_persist_directory passed
    Requires persist directory (no description in the source; shown from the name)
  • test_requires_uri passed
    Requires uri (no description in the source; shown from the name)
  • test_uri_from_env_var passed
    Uri from env var (no description in the source; shown from the name)
  • test_vector_key_construction passed
    Vector key construction (no description in the source; shown from the name)
  • test_build_filter_expr_single_condition passed
    Build filter expr single condition (no description in the source; shown from the name)
  • test_build_filter_expr_multiple_conditions_anded passed
    Build filter expr multiple conditions anded (no description in the source; shown from the name)
  • test_build_filter_expr_numeric_value_not_quoted passed
    Build filter expr numeric value not quoted (no description in the source; shown from the name)
  • test_build_filter_expr_empty_when_no_filter passed
    Build filter expr empty when no filter (no description in the source; shown from the name)
  • test_insert_document_and_list_documents passed
    Insert document and list documents (no description in the source; shown from the name)
  • test_insert_chunk_calls_insert_with_correct_row passed
    Insert chunk calls insert with correct row (no description in the source; shown from the name)
  • test_search_dense_returns_chunks_with_text_from_sqlite passed
    Search dense returns chunks with text from sqlite (no description in the source; shown from the name)
  • test_search_dense_normalizes_cosine_similarity_to_unit_range passed
    Regression test for the similarity_score normalization fix. Milvus
  • test_search_dense_skips_orphaned_hits passed
    Search dense skips orphaned hits (no description in the source; shown from the name)
  • test_delete_document_calls_delete_with_vector_keys passed
    Delete document calls delete with vector keys (no description in the source; shown from the name)
  • test_supports_sparse_is_false passed
    Supports sparse is false (no description in the source; shown from the name)
  • test_milvus_backend_importable_from_vectorstores passed
    Milvus backend importable from vectorstores (no description in the source; shown from the name)
test_net_guard.py — 15 tests 18 passed, 0 failed, 0 skipped

Attack-case tests for ragleap._net, ragleap.web and ingest_url(). A loopback HTTP server stands in for the network; `pretend_public` makes names ending in .test resolve to it while still passing the public-address check, so redirect re-validation and IP pinning are exercised for real.

  • test_non_public_targets_are_refused passed
    Non public targets are refused (no description in the source; shown from the name)
  • test_odd_address_forms_are_refused_or_unresolvable passed
    Odd address forms are refused or unresolvable (no description in the source; shown from the name)
  • test_bad_schemes_credentials_and_missing_host_are_refused passed
    Bad schemes credentials and missing host are refused (no description in the source; shown from the name)
  • test_allow_private_fetches_loopback passed
    Allow private fetches loopback (no description in the source; shown from the name)
  • test_connects_to_validated_ip_and_sends_original_host passed
    Connects to validated ip and sends original host (no description in the source; shown from the name)
  • test_redirect_to_private_address_is_refused passed
    Redirect to private address is refused (no description in the source; shown from the name)
  • test_redirect_loop_returns_none passed
    Redirect loop returns none (no description in the source; shown from the name)
  • test_oversized_body_returns_none passed
    Oversized body returns none (no description in the source; shown from the name)
  • test_non_200_and_encoded_responses_return_none passed
    Non 200 and encoded responses return none (no description in the source; shown from the name)
  • test_slow_server_hits_the_timeout passed
    Slow server hits the timeout (no description in the source; shown from the name)
  • test_fetch_url_text_refuses_loopback_by_default passed
    Fetch url text refuses loopback by default (no description in the source; shown from the name)
  • test_fetch_url_text_extracts_when_private_is_allowed passed
    Fetch url text extracts when private is allowed (no description in the source; shown from the name)
  • test_ingest_url_refuses_loopback_by_default passed
    Ingest url refuses loopback by default (no description in the source; shown from the name)
  • test_compressed_responses_are_decoded passed
    Compressed responses are decoded (no description in the source; shown from the name)
  • test_decompression_bomb_returns_none passed
    Decompression bomb returns none (no description in the source; shown from the name)
test_observability.py — 7 tests 7 passed, 0 failed, 0 skipped

Tests for on_ingest/on_query/on_answer observability hooks - fire- and-forget event emission that never breaks the actual RAG operation, even when a hook raises.

  • test_on_ingest_fires_with_correct_event_shape passed
    On ingest fires with correct event shape (no description in the source; shown from the name)
  • test_on_query_and_on_answer_fire_on_ask passed
    On query and on answer fire on ask (no description in the source; shown from the name)
  • test_on_query_and_on_answer_fire_on_ask_stream passed
    On query and on answer fire on ask stream (no description in the source; shown from the name)
  • test_multiple_handlers_all_fire_in_order passed
    Multiple handlers all fire in order (no description in the source; shown from the name)
  • test_broken_hook_does_not_break_ingestion passed
    Broken hook does not break ingestion (no description in the source; shown from the name)
  • test_broken_hook_does_not_break_ask passed
    Broken hook does not break ask (no description in the source; shown from the name)
  • test_no_hooks_configured_is_a_true_no_op passed
    No hooks set (the default) - fire_event should be a silent
test_package_metadata.py — 1 tests 1 passed, 0 failed, 0 skipped

No file-level description in the source.

  • test_version_matches_pyproject passed
    Version matches pyproject (no description in the source; shown from the name)
test_parser_limits.py — 5 tests 5 passed, 0 failed, 0 skipped

No file-level description in the source.

  • test_zip_within_limits_still_extracts passed
    Zip within limits still extracts (no description in the source; shown from the name)
  • test_too_many_members_is_rejected passed
    Too many members is rejected (no description in the source; shown from the name)
  • test_declared_size_over_limit_is_rejected passed
    Declared size over limit is rejected (no description in the source; shown from the name)
  • test_container_formats_are_checked_too passed
    Container formats are checked too (no description in the source; shown from the name)
  • test_running_budget_applies_even_if_the_precheck_is_bypassed passed
    Running budget applies even if the precheck is bypassed (no description in the source; shown from the name)
test_parsers.py — 5 tests 5 passed, 0 failed, 0 skipped

Per-extension extraction tests for ragleap.parsers.extract_text().

  • test_every_supported_extension_has_a_sample passed
    Every supported extension has a sample (no description in the source; shown from the name)
  • test_extracts_marker_from_minimal_sample passed
    Extracts marker from minimal sample (no description in the source; shown from the name)
  • test_legacy_office_formats_are_rejected_with_a_conversion_hint passed
    Legacy office formats are rejected with a conversion hint (no description in the source; shown from the name)
  • test_unknown_extension_is_rejected passed
    Unknown extension is rejected (no description in the source; shown from the name)
  • test_parquet_without_pandas_raises_value_error passed
    Parquet without pandas raises value error (no description in the source; shown from the name)
test_pinecone_backend.py — 20 tests 20 passed, 0 failed, 0 skipped

Tests for PineconeBackend - NOT live-verified against a real Pinecone account (same honest caveat as mistral/together/cohere/voyage embedding providers). These tests mock the actual Pinecone client entirely and verify: (1) constructor validation, (2) pure helper functions, (3) the SQLite sidecar's CRUD correctness in isolation, and (4) that the right Pinecone SDK methods get called with the right arguments and the right attribute-access pattern (not dict-style) - every attribute path used here (IndexList.names, IndexStatus.ready, ScoredVector.id/.score) was verified against the actual installed pinecone==9.1.0 package's real source code during development, not assumed from documentation alone.

  • test_requires_persist_directory passed
    Requires persist directory (no description in the source; shown from the name)
  • test_requires_api_key passed
    Requires api key (no description in the source; shown from the name)
  • test_api_key_from_env_var passed
    Api key from env var (no description in the source; shown from the name)
  • test_default_index_name_and_region passed
    Default index name and region (no description in the source; shown from the name)
  • test_creates_sqlite_tables_on_init passed
    Creates sqlite tables on init (no description in the source; shown from the name)
  • test_vector_id_construction passed
    Vector id construction (no description in the source; shown from the name)
  • test_build_filter_none_when_no_filter passed
    Build filter none when no filter (no description in the source; shown from the name)
  • test_build_filter_translates_to_pinecone_eq_syntax passed
    Build filter translates to pinecone eq syntax (no description in the source; shown from the name)
  • test_insert_document_and_list_documents passed
    Insert document and list documents (no description in the source; shown from the name)
  • test_get_document_filename passed
    Get document filename (no description in the source; shown from the name)
  • test_init_schema_creates_index_when_not_existing passed
    Init schema creates index when not existing (no description in the source; shown from the name)
  • test_init_schema_skips_create_when_index_exists passed
    Init schema skips create when index exists (no description in the source; shown from the name)
  • test_insert_chunk_upserts_to_pinecone_with_correct_id_and_metadata passed
    Insert chunk upserts to pinecone with correct id and metadata (no description in the source; shown from the name)
  • test_search_dense_returns_chunks_with_text_from_sqlite passed
    Search dense returns chunks with text from sqlite (no description in the source; shown from the name)
  • test_search_dense_skips_orphaned_vectors passed
    A vector exists in Pinecone (per the mocked response) but has no
  • test_search_dense_wrong_dimensions_returns_empty passed
    Search dense wrong dimensions returns empty (no description in the source; shown from the name)
  • test_delete_document_deletes_from_pinecone_and_sqlite passed
    Delete document deletes from pinecone and sqlite (no description in the source; shown from the name)
  • test_delete_document_returns_false_when_not_found passed
    Delete document returns false when not found (no description in the source; shown from the name)
  • test_supports_sparse_is_false passed
    Supports sparse is false (no description in the source; shown from the name)
  • test_pinecone_backend_importable_from_vectorstores passed
    Confirms it's actually wired into vectorstores/__init__.py, not
test_qdrant_backend.py — 15 tests 15 passed, 0 failed, 0 skipped

Tests for QdrantBackend - NOT live-verified against a real Qdrant instance (same honest caveat as PineconeBackend/WeaviateBackend). Mocks the actual Qdrant client entirely; every method signature and pydantic model field used was verified against the actual installed qdrant-client==1.18.0 package's real source during development.

  • test_requires_persist_directory passed
    Requires persist directory (no description in the source; shown from the name)
  • test_requires_url passed
    Requires url (no description in the source; shown from the name)
  • test_url_from_env_var passed
    Url from env var (no description in the source; shown from the name)
  • test_vector_key_construction passed
    Vector key construction (no description in the source; shown from the name)
  • test_deterministic_uuid_is_stable passed
    Deterministic uuid is stable (no description in the source; shown from the name)
  • test_build_filter_translates_correctly passed
    Build filter translates correctly (no description in the source; shown from the name)
  • test_build_filter_none_when_empty passed
    Build filter none when empty (no description in the source; shown from the name)
  • test_insert_document_and_list_documents passed
    Insert document and list documents (no description in the source; shown from the name)
  • test_insert_chunk_calls_upsert_with_correct_point passed
    Insert chunk calls upsert with correct point (no description in the source; shown from the name)
  • test_search_dense_returns_chunks_with_text_from_sqlite passed
    Search dense returns chunks with text from sqlite (no description in the source; shown from the name)
  • test_search_dense_skips_orphaned_points passed
    Search dense skips orphaned points (no description in the source; shown from the name)
  • test_delete_document_calls_delete_with_point_ids passed
    Delete document calls delete with point ids (no description in the source; shown from the name)
  • test_search_dense_normalizes_cosine_similarity_to_unit_range passed
    Regression test for the similarity_score normalization fix. Qdrant
  • test_supports_sparse_is_false passed
    Supports sparse is false (no description in the source; shown from the name)
  • test_qdrant_backend_importable_from_vectorstores passed
    Qdrant backend importable from vectorstores (no description in the source; shown from the name)
test_qdrant_backend_live.py — 6 tests 0 passed, 0 failed, 6 skipped

Live tests for QdrantBackend against a real running Qdrant instance. Gated on QDRANT_TEST_URL (e.g. "http://localhost:6333"), same pattern as the OpenSearch/Upstash live-gated tests in ragleap-vectorstores - skip cleanly (not fail) when no real instance is configured.

  • test_init_schema_is_idempotent skipped
    Init schema is idempotent (no description in the source; shown from the name)
  • test_search_dense_orders_and_normalizes_score skipped
    Search dense orders and normalizes score (no description in the source; shown from the name)
  • test_search_dense_metadata_filter skipped
    Search dense metadata filter (no description in the source; shown from the name)
  • test_list_documents_and_get_filename skipped
    List documents and get filename (no description in the source; shown from the name)
  • test_delete_document_removes_vectors skipped
    Delete document removes vectors (no description in the source; shown from the name)
  • test_supports_sparse_is_false skipped
    Supports sparse is false (no description in the source; shown from the name)
test_query_rewrite.py — 19 tests 19 passed, 0 failed, 0 skipped

Tests for query rewriting/expansion (contextual, hyde, multi_query). Live semantic verification (real Gemini calls proving contextual rewrite correctly resolves pronouns, HyDE generates on-topic passages, multi_query produces genuinely distinct phrasings) was done manually this session and is documented in CHANGELOG - not automated into CI since it needs a real, currently-valid API key CI doesn't have.

  • test_contextual_rewrite_no_history_returns_original_query_no_call passed
    Contextual rewrite no history returns original query no call (no description in the source; shown from the name)
  • test_contextual_rewrite_with_history_calls_generator_and_returns_rewrite passed
    Contextual rewrite with history calls generator and returns rewrite (no description in the source; shown from the name)
  • test_contextual_rewrite_fails_open_on_generator_error passed
    Contextual rewrite fails open on generator error (no description in the source; shown from the name)
  • test_contextual_rewrite_empty_answer_falls_back_to_original passed
    Contextual rewrite empty answer falls back to original (no description in the source; shown from the name)
  • test_hyde_document_returns_hypothetical_passage passed
    Hyde document returns hypothetical passage (no description in the source; shown from the name)
  • test_hyde_document_fails_open_on_generator_error passed
    Hyde document fails open on generator error (no description in the source; shown from the name)
  • test_multi_query_variants_includes_original_first passed
    Multi query variants includes original first (no description in the source; shown from the name)
  • test_multi_query_variants_deduplicates_case_insensitively passed
    Multi query variants deduplicates case insensitively (no description in the source; shown from the name)
  • test_multi_query_variants_fails_open_to_original_only passed
    Multi query variants fails open to original only (no description in the source; shown from the name)
  • test_rrf_ranks_items_in_multiple_lists_higher passed
    Rrf ranks items in multiple lists higher (no description in the source; shown from the name)
  • test_rrf_deduplicates_by_chunk_id passed
    Rrf deduplicates by chunk id (no description in the source; shown from the name)
  • test_rrf_falls_back_to_document_id_chunk_index_when_no_chunk_id passed
    Rrf falls back to document id chunk index when no chunk id (no description in the source; shown from the name)
  • test_rrf_empty_lists_returns_empty passed
    Rrf empty lists returns empty (no description in the source; shown from the name)
  • test_ask_without_query_rewrite_has_no_query_rewrite_key passed
    Ask without query rewrite has no query rewrite key (no description in the source; shown from the name)
  • test_ask_with_contextual_rewrite_needs_session_id_to_do_anything passed
    No session_id means no history means contextual_rewrite makes no
  • test_ask_with_hyde_returns_hyde_document_field passed
    Ask with hyde returns hyde document field (no description in the source; shown from the name)
  • test_ask_with_multi_query_returns_variants_and_retrieves passed
    Ask with multi query returns variants and retrieves (no description in the source; shown from the name)
  • test_ask_final_answer_always_uses_original_query_not_rewritten passed
    The rewrite only affects what gets retrieved, never what's shown
  • test_ask_query_rewrite_cost_contributes_to_cumulative_spend passed
    The extra rewrite LLM call should be recorded into cumulative
test_reranking.py — 4 tests 3 passed, 0 failed, 1 skipped

No file-level description in the source.

  • test_rerank_true_calls_reranker_and_reorders passed
    Rerank true calls reranker and reorders (no description in the source; shown from the name)
  • test_rerank_false_never_calls_reranker passed
    Rerank false never calls reranker (no description in the source; shown from the name)
  • test_rerank_expands_candidate_pool_before_reranking passed
    rerank=True should retrieve top_k*4 candidates for the reranker
  • test_reranking_real_onnx_model_ranks_correctly skipped
    Real end-to-end test of the actual ONNX reranker (no mocking) -
test_retrieval.py — 14 tests 13 passed, 0 failed, 1 skipped

No file-level description in the source.

  • test_sparse_search_finds_real_keyword_matches passed
    Sparse search finds real keyword matches (no description in the source; shown from the name)
  • test_sparse_search_no_match_returns_empty passed
    Sparse search no match returns empty (no description in the source; shown from the name)
  • test_sparse_search_respects_metadata_filter passed
    Sparse search respects metadata filter (no description in the source; shown from the name)
  • test_dense_search_respects_embedding_dimension_mismatch passed
    A query embedding of the wrong dimension should be rejected or
  • test_dense_search_empty_embedding_returns_empty passed
    Dense search empty embedding returns empty (no description in the source; shown from the name)
  • test_hybrid_search_prefers_keyword_match_via_sparse_signal passed
    Dense (fake) embeddings carry no real semantic signal, so a
  • test_hybrid_search_combines_dense_and_sparse_result_sets passed
    Hybrid search combines dense and sparse result sets (no description in the source; shown from the name)
  • test_backend_reports_sparse_support passed
    PgVectorBackend (the default) genuinely supports sparse search -
  • test_retrieve_returns_chunks_without_generating_answer passed
    retrieve() must never call generation - verify by making the
  • test_retrieve_respects_top_k passed
    Retrieve respects top k (no description in the source; shown from the name)
  • test_retrieve_respects_metadata_filter passed
    Retrieve respects metadata filter (no description in the source; shown from the name)
  • test_retrieve_dense_only_when_hybrid_false passed
    Retrieve dense only when hybrid false (no description in the source; shown from the name)
  • test_retrieve_no_documents_returns_empty passed
    Note: NOT testing "no semantic match" with an ingested document -
  • test_retrieve_with_rerank_does_not_crash skipped
    rerank=True lazily constructs a RerankerService - just verify
test_sanitization.py — 10 tests 10 passed, 0 failed, 0 skipped

Pure unit tests for sanitization — no DB needed, but autouse fixtures still run (harmless, DB is up regardless).

  • test_sanitize_removes_null_bytes passed
    Sanitize removes null bytes (no description in the source; shown from the name)
  • test_sanitize_removes_control_chars_but_keeps_newline_and_tab passed
    Sanitize removes control chars but keeps newline and tab (no description in the source; shown from the name)
  • test_sanitize_empty_string passed
    Sanitize empty string (no description in the source; shown from the name)
  • test_sanitize_normal_text_unaffected passed
    Sanitize normal text unaffected (no description in the source; shown from the name)
  • test_detect_injection_risk_finds_known_pattern passed
    Detect injection risk finds known pattern (no description in the source; shown from the name)
  • test_detect_injection_risk_case_insensitive passed
    Detect injection risk case insensitive (no description in the source; shown from the name)
  • test_detect_injection_risk_no_match_on_clean_text passed
    Detect injection risk no match on clean text (no description in the source; shown from the name)
  • test_detect_injection_risk_empty_text passed
    Detect injection risk empty text (no description in the source; shown from the name)
  • test_check_length_within_limit passed
    Check length within limit (no description in the source; shown from the name)
  • test_check_length_exceeds_limit passed
    Check length exceeds limit (no description in the source; shown from the name)
test_smoke.py — 2 tests 2 passed, 0 failed, 0 skipped

Minimal smoke test — proves the fixtures, fake providers, and real Postgres schema all work together before building out the full suite.

  • test_ingest_and_ask_roundtrip passed
    Ingest and ask roundtrip (no description in the source; shown from the name)
  • test_schema_actually_has_pgvector_and_halfvec passed
    Schema actually has pgvector and halfvec (no description in the source; shown from the name)
test_streaming.py — 6 tests 6 passed, 0 failed, 0 skipped

Tests for ask_stream() — sync generator, incremental piece yielding, and history storage once streaming completes.

  • test_ask_stream_yields_and_assembles_full_answer passed
    Ask stream yields and assembles full answer (no description in the source; shown from the name)
  • test_ask_stream_stores_full_answer_to_history_when_session_id_given passed
    Ask stream stores full answer to history when session id given (no description in the source; shown from the name)
  • test_ask_stream_no_session_id_stores_nothing passed
    Ask stream no session id stores nothing (no description in the source; shown from the name)
  • test_ask_stream_respects_metadata_filter passed
    Ask stream respects metadata filter (no description in the source; shown from the name)
  • test_ask_stream_rerank_true_calls_reranker passed
    Ask stream rerank true calls reranker (no description in the source; shown from the name)
  • test_ask_stream_rerank_false_never_calls_reranker passed
    Ask stream rerank false never calls reranker (no description in the source; shown from the name)
test_structured.py — 11 tests 11 passed, 0 failed, 0 skipped

Tests for structured/JSON output mode (v0.8.0). Split into two groups: unit tests for ragleap.structured's parse/validate logic (including a forced no-jsonschema fallback path via monkeypatching, since jsonschema IS installed in this test environment), and plumbing tests proving response_format= flows correctly through ask() -> generate_answer() -> the provider call and back, using the existing fake_call_provider fixture (no live network calls - live Gemini/ Anthropic verification was done manually this session, documented in CHANGELOG, since it needs a real committed API key CI doesn't have).

  • test_parse_and_validate_valid_json_matching_schema passed
    Parse and validate valid json matching schema (no description in the source; shown from the name)
  • test_parse_and_validate_valid_json_not_matching_schema passed
    Missing the required 'name' field - valid JSON, invalid per schema.
  • test_parse_and_validate_malformed_json_returns_none passed
    Parse and validate malformed json returns none (no description in the source; shown from the name)
  • test_parse_and_validate_object_valid passed
    Mirrors Anthropic's tool-use path, which hands back an already-
  • test_parse_and_validate_without_jsonschema_installed passed
    Forces the basic-type-check-only fallback path by monkeypatching
  • test_parse_and_validate_invalid_schema_itself passed
    A malformed schema (invalid 'type' value) should be caught as a
  • test_ask_without_response_format_has_no_structured_fields passed
    Backward compatibility - existing callers who never pass
  • test_ask_with_response_format_returns_structured_fields passed
    Uses a permissive schema matching fake_call_provider's canned
  • test_ask_with_response_format_catches_real_schema_mismatch passed
    SCHEMA requires a "name" field that the fake fixture's canned
  • test_ask_with_response_format_array_schema passed
    Confirms the fake fixture (and real plumbing) respects the
  • test_ask_with_response_format_and_guardrails_still_work passed
    response_format's JSON-string answer still passes through the
test_weaviate_backend.py — 12 tests 12 passed, 0 failed, 0 skipped

Tests for WeaviateBackend - NOT live-verified against a real Weaviate instance (same honest caveat as PineconeBackend). Mocks the actual Weaviate client entirely; every attribute path verified against the actual installed weaviate-client==4.22.0 package's real source during development (self.collections/self.data/self.query are real instance attributes set in __init__, not class-level methods - confirmed by reading the actual source, not assumed).

  • test_requires_persist_directory passed
    Requires persist directory (no description in the source; shown from the name)
  • test_collection_name_normalized_to_uppercase_first_letter passed
    Collection name normalized to uppercase first letter (no description in the source; shown from the name)
  • test_vector_key_construction passed
    Vector key construction (no description in the source; shown from the name)
  • test_deterministic_uuid_is_stable passed
    Deterministic uuid is stable (no description in the source; shown from the name)
  • test_insert_document_and_list_documents passed
    Insert document and list documents (no description in the source; shown from the name)
  • test_insert_chunk_calls_data_insert_with_correct_args passed
    Insert chunk calls data insert with correct args (no description in the source; shown from the name)
  • test_search_dense_converts_distance_to_similarity_score passed
    Search dense converts distance to similarity score (no description in the source; shown from the name)
  • test_search_dense_skips_orphaned_objects passed
    Search dense skips orphaned objects (no description in the source; shown from the name)
  • test_delete_document_calls_delete_by_id_for_each_chunk passed
    Delete document calls delete by id for each chunk (no description in the source; shown from the name)
  • test_search_dense_normalizes_cosine_distance_to_unit_range passed
    Regression test for the similarity_score normalization fix.
  • test_supports_sparse_is_false passed
    Supports sparse is false (no description in the source; shown from the name)
  • test_weaviate_backend_importable_from_vectorstores passed
    Weaviate backend importable from vectorstores (no description in the source; shown from the name)
test_weaviate_backend_live.py — 6 tests 0 passed, 0 failed, 6 skipped

Live tests for WeaviateBackend against a real running Weaviate instance. Gated on WEAVIATE_TEST_URL (e.g. "http://localhost:8081") and WEAVIATE_TEST_GRPC_PORT (defaults to 50051), same pattern as the OpenSearch/Upstash/Qdrant live-gated tests - skip cleanly (not fail) when no real instance is configured.

  • test_init_schema_is_idempotent skipped
    Init schema is idempotent (no description in the source; shown from the name)
  • test_search_dense_orders_and_normalizes_score skipped
    Search dense orders and normalizes score (no description in the source; shown from the name)
  • test_search_dense_metadata_filter skipped
    Search dense metadata filter (no description in the source; shown from the name)
  • test_list_documents_and_get_filename skipped
    List documents and get filename (no description in the source; shown from the name)
  • test_delete_document_removes_vectors skipped
    Delete document removes vectors (no description in the source; shown from the name)
  • test_supports_sparse_is_false skipped
    Supports sparse is false (no description in the source; shown from the name)
test_web_import_error.py — 1 tests 1 passed, 0 failed, 0 skipped

No file-level description in the source.

  • test_import_error_message_includes_the_real_cause passed
    Import error message includes the real cause (no description in the source; shown from the name)

Compiled by script from the public repository, its CI logs and the GitHub API. To correct something, open an issue.

All tests