RagLeap
Blog

How RagLeap detects the language of very short queries

A chat question like "Quale API usa RagLeap?" has only a handful of words. That is very little evidence for a language detector, and closely related languages such as Italian, French, Spanish and Portuguese are easy to confuse. This post explains how RagLeap Core handles that, what the contributor who built it measured, and how far those numbers can be trusted.

The problem

The issue behind this work (#24) reported that a French query such as "Quel framework utilise RagLeap Core?" could be detected as Italian. The earlier implementation also returned the default language for any input shorter than 10 characters, so short Japanese, Korean and Chinese queries never reached script-based detection at all.

What changed

The fix is pull request #452 by Hardik Anand, merged on 2026-09-20 and included in ragleap-core v0.7.5. It touches three files: core/language.py, requirements.txt and a new tests/original_benchmark.py. It has four parts:

  • FastText as the primary detector. The fast-langdetect package provides a pretrained FastText model, so no large custom classifier has to be maintained.
  • A character n-gram fallback. A lightweight second detector uses character n-grams of 2 to 5 characters, trained on 224 labeled examples across 9 languages. It gives an independent signal for difficult short queries.
  • Confidence-gated disagreement handling. If both detectors agree, that answer is used. If they disagree, the n-gram result overrides FastText only when its confidence reaches 0.55; otherwise FastText wins. The threshold is the environment setting LANGUAGE_DETECTION_NGRAM_CONFIDENCE_THRESHOLD, with 0.55 as its default.
  • Short-script handling. Text shorter than 10 characters is no longer rejected before script inference, so queries such as こんにちは (ja), 안녕하세요 (ko) and 你好 (zh) can still be identified by their script.

Where FastText alone went wrong

On the benchmark, FastText alone got three of 32 queries wrong, and the n-gram detector got all three right:

  • "Quale API usa RagLeap?" was detected as German; the correct answer is Italian.
  • "Qual framework o RagLeap usa?" was detected as English; the correct answer is Portuguese.
  • "Como funciona o RagLeap?" was detected as Spanish; the correct answer is Portuguese.

The benchmark

The benchmark is 32 short queries across 9 languages: French, Italian, Spanish and English (5 each), Portuguese and German (3 each), and Japanese, Korean and Chinese (2 each). It is separate from the 224 training examples.

ApproachCorrectAccuracy
Original langdetect14 / 3243.75%
FastText alone29 / 3290.63%
FastText + n-gram fallback + routing + short-script handling32 / 32100%

To reproduce it, run python -m tests.original_benchmark from the ragleap-core repository. The expected last line is Accuracy: 32/32 = 100.00%.

What the 100% does and does not mean

A perfect score on 32 queries is encouraging, but it is not a general accuracy figure, and the pull request says so itself. Three things to keep in mind:

  • The sample is small: 32 queries, and only two each for Japanese, Korean and Chinese.
  • The 0.55 threshold was chosen by sweeping values against this same benchmark (0.50 scored 31/32, 0.55 scored 32/32). The number was not tested on separate held-out queries.
  • The confidence value is a routing signal, not a calibrated probability. A score of 0.8 does not mean an 80% chance of being right, because short queries can produce confident but wrong predictions.

So read the result as evidence that the approach fixes the reported failure cases and the closely related ones. It is not a promise about every language or every kind of query.

What the people involved said

Hardik wrote about this work on LinkedIn as his fifth merged open-source pull request, and described the main lesson as handling the grey zone when two models disagree, and deciding when the evidence is strong enough to trust one signal over the other. Read his post on LinkedIn.

Separately, Dragan Petkovic, an AI solution builder, named RagLeap Core in a German-language LinkedIn post about open-source alternatives, describing it as offering 46 role-based AI employees that are self-hosted and MIT-licensed, with no license key. Read his post on LinkedIn. We link to it as an outside mention; we do not vouch for the other claims in that post. Both posts are also shown on the homepage.

Take part

RagLeap is built in the open. If you find a query that is detected wrongly, open an issue with the text and the language you expected. For code changes, the contributing guide explains how, and every contributor is listed on the contributors page.

Back to the blog