Quality evals

Validate translation quality across models and check how well voice recognition hears you.

Translation quality (chrF + LLM judge)

Runs the built-in DE→UK / EN→UK sentence set through the selected models. Scores each output with chrF against a reference translation; optionally an LLM judge rates adequacy 1–5.

Voice recognition validation (WER)

Read the phrase aloud; Loqui compares what the recognizer heard against the reference and computes the word error rate. Lower is better (0% = perfect recognition).

🇩🇪 Read aloud in German · 1/6

Ich möchte morgen früh nach München fahren.