Avalia
Independent evaluations of AI models in Brazilian Portuguese. We compare limits and risks where they matter. In real Brazil.
Loading results…
How we evaluate
IDJÉ Avalia rankings run on Régua: our proprietary multiple-choice evaluation harness. Same answer reading, same scoring rule, with care for Brazilian Portuguese.
Question
Each test item is a multiple-choice question with two to ten options.
Answer reading
Régua identifies which option the model chose, the same way every time.
↓ offline scoring, no LLM judge ↓
What counts
Splits valid answers, refusals, invalid replies and errors. Clear what enters each rate.
Test metric
Accuracy, prudence, bias: each evaluation defines what to measure. Counting follows Régua.