Human experts remain essential for checking clinical AI outputs
In a Rwanda-based evaluation, LLM judges produced consistent and far cheaper ratings of clinical AI responses but matched local clinician ratings on only a minority of assessment criteria. They consistently…