Skip to main content

Goal

Evaluate model predictions for a text QA task and compute accuracy.

Full example

Optional: repeat for uncertainty