Reproducible per-example results from bench/run_eval.py. No numbers here are estimated — this table is empty until a real run has been committed. See README for the harness and how to add your own hardware's results.
| Checkpoint | Language | Metric | Value | n | Hardware | Date |
|---|