Benchmark
Every configuration runs the same 30-question golden set through the same pipeline the API serves. Each config changes exactly one thing from the one above it, so a difference in the table is attributable to that change and nothing else.
loading…