Three providers. One useful test.
Picking a provider from a general leaderboard is tempting. Running the models on your own product's jobs gives you a much better basis for the choice.
Agree on what good means
Choose a test set with ordinary and difficult cases. Use the same expected outcome, even if each provider needs a slightly different prompt format. Keep the grading criteria fixed and review outputs without provider labels when possible.
Include production constraints
Measure response time, available context, rate limits, and the features your integration needs. Test the actual service route and account tier. Record exact model IDs so the comparison can be repeated.
Allow more than one winner
Different features may justify different models. Keep routing simple enough to maintain and monitor the quality of any fallback. Compare the final bill after retries and eligible credits, with current pricing linked in your working notes.