Compare the workload. Then the price.
A price per million tokens tells you something, but it doesn't tell you what your feature costs. Start with a repeatable set of real tasks and build the comparison from there.
Use the same jobs
Choose representative requests: a short classification, a document summary, and your hardest customer workflow. Keep the desired outcome the same for every model. Save results so a cheap but incorrect answer doesn't win by accident.
Count the full request
Include input and output charges, billed reasoning where applicable, tool calls, retries, and hosting. Add latency and rate limits beside the totals. Record the date and exact model ID whenever you collect pricing.
Choose on cost per useful result
Divide spend by the number of outputs that meet your acceptance criteria. A more expensive model can be cheaper if it succeeds in one attempt. A smaller model can be enough for simpler jobs. Link your spreadsheet to current provider rate cards instead of freezing today's prices into a long-term plan.