Grok 4.6 scored 61 on the AI ranking. What does that change for you?
Every week a message lands on my WhatsApp with a ranking screenshot and the question: "Diego, should I switch to this one?". This time the screenshot was about Grok 4.6, xAI's new model, which scored 61 points on the Artificial Analysis Intelligence Index. It's a good number. It puts the model in the elite group, competing with the best on the market. But before you pull out your credit card, it's worth understanding what that 61 measures, what it doesn't measure, and how it connects to your revenue.
What the Artificial Analysis Intelligence Index is
Artificial Analysis is an independent company that tests AI models in the most comparable way possible. They take a set of standardized tests: reasoning, math, programming, general knowledge, the ability to follow long instructions, tool use. They run everything under the same conditions and combine it into a single score, from 0 to 100.
Grok 4.6 scored 61. For scale: today's top models sit in the low 60s. Good, cheap models land between 40 and 50. Models from a year and a half ago, which felt like magic at the time, sit below 40.
So yes, it's a top-shelf result. xAI went from "the Twitter chatbot" to a real contender in a little over two years.
But there's a detail the screenshot doesn't show.
What the score hides
The index is an average. And averages hide things.
One model can be brilliant at math and mediocre at following formatting instructions. Another can be excellent at writing text in Portuguese and weak at code. Both can end up with the same final score.
In the case of Grok 4.6, what stood out in the report was the combination of two things: high score and high speed. It answers fast and thinks well. That's rare. Usually, the more a model "thinks" before answering, the slower and more expensive it gets.
But the report also shows where it spends more tokens than the competition to reach the same answer. In business-owner language: the bill at the end of the month can be bigger than the price per million tokens suggests, because the model talks too much to solve the same task.
Rankings measure the test. Your business measures the result.
That distinction is worth more than any benchmark.
Why the best model on the ranking can be the worst choice for you
Let me tell you a real case, with names changed.
A dental clinic reached out to me earlier this year. They wanted to automate message triage on WhatsApp: identify whether it's a booking, an insurance question, a quote request or a complaint, and route it to the right place. Volume: about 400 messages a day.
One of the partners had seen a ranking and wanted "the best". The best at the time cost around 15 dollars per million output tokens.
We tested three models with 200 real messages from the clinic. Result:
- Top-tier model: 97% of classifications correct.
- Mid-tier model: 96% correct.
- Cheap model: 94% correct.
The difference between first and second was two messages out of 200. The cost was seven times higher. We went with the mid-tier. Savings of over R$ 600 a month in a small operation, with no noticeable loss.
Classifying a patient's message is not the same as solving an olympiad math problem. The ranking tests the second. You need the first.
When the ranking really matters
I don't want you to leave here thinking benchmarks are useless. They matter in three situations:
When the task is genuinely hard. Contract analysis with cross-referenced clauses, complex code, reports that require combining several sources. That's where the difference between 55 and 61 shows up in practice, and it shows up as expensive mistakes.
When you're going to use agents. An agent is a system that runs multi-step tasks on its own: opens a spreadsheet, queries an API, makes a decision, moves on. A mistake at step 2 contaminates everything that follows. Stronger models make fewer mistakes along the way, and that adds up.
When you have no way to test first. If you can't run a pilot with your own data, the ranking is the best proxy there is. Better than the vendor's marketing, for sure.
Outside of that, the ranking is just curiosity. Good curiosity, but curiosity.
How to pick the right model without becoming an expert
Here's the process I use with clients, boiled down to the essentials:
- Define the task precisely. "Use AI in customer service" is not a task. "Classify messages into 5 categories" is.
- Gather 100 to 200 real examples. From your data, not examples off the internet. Your customers' messages, your documents, your spreadsheets.
- Test 3 models from different price tiers. One top-tier, one mid-tier, one cheap.
- Measure accuracy and cost side by side. If the cheap one gets 94% and the expensive one gets 97%, ask: how much does each mistake cost? Sometimes it's worth paying. Most of the time, it isn't.
- Make switching easy. Use a layer that lets you change models without rewriting everything. Three months from now the ranking will change again and you don't want to rebuild the system.
That last point is what saves the most money in the long run. Today's Grok 4.6 is tomorrow's mid-tier model. I've seen it happen with every model that was once "the best".
So, is Grok 4.6 worth it?
My honest opinion: it's worth testing if you already have something running and want to compare. It's not worth switching on impulse.
The model is competent, fast, and xAI has been investing heavily in infrastructure. But it's a young company, with a history of abrupt changes in direction, and the tooling ecosystem around it is still smaller than its competitors'. For a company that depends on the system working every day, vendor stability weighs as much as the ranking score.
If your AI operation depends on a number that changes every week, the problem isn't the model. It's the lack of a method for choosing.
That's exactly the kind of decision I help companies make: test with real data, measure what matters, and build automation that doesn't break when the ranking changes. If you'd like to talk about your case, get to know me better here.
LinkedIn summary
Grok 4.6 scored 61 on the AI ranking. So what? Every week I get a ranking screenshot on WhatsApp with the same question: "Diego, should I switch?" Rankings measure the test. Your business measures the result. I tested three models with 200 real messages from a clinic: the top one got 97% right, the mid-tier 96%, the cheap one 94%. The most expensive cost seven times more for a two-message difference. If your AI operation depends on a number that changes every week, the problem isn't the model. It's the lack of a method for choosing. Want to know how to choose with real data instead of screenshots? Message me. #ArtificialIntelligence #Automation #AIforBusiness #Grok #Technology