Before buying an AI service because it has an impressive energy grade, I would want to know which job earned it. Transcribing a recording, sorting support requests and writing a long answer are different tasks. A label that ignores the difference would make shopping easier in the way a missing price tag makes arithmetic easier. AI should carry an energy label, but the useful version needs to identify the work, the required quality and the conditions of the test.
We already have a start. AI Energy Score compares models under controlled conditions. The next challenge is making that information useful without turning a limited measurement into a promise about the entire product.
Start with the job being measured
The first line of a label should describe a task a buyer recognizes. It should tell someone choosing a transcription service that the test involved transcription, then make clear how well the system performed. An energy-saving system that produces unusable transcripts has saved the supplier some electricity and handed the work back to the customer.
That is why a quality requirement belongs beside the energy measurement. The buyer should define an acceptable result, then compare the resources needed to deliver it. There may be several acceptable choices. The most accurate system does not automatically deserve every contract, and the least demanding one does not automatically meet the need.
AI Energy Score’s documented tests control the task, hardware and batching. The score focuses on GPU energy. Its guidance asks users to weigh that result against how well and how fast a model does the work, among other measures. Those controls let customers repeat a test and check a supplier’s claim that a model is efficient.
A label could state:
- How much electricity a model used for a set of jobs.
- How long the inputs and outputs were.
- Which model version was used.
- Which reasoning setting was used.
A buyer could then ask whether those conditions resemble the intended work. That would tell a buyer more than a debate about which average is greener when neither has a clear source.
The test would also need to count failed jobs. Otherwise a supplier could look efficient by measuring quick answers while leaving the customer to discover how often those answers need correction. The label should record the quality threshold and failure rate.
A grade needs a number beside it
A star or a letter helps people make a quick comparison. It also invites them to stop reading. I would keep the grade, because a label too elaborate to use has its own problems, but put the measured energy beside it in ordinary units.
AI Energy Score’s stars are relative to a task and a particular group of tested models. Text-generation ratings also use model classes. The ratings can change when the comparison group changes. A high rating therefore needs its context. It cannot safely be read as a promise that the model uses less electricity than every lower-rated model doing some other job.
The number lets the buyer track energy use as the ratings change. If a product moves from a higher grade to a lower one because its peers improved, the number helps explain that change. If the product itself starts using more electricity for the same work, the number makes that visible too.
There must also be a short explanation of what was measured. GPU energy counts one part of the equipment. MLPerf’s power measurements cover the tested system at the wall and apply to the accompanying benchmark. Neither measure, on its own, counts a job’s share of all the power used at the site. A buyer should be able to see what each test counts without having to read a long report.
A compact label might have a plain statement underneath the grade explaining:
- Which equipment was counted.
- When the test ran.
- Who checked it.
Further details could go in the test report, while the short version would make only claims the test supports.
When a service changes its models
An AI product can keep its name while changing how it does the work. That makes a permanent badge a poor fit.
MLCommons’ September 16 release of MLPerf Inference v6.1 added tests for end-to-end retrieval-augmented generation and AI agents running on user devices. These reflect work that happens in stages. A system might find source material, rank it, and send it to a language model before producing the answer. An agent may make several calls while working toward a result.
Miro Hodak, a co-chair of the benchmark group, argues that testing has to reflect those systems. He is right, and his point is a serious challenge to the simplest version of an AI label. The user bought the completed service. A favorable test of one component does not establish the energy cost of everything the service does.
I would give a model’s lab rating its own label, apart from the report supplied by a hosted service. The first helps compare models in specified conditions. The second should explain the mixture of models and other steps used in production. If the supplier changes that mixture enough to affect the result, its published account should change too.
Mistral has raised a related concern: a poorly designed standard can reward performance on the test while missing ordinary use. This should worry anyone who has watched a product specification become a sales argument. A supplier will have an incentive to look good under the published test conditions. We should make those conditions useful and leave room for independent checks with representative work.
The answer cannot be an unannounced test that nobody can reproduce. It should be a published method, sensible workload samples and a way to challenge results that do not resemble the product being delivered. Readers should be able to see what changed. Otherwise the buyer loses the ability to tell whether the service improved or the test became easier.
Make the label useful to buyers
Buyers also need a say in the label’s design. In September, the European Commission published a study of proposed energy labels for computers. Its recommendations addressed different types of use and the need to explain computing performance. The study concerned computers, so it cannot tell us what AI buyers would understand. An AI label should get its own test with the people who would use it.
The label should also stay modest about what it certifies. Lower measured electricity for a useful task does not establish good privacy practices, reliable answers or a small water footprint. Those questions deserve their own evidence. The energy label earns trust by being clear about its subject.
AI Energy Score’s December 2025 update reported that Salesforce was using energy data in its own model tests. That is a promising place for this work to begin. A buyer can ask for results it can compare when it chooses a product, then require a new report as that product changes.
The individual opening a chat window often cannot see which model will answer. Asking that person to make an informed environmental choice without giving them any control amounts to assigning homework with no materials.
A purchasing team has a better opportunity. It can bring a sample of the work it needs done, specify what counts as success and ask the supplier to reproduce the result under an agreed test. A useful label would make that request routine. The next person using the service could get on with the job, knowing someone had asked how much electricity it took to finish it.





