Energy methodology
Every message in TreadLightlyAI carries an energy figure. Here is how it is worked out, how accurate it is, and what it leaves out.
Last updated 2026-08-27
The short version
- The formula
energy_wh = total_tokens × rate × 1.2Input and output tokens added together, times a per-token rate in microwatt-hours, times 1.2 for infrastructure overhead.- Measured or estimated
- One model reports its own energy use. Everything else is estimated, and every estimate in the product is marked with a tilde and set in italics.
- How accurate
- Order of magnitude. Good enough to tell a cheap conversation from an expensive one. Not good enough to carry into a carbon report.
- Confidence intervals
- None. We publish point estimates and we have not computed a confidence interval, because doing it honestly needs repeated measurements of the models nobody measures.
- Not counted at all
- Training, water, carbon for models we do not host, and diagram rendering.
- Last reviewed
- 4 August 2026. Nothing changed as a result.
What the formula means
Tokens, not words. Input and output tokens are added together and multiplied by a single rate. The rates are blended, meaning a token you send and a token the model writes are charged the same, because none of the measured data we have separates the two.
The 1.2 overhead multiplier. Inference is not the only thing drawing power when you send a message. The multiplier covers networking, cooling, storage, and load balancing, and is in line with the overhead figures used in ordinary data center accounting. It is applied to every estimate on this page.
Because energy scales with tokens, a longer answer costs more than a short one, and re-sending a long conversation costs more than starting a fresh one. That is the behaviour the figure is meant to make visible.
Rates by model
| Model | Rate (μWh/token) | With overhead | Source | Why that number |
|---|---|---|---|---|
| Green Tiny and Green Small | 90 | 108 | Estimated | Bounded by the two measurements below |
| Green Large | n/a | 115.31 observed | Measured | Reported by the provider, not estimated at all |
| Claude Haiku | 65 | 78 | Estimated | Above the 56.65 measured for a retired small model |
| Claude Sonnet | 120 | 144 | Estimated | Above Green Large, spaced by published API prices |
| Green Docs (document search) | 50 | 60 | Estimated | Indicative only, and the widest uncertainty on this page |
Worth saying plainly, because it cuts against our own marketing: our estimate for Claude Haiku is lower than our estimate for the green models. The green tiers are chosen for where their electricity comes from and who trained the model, not for using less of it.
The full reasoning behind each rate
- Green Tiny and Green Small
- One engine (Gemma 4 31B Instruct at Infomaniak) serves both tiers, so both carry the same per-token rate. The figure is bounded by the two measurements below rather than scaled from parameter count, and it is deliberately round because the inputs support one significant figure.
- Green Large
- GreenPT reports the energy of each request in its API response, so this model is not estimated at all. The figure shown is the average across 95 calls and about 89,000 tokens, and it is what the other rates are calibrated against.
- Claude Haiku
- Set above the 56.65 measured for a small model on GreenPT that has since left the lineup. GPU inference in US data centers is expected to cost more per token than the CPU-prioritized Swiss infrastructure that measurement came from.
- Claude Sonnet
- Set above the 115.31 measured for Green Large. The gap between Haiku and Sonnet follows the ratio of their published API prices, used as a stand-in for compute intensity, at roughly 1.85 times Haiku.
- Green Docs (document search)
- Scaled from a chat baseline that has since been retired, by active parameter count. Embedding is a single encode pass rather than word-by-word generation, so this rate is indicative and carries wider uncertainty than the rows above.
Methodology version:
Search, page reads, and documents
| Action | Estimate | Basis |
|---|---|---|
| Web search | 0.0108 Wh per search | A placeholder with a stated reason, not a measurement. Pegged to what a short chat message of about a hundred tokens costs on our default model. |
| Reading a page | 0.005 Wh, plus 0.00003 Wh per kilobyte | Connection setup and parsing, plus transfer at roughly 0.03 kWh per gigabyte. A typical page lands near 0.0074 Wh. |
| Documents you upload | 50 µWh per token, 60 with overhead | A separate model makes the document searchable. Token counts come from the provider, or from a count of about four characters per token when it reports none. |
The thinking the model does about a page it reads is not counted here. That is already billed through the message's own tokens, and counting it twice would be its own transparency failure.
How uncertain this is
- The one spread we can observe is wide. Two similar models on the same host differed by roughly four to six times per token.
- Model size does not predict energy. We used to scale rates by parameter count and dropped it, because it fails against every measurement we can test it on.
- The precision is one significant figure. The 90 rate is round on purpose.
- We have shipped a wrong number before. The per-search estimate was a thousand times too low until 4 August 2026.
- Small totals are the least reliable. For a single short message the estimate and the rounding are the same size.
The longer version of each
The observable spread. Two models from the same family, of similar size, measured on the same host, differed by roughly four to six times per token. That comparison rests on a single run of one of them, so treat it as an indication of scale rather than a figure. It is still the best answer we have to how far off an estimate can be.
Parameter count. We dropped the size-based rule because it fails on every pair we can test against a measurement: it does not put the measured rates in the right order. Anyone estimating AI energy from parameter counts alone, including our own earlier method, is guessing.
Precision. Our inputs support one significant figure. A number like 95.9 would imply an accuracy we do not have.
The correction. The per-search energy estimate was a thousand times too low for a period, a unit slip corrected on 4 August 2026. We name it here because a page about honest measurement that hides its own corrections is not one.
Small totals. The figure is more useful as a relative signal, comparing one conversation or one model against another, than as an absolute quantity.
What we do not estimate
These are gaps, named rather than filled. A blank is more useful than a number we made up.
- Carbon, for models we do not host. Energy scales with tokens; carbon depends on the grid behind the data center, hour by hour.
- Water. No provider reports it, and it is widely considered the largest missing measure in AI environmental accounting.
- Training. Every figure here is the cost of running a model, not of building it.
- Diagram rendering. The message that writes a diagram is counted. The step that turns it into an image is not.
Why each of these stays blank
Carbon. GreenPT reports emissions alongside energy and we pass those through. Anthropic publishes neither its data center locations nor its power sourcing, so the field stays empty for Claude calls rather than being guessed.
Water. We would rather name the gap than invent a figure. Infomaniak's Geneva data center draws no water for cooling, which is a property of that site rather than a number we can attribute to your message.
Training. Training is a large share of the total footprint of any AI system, and we are not aware of a large model trained on renewable power. What we show you is the smaller half of the picture, and we would rather say so than let the number stand for more than it covers.
The everyday comparisons
Watt-hours mean nothing to most people, so the product also shows a plain-language comparison. Every screen draws from this one table, so two screens can never describe the same energy differently.
| Energy | What we say |
|---|---|
| Under 1 mWh | Less than a second of a 10W bulb |
| 1 to 10 mWh | A few seconds of a 10W bulb |
| 10 to 50 mWh | Roughly 10 seconds of a 10W bulb |
| 50 to 200 mWh | About a minute of a 10W bulb |
| 200 mWh to 1 Wh | A few minutes of a 10W bulb |
| 1 to 5 Wh | Up to half an hour of a 10W bulb |
| 5 to 15 Wh | About one phone charge |
| 15 to 40 Wh | A few phone charges |
| 40 to 120 Wh | An hour or two of laptop use |
| 120 to 500 Wh | Several hours of laptop use |
| 500 Wh to 2 kWh | Roughly a full day of laptop use |
| Above 2 kWh | More than a day of nonstop laptop use |
Assumptions: a 10W LED bulb, a phone charge of about 15 Wh, and a laptop drawing about 50 W in active use.
The comparison we removed
We used to compare against a web search. The figure everyone quotes for a search traces back to a 2009 blog post and no longer describes anything current, and a comparison that fails a common-sense check does more damage than no comparison at all.
Where the numbers come from
Two real measurements sit underneath every estimate on this page, both from our own production logs of GreenPT calls, where the provider returned the energy it actually used.
- 115.31 μWh/token for Green Large, across 95 calls and about 89,000 tokens. Still in the lineup.
- 56.65 μWh/token for a smaller GreenPT model, across 85 calls and about 45,000 tokens. It left the lineup in August 2026; the measurement is still a valid reference point.
Every estimated rate is placed relative to those two numbers. The resulting ordering is a constraint we chose and are disclosing, not a result the method derived on its own.
Why most of the lineup is estimated rather than measured
Green Large runs at GreenPT, which returns the energy of each request in the API response, so its numbers are measurements passed straight through. Infomaniak, which hosts our default model, reports no per-request energy. Anthropic reports none for Claude. Since August 2026 that includes the model most conversations actually use, so most of what you see in the product is an estimate.
That is a step backwards in data quality and it was a deliberate trade. The default model moved to the host with the strongest environmental record of any provider we use: Swiss renewable power, all of the data center's waste heat recovered into Geneva's district heating network, and no water drawn for cooling. None of that turns an estimate into a measurement, which is why every estimated figure carries a tilde. If Infomaniak starts reporting per-request energy we will use it, and say when we switched.
How the estimates are bracketed, and what we do not use
Provider-reported energy is real data center consumption, wall-clock time against power draw, which already includes cooling, networking, and storage. Published academic GPU benchmarks measure the chip and not the building, so they run low. We use one of them, Luccioni et al. (2023), only as an order-of-magnitude sanity check, never as a rate.
There is one supporting signal for the 90 figure that does not come from ordering. A same-family model at GreenPT measured several times the small-model baseline per token, and Gemma 4 31B at Infomaniak serves roughly 2.3 times faster than that baseline model in our own timing probes, at about 105 tokens per second against about 46. At comparable rack power, faster serving means fewer watt-seconds per token. Those two facts together bracket a range the 108 effective figure sits inside.
The rates get a scheduled review roughly every three months. The last one was 4 August 2026: no provider had begun publishing model-level per-token energy, the third-party estimates available were per-query rather than per-token, and the general inference literature placed per-token energy inside a range these rates already sit in. Nothing changed as a result.
Checking our work
- The AI Transparency Card is the one-glance summary: models, hosts, data center grades, safety scores, and the commitments we are bound to.
- How we compare puts our model roster and its energy rates next to what other assistants publish, which is usually nothing.
- Inside the product, every conversation shows its own running total, and your account page breaks the month down by model and by whether the power behind it was renewable.
We will never claim to be carbon neutral or zero emission without data that supports it. Measured where we can, estimated where we cannot, labelled either way.