Home / Guides / Performance per Watt for Local AI: Why Efficiency Matters

guide

Performance per Watt for Local AI: Why Efficiency Matters

Updated 2026-09-17

The fastest GPU is not always the cheapest to operate. This guide shows how to compare local AI hardware using energy per task, total cost of ownership, cooling, maintenance, and resale value.

Performance per watt tells you how much useful AI work a GPU delivers for the electricity it consumes. For frequent local inference, it can matter as much as peak speed—but it should not be used alone.

The practical comparison is:

Useful work per watt = throughput / power draw

For a specific workload, a more useful metric is:

Energy per task (kWh) = wall power (W) × task runtime (seconds) / 3,600,000

A GPU that is slower but uses much less power may have lower operating cost. A faster GPU can still be the better purchase if its higher throughput saves enough time or allows you to complete more work during your normal usage window.

The right decision depends on five cost categories:

  • Purchase price
  • Electricity
  • Cooling and system overhead
  • Maintenance and downtime
  • Resale value

Performance per watt is a workload metric

Performance per watt is not a permanent characteristic of a GPU. It changes with:

  • Model architecture and size
  • Quantization format
  • Context length
  • Batch size
  • Sequence length
  • Software stack and drivers
  • GPU utilization
  • CPU and storage bottlenecks
  • Whether the workload is inference, fine-tuning, or training

Compare GPUs using the same model, precision, prompt or dataset, batch size, and measurement method. For local inference, measure wall-clock throughput such as tokens per second or completed jobs per hour.

Examples:

Tokens per second per watt = tokens per second / average wall power (W)

Jobs per hour = 3,600 / average seconds per job

Jobs per watt-hour = jobs per hour / average wall power (W)

The last metric is particularly useful because it expresses useful work relative to energy rather than instantaneous power.

GPU power is not the same as system power

A GPU's reported board power does not include every watt used by the computer. The actual electricity cost may also include:

  • CPU and motherboard
  • Memory
  • Storage
  • Fans and pumps
  • Power-supply losses
  • Displays or attached peripherals
  • Networking equipment
  • Cooling equipment

For cost calculations, wall power measured with a plug-in power meter is preferable. If you only have a GPU power estimate, label the result as an estimate and add a system-overhead assumption.

Build a total cost model

A simple local AI cost model separates fixed costs from variable operating costs.

Net hardware cost = purchase price - expected resale value

If you plan to use the system for a defined period:

Annualized hardware cost = (purchase price - expected resale value) / useful life in years

Energy cost is:

Electricity cost = energy used (kWh) × electricity price ($/kWh)

For a repeated task:

Energy per task (kWh) = average wall power (W) × runtime (seconds) / 3,600,000

Operating cost per task = energy cost per task + cooling cost per task + maintenance cost per task

A broader total-cost calculation is:

TCO = hardware cost - resale value + electricity + cooling + maintenance

And for a workload:

Cost per task = TCO / completed tasks

This model can be calculated over a month, year, or the expected service life of the system. Use the period that matches your decision. A person running a model for a few hours each week should not evaluate hardware like an always-on inference server.

The five costs that matter

1. Purchase cost

Purchase price is usually the largest cost for moderate-use systems. It is also the easiest cost to see, which can cause buyers to overemphasize electricity savings.

A more useful number is the net hardware cost:

Net hardware cost = purchase price - resale value

A GPU that costs more initially may retain more resale value, reducing the effective cost of ownership. Conversely, a low-cost used GPU may be attractive even if it is less efficient, provided it meets your performance and memory requirements.

2. Electricity

Electricity matters most when the system runs frequently, operates at high power, or is located where energy prices are high.

Annual energy use can be estimated as:

Annual energy (kWh) = average wall power (W) × operating hours per year / 1,000

If you think in completed jobs instead of hours:

Annual energy (kWh) = energy per task (kWh) × tasks per year

Do not use maximum rated power as if it were average consumption unless the workload actually keeps the GPU near that level. Measure representative runs when possible.

Also account for idle time. An always-on system can consume meaningful energy while waiting for requests:

Total daily energy = active energy + idle energy

Idle energy (kWh) = idle wall power (W) × idle hours / 1,000

3. Cooling and system overhead

Every watt consumed by the computer eventually becomes heat in the room. In a small office, that can increase air-conditioning use. In a server room, cooling may be a major operating cost.

A simple estimate is:

Cooling energy = computer energy × cooling overhead factor

The factor depends on the room, HVAC system, climate, ventilation, and whether the computer's waste heat would have been offset by heating demand. Treat it as an estimate, not a universal constant.

At minimum, compare GPUs using wall power rather than GPU-only power. If you need a detailed facility estimate, model the computer and cooling systems separately.

4. Maintenance and downtime

Maintenance costs may include:

  • Replacement fans
  • Thermal-interface service
  • Dust cleaning
  • Storage or power-supply replacement
  • Troubleshooting time
  • Monitoring and backup hardware
  • Lost productivity during failures

Downtime is difficult to price, but it can outweigh modest electricity differences for a business or production workload. A more efficient GPU is not necessarily the better choice if it has inadequate memory, poor software support, or cannot complete the job reliably.

5. Resale value

Resale value reduces the effective cost of ownership, but it is uncertain. It depends on:

  • Remaining warranty
  • Condition
  • Memory capacity
  • Market demand
  • New-generation performance
  • Power efficiency
  • Whether the card can still run current software

Use a conservative resale estimate. Do not treat resale proceeds as guaranteed.

Worked example: comparing two hypothetical GPUs

The following example uses invented values to demonstrate the method. They are not benchmark results or market recommendations.

Suppose two GPUs are measured on the same local AI task:

VariableGPU AGPU B
Purchase price$1,200$900
Expected resale value after three years$300$300
Average wall power during the task500 W330 W
Runtime per task6 seconds14 seconds
Annual maintenance allowance$60$40
Electricity price$0.30/kWh$0.30/kWh
Workload300,000 tasks/year300,000 tasks/year

The power and runtime values are hypothetical measurements for calculation purposes.

Step 1: Calculate energy per task

For GPU A:

Energy per task = 500 × 6 / 3,600,000

Energy per task = 0.000833 kWh

For GPU B:

Energy per task = 330 × 14 / 3,600,000

Energy per task = 0.001283 kWh

GPU A uses less energy per completed task despite drawing more power while active because it finishes substantially faster.

Step 2: Calculate performance per watt

GPU A completes:

1 / 6 = 0.1667 tasks per second

Its task throughput per watt is:

0.1667 / 500 = 0.000333 tasks per second per watt

GPU B completes:

1 / 14 = 0.0714 tasks per second

Its task throughput per watt is:

0.0714 / 330 = 0.000216 tasks per second per watt

GPU A is more efficient for this particular workload under these assumptions.

Step 3: Calculate annual electricity cost

GPU A:

0.000833 kWh × 300,000 tasks = 250 kWh/year

250 × $0.30 = $75/year

GPU B:

0.001283 kWh × 300,000 tasks = 385 kWh/year

385 × $0.30 = $115.50/year

Step 4: Calculate three-year TCO

GPU A's net hardware cost is:

$1,200 - $300 = $900

Three-year electricity cost is:

$75 × 3 = $225

Three-year maintenance allowance is:

$60 × 3 = $180

Therefore:

GPU A TCO = $900 + $225 + $180 = $1,305

GPU B's net hardware cost is:

$900 - $300 = $600

Three-year electricity cost is:

$115.50 × 3 = $346.50

Three-year maintenance allowance is:

$40 × 3 = $120

Therefore:

GPU B TCO = $600 + $346.50 + $120 = $1,066.50

Over three years and 900,000 tasks:

GPU A cost per task = $1,305 / 900,000 = $0.00145

GPU B cost per task = $1,066.50 / 900,000 = $0.00119

GPU A is faster and more energy-efficient in this example, but GPU B still has the lower total cost because its initial net hardware cost is $300 lower.

This is the central trade-off: better performance per watt does not automatically mean better financial value. The high-efficiency option must run often enough, save enough electricity, or create enough time savings to justify its additional purchase cost.

Finding the break-even point

To determine when a more expensive GPU becomes worthwhile, separate fixed and variable costs.

Let:

  • Fixed-cost difference = additional net hardware and other fixed costs
  • Variable-cost savings per task = operating cost of the less efficient option minus operating cost of the more efficient option

Then:

Break-even tasks = fixed-cost difference / variable-cost savings per task

In the example, GPU A costs $300 more in net hardware cost. Its electricity savings are:

$0.30 × (0.001283 - 0.000833) = $0.000135 per task

Ignoring maintenance differences:

Break-even tasks = $300 / $0.000135 = approximately 2,222,000 tasks

At 300,000 tasks per year, that is approximately 7.4 years. Including the maintenance difference changes the result, but the conclusion remains: electricity savings alone do not quickly repay the higher purchase price at this workload volume.

The calculation changes substantially if:

  • The electricity price is higher
  • The system runs more hours per day
  • The faster GPU enables more billable work
  • The runtime difference is larger
  • The price premium is smaller
  • Cooling costs are significant
  • The system remains in service for longer
  • The cheaper GPU cannot meet your latency or memory requirements

Break-even based on time value

Electricity is only part of the value of speed. If faster completion saves paid labor time or increases production capacity, assign a value to that time:

Time value per task = seconds saved per task / 3,600 × value of operator time per hour

Then:

Total savings per task = energy savings per task + time value per task

Use this only when the saved time is genuinely useful. If the GPU finishes early but the operator, pipeline, or downstream service remains idle, the theoretical speed advantage may not translate into economic value.

When performance per watt should dominate the decision

Efficiency deserves extra weight when:

  • The system runs continuously or nearly continuously
  • Electricity prices are high
  • The computer operates in a hot or poorly ventilated room
  • Cooling capacity is limited
  • You are deploying several GPUs
  • Noise is a constraint
  • The workload is repetitive and predictable
  • Power availability limits the number of systems you can operate
  • The GPU will be used for several years

For multi-GPU systems, power efficiency can also affect the power supply, electrical circuit, cooling design, and chassis choice. Those secondary costs may be more important than the electricity bill alone.

When purchase price or memory matters more

A lower-power GPU is not a good value if it cannot run the model you need. Before optimizing energy use, verify:

  • VRAM capacity
  • Required quantization and context length
  • Framework and driver support
  • Multi-GPU requirements
  • Target latency
  • Storage and system-memory requirements
  • Reliability under sustained load

A GPU that repeatedly runs out of memory, requires severe quantization, or forces an inefficient offload path may consume more energy per useful result despite having a lower nominal power draw.

How to measure your own system

For a practical comparison:

  1. Use the same model and software configuration on each GPU.
  2. Warm up the model before recording results.
  3. Measure several runs, not a single request.
  4. Record wall power with a plug-in meter when possible.
  5. Record runtime, throughput, and idle power separately.
  6. Test the workload at the batch size and context length you actually use.
  7. Calculate energy per completed task.
  8. Add hardware, maintenance, cooling, and resale assumptions.

A useful measurement table looks like this:

MeasurementGPU AGPU B
Average wall power while active
Idle wall power
Runtime per task
Throughput
Energy per task
Electricity cost per task
Purchase price
Expected resale value
Estimated maintenance

Record the assumptions with the results. A performance-per-watt number without the workload, power measurement point, and software configuration is difficult to reproduce or use for a buying decision.

Use total value, not a single efficiency number

Performance per watt answers:

How much useful work do I get for the electricity consumed?

Total cost of ownership answers:

What will this system cost to own and operate for my workload and time horizon?

Both questions matter. A GPU can be:

  • Fast but expensive to buy
  • Efficient but too slow for the required latency
  • Cheap but costly to run continuously
  • Efficient at one batch size and inefficient at another
  • Valuable because it avoids buying a second system
  • Economical only when used frequently

For a structured purchase comparison, use the GPU value calculator to combine hardware cost and expected usage rather than comparing specifications in isolation. If your usage is intermittent, also compare ownership with hosted capacity using the buy-versus-rent analysis.

The best choice is the GPU that meets your model and latency requirements at the lowest relevant total cost—not necessarily the one with the lowest wattage or the highest benchmark score.

Related guides