Wattproof

Open research project · Kubernetes · wall energy

Fewer watts at the wall. Proof anyone can re-run.

Abstract

Wattproof lowers the electricity a Kubernetes cluster draws at the wall, on hardware whose power bill you pay, without breaking the service levels of the workloads on it. It proves the saving with a randomised, pre-registered experiment that anyone can re-run from the published data.

The project is in its design phase. Nothing is ready to install. The design, the protocol, the code and every result, negative ones included, are public as they are made.

Status
Design and simulation
Licence
Apache-2.0
Registered results
None yet
Protocol
Public draft
Figure 1 one week, six servers sketch, not a measurement
Time (UTC)
Mon 00:00
Servers on
6 / 6
Wall power
— kW
Saved now, vs all on
—
Figure 1. The power-state lever, sketched. Servers powered to demand, all servers on, the difference. Over the whole week this rule draws 24% less than all servers on. Demand follows German Wikipedia's hourly pageviews for 1–7 September 2025 (CC0), peaking at 70% of capacity. Each server draws along the published SPECpower curve of a Dell PowerEdge R7725 (138 W idle, 861 W at full load), and 10 W when off. The rule here is deliberately simple: keep enough servers on to carry the next hour at 80% load. Worker power only. This shows the mechanism; it is neither Wattproof's planner nor a result. The simulated results are in §4. At rest it shows the whole week: how much of it each server was on, and the averages. Press Play to run the week, or drag across the chart to read any hour.
1

Three levers, and only stock Kubernetes

Most servers in a cluster are on all the time, and an idle server still draws a large share of its peak power. Kubernetes already knows how much capacity each workload reserves and uses. Wattproof uses that knowledge to switch off what is not needed and to bring it back before it is.

  1. 01

    Rightsize

    Lower CPU requests to what services use, hour by hour, through in-place pod resize. On hardware you own this alone saves nothing: the freed capacity stays powered.

  2. 02

    Power state

    Keep the most efficient servers on, drain and power off the rest, and power them on before demand returns. Taints, the eviction API and PodDisruptionBudgets only.

  3. 03

    Verify

    A controlled experiment keeps running, so the saving is measured against a baseline, not estimated from a model.

Models of each server's power draw, boot time and the cluster's demand are fitted continuously, and a new model is promoted only when it beats the one it replaces.

2

Measured at the wall, with the meter named

A saving is only as good as the meter behind it. Wattproof measures wall power per server, from metered PDU outlets calibrated against a reference power analyser. Every published number states the class of meter it came from, and headline claims use class A or B only.

ClassDevicesAccuracyRateUsed for
AReference power analyser on SPEC's PTDaemon list±0.1–0.2%10 HzCalibration, spot checks
BOutlet-metered PDU, billing grade±1%~1 HzContinuous per-server power
CBMC (Redfish, IPMI DCMI)coarse0.1–1 HzField installs; never headline claims
DComponent estimates (RAPL, Kepler, NVML)partialmsModel features; never totals

Table 1. Meter classes. Component estimates such as RAPL or NVML cover only part of a server, so they feed models and attribution, never totals. Full procedure in measurement.md.

3

A saving is a result, not an estimate

The claim is tested in a randomised experiment on a metered testbed, against two baselines. Beating stock Kubernetes is the headline. Beating a competent open-source setup is the claim an expert will check, so both are pre-registered.

B0
Stock Kubernetes. Default scheduler, every server on all the time.
B1
Tuned open source. Tight packing, the descheduler emptying lightly used servers, and Cluster Autoscaler powering empty servers off and on through the same power driver.
T
Wattproof, with its parameters frozen at registration.
  • The protocol is registered with a timestamp (OSF) before the first Wattproof block runs.
  • Arms run in randomised blocks, balanced for order, each starting at the daily peak.
  • The number of blocks and the analysis are fixed in advance. The experiment does not stop early.
  • The result is published whatever it shows, on the date the protocol states, with the raw data and the code behind every number.
  • A power-measurement scientist and a statistician review the design before it is frozen; every comment gets a written answer.
  • The testbed has two hardware generations, because identical servers would hide one of the few places a planner can beat B1.

Service level is measured next to energy: offered load, p99 latency and the minutes pods wait for capacity. A saving that breaks service is not counted as one.

4

What the simulations say, including what did not work

Before any hardware time is spent, a simulator estimates the effect on public traces: Google's 2019 cluster trace, Microsoft Azure's 2024 LLM inference trace, and published SPECpower curves. These are simulations, not measurements. They decide what the testbed has to show.

QuestionSimulated resultReads asEntry
Powering empty servers off, tuned open source against all servers on 13–27% less energythe lever Google trace
Wattproof against tuned open source, identical servers 5–11% more energyagainst Google trace
The same, two hardware generations 9% less to 9% moredepends Google trace
Minutes a week pods wait for a server, Wattproof against tuned open source 0 vs up to 220for First run
CPU-only rightsizing against VPA, once memory counts 4–58% more energyagainst Google trace
GPU inference servers powered to demand, against all on 8–52% less energyfor Azure LLM
The same, against a reactive autoscaler 2–12% more energytrade-off Azure LLM
Hours a week GPU replicas wait for a server, Wattproof against the reactive autoscaler 0–4.5 vs 11–29for Azure LLM

Table 2. Simulated results to date. “Tuned open source” is the B1 baseline of §3, approximated; “Wattproof” is its planner (T). Fluid models: the Google rows replay 13,469 services over four scored weeks with memory as a second limit, the Azure rows one week of inference traffic, the waiting row every CPU run at boot times of 2–10 minutes. Ranges span the idle powers and load times swept. Each row links to its notebook entry; every number reproduces with go run ./cmd/wattproof-sim.

The energy claim that survives so far is narrow: power state on mixed hardware. With the idle power the servers' SPECpower results report, Wattproof uses the same as tuned open source there, or up to 9% less. If idle power is 30–50% of peak, the services decide: from 9% less to 9% more. The power profiles measured on the testbed decide which case applies to its servers.

On identical servers, a planner that keeps headroom and powers on before demand returns uses 5–11% more energy than a reactive setup. What it buys is service: in the CPU runs no pod ever waits for a server. On the GPU week Wattproof's replicas wait at most 4.5 hours a week, against 11–29 hours for the reactive autoscaler, which uses 2–12% less energy.

Rightsizing CPU alone lost to the Vertical Pod Autoscaler once memory was modelled: memory then decides how many servers stay on. The design changed accordingly (ADR-0013). Earlier entries that ignored memory carry a correction at the top.

5

What is about to be tested

Each technique below is a hypothesis, not a feature. The simulations give a first estimate, and some of it points against Wattproof. Each question is tested on the metered testbed, with the test and what counts as failure fixed before it runs. Any of them can come out against Wattproof, and the result is published either way.

  1. Q1

    Does powering servers to demand use less wall energy than stock Kubernetes, without worse service?

    Test
    The registered experiment of §3: Wattproof against stock Kubernetes, in randomised blocks.
    Simulations expect
    A saving of 9–21% on the Google trace, larger where servers draw more when idle (§4).
    Counts against it
    Wall energy not lower, or p99 latency or minutes of SLO violation worse than a margin set in advance from the service level.
    When
    T0 + 6–9 weeks
  2. Q2

    Does it also beat a tuned open-source setup?

    Test
    The same experiment, against B1. Tested only if Q1 holds.
    Simulations expect
    Nothing on identical servers; a small difference either way on mixed hardware (§4).
    Counts against it
    The same tests as Q1, against B1.
    When
    T0 + 6–9 weeks
  3. Q3

    Does rightsizing that follows the hour of day add anything once power state runs?

    Test
    Experiment 2: Wattproof with resizing against Wattproof without it, and against the Vertical Pod Autoscaler. Memory is resized too if ADR-0013 is accepted.
    Simulations expect
    CPU-only rightsizing lost to VPA in every scenario once memory counted.
    Counts against it
    No less wall energy than VPA, or more CPU throttling, memory kills or evictions.
    When
    After Q1 and Q2, planned Feb 2027
  4. Q4

    Is switching GPU inference servers off worth it?

    Test
    One GPU server measured at the wall: off, idle and serving, boot and model load timed, 20 power cycles each followed by a health check. The measured values go back into the simulation.
    Counts against it
    Fixed in advance: if the simulated saving with these inputs is under 5% of all servers on, GPU power state is dropped.
    When
    With the registered experiment, where an operator gives access
  5. Q5

    Do GPU energy settings save as much at the wall as on the GPU board?

    Test
    On the same server: default settings, a vendor's published energy profile, and locked clocks per serving phase, compared by wall energy per generated token.
    Counts against it
    No lower wall energy per token, or token latency worse than a margin set in advance. The wall saving is reported next to the board saving.
    When
    With Q4
6

Roadmap

Dates after the testbed counts from T0, the day a metered testbed is handed over. That date is the critical path, and it is not the project's to set.

  1. Sep–Oct 2026

    Design in public

    Architecture, decision records, related work, the measurement method and the protocol draft.

  2. Oct 2026

    Simulator on public traces

    Google 2019 and Azure 2024 traces as inputs; the results in §4.

  3. Now

    Report and measurement software

    A rightsizing report that works on any cluster, the Redfish meter driver, and the power-profiling tools.

  4. T0 · target Nov 2026

    Metered testbed

    Two hardware generations, every server metered at class B or better. Calibration, profiles, an A/A pilot, outside review.

  5. T0 + 6–9 weeks

    Registered experiment

    Wattproof against B0 and B1, as registered. One GPU server is also measured at the wall.

  6. planned Jan 2027

    v0.1 and the report

    Signed release, the full report with confidence intervals, and the data with a DOI.

  7. planned Feb 2027

    Experiment 2

    Rightsizing that follows the day, against the Vertical Pod Autoscaler.

8

Take part

Operators running Kubernetes on their own hardware: the rightsizing report will run in observe mode, change nothing, and tell you in servers, kWh and peak power what could be saved. Write if you would like to try it first.

Researchers in power measurement, statistics or scheduling: the protocol is open for comments until it is registered.

Contributors: the code is Apache-2.0, contributions are signed off under the Developer Certificate of Origin (CONTRIBUTING.md).