Open research project · Kubernetes · wall energy
Fewer watts at the wall. Proof anyone can re-run.
Abstract
Wattproof lowers the electricity a Kubernetes cluster draws at the wall, on hardware whose power bill you pay, without breaking the service levels of the workloads on it. It proves the saving with a randomised, pre-registered experiment that anyone can re-run from the published data.
The project is in its design phase. Nothing is ready to install. The design, the protocol, the code and every result, negative ones included, are public as they are made.
- Status
- Design and simulation
- Licence
- Apache-2.0
- Registered results
- None yet
- Protocol
- Public draft
- Time (UTC)
- Mon 00:00
- Servers on
- 6 / 6
- Wall power
- — kW
- Saved now, vs all on
- —
Three levers, and only stock Kubernetes
Most servers in a cluster are on all the time, and an idle server still draws a large share of its peak power. Kubernetes already knows how much capacity each workload reserves and uses. Wattproof uses that knowledge to switch off what is not needed and to bring it back before it is.
- 01
Rightsize
Lower CPU requests to what services use, hour by hour, through in-place pod resize. On hardware you own this alone saves nothing: the freed capacity stays powered.
- 02
Power state
Keep the most efficient servers on, drain and power off the rest, and power them on before demand returns. Taints, the eviction API and PodDisruptionBudgets only.
- 03
Verify
A controlled experiment keeps running, so the saving is measured against a baseline, not estimated from a model.
Models of each server's power draw, boot time and the cluster's demand are fitted continuously, and a new model is promoted only when it beats the one it replaces.
Measured at the wall, with the meter named
A saving is only as good as the meter behind it. Wattproof measures wall power per server, from metered PDU outlets calibrated against a reference power analyser. Every published number states the class of meter it came from, and headline claims use class A or B only.
| Class | Devices | Accuracy | Rate | Used for |
|---|---|---|---|---|
| A | Reference power analyser on SPEC's PTDaemon list | ±0.1–0.2% | 10 Hz | Calibration, spot checks |
| B | Outlet-metered PDU, billing grade | ±1% | ~1 Hz | Continuous per-server power |
| C | BMC (Redfish, IPMI DCMI) | coarse | 0.1–1 Hz | Field installs; never headline claims |
| D | Component estimates (RAPL, Kepler, NVML) | partial | ms | Model features; never totals |
Table 1. Meter classes. Component estimates such as RAPL or NVML cover only part of a server, so they feed models and attribution, never totals. Full procedure in measurement.md.
A saving is a result, not an estimate
The claim is tested in a randomised experiment on a metered testbed, against two baselines. Beating stock Kubernetes is the headline. Beating a competent open-source setup is the claim an expert will check, so both are pre-registered.
- B0
- Stock Kubernetes. Default scheduler, every server on all the time.
- B1
- Tuned open source. Tight packing, the descheduler emptying lightly used servers, and Cluster Autoscaler powering empty servers off and on through the same power driver.
- T
- Wattproof, with its parameters frozen at registration.
- The protocol is registered with a timestamp (OSF) before the first Wattproof block runs.
- Arms run in randomised blocks, balanced for order, each starting at the daily peak.
- The number of blocks and the analysis are fixed in advance. The experiment does not stop early.
- The result is published whatever it shows, on the date the protocol states, with the raw data and the code behind every number.
- A power-measurement scientist and a statistician review the design before it is frozen; every comment gets a written answer.
- The testbed has two hardware generations, because identical servers would hide one of the few places a planner can beat B1.
Service level is measured next to energy: offered load, p99 latency and the minutes pods wait for capacity. A saving that breaks service is not counted as one.
What the simulations say, including what did not work
Before any hardware time is spent, a simulator estimates the effect on public traces: Google's 2019 cluster trace, Microsoft Azure's 2024 LLM inference trace, and published SPECpower curves. These are simulations, not measurements. They decide what the testbed has to show.
| Question | Simulated result | Reads as | Entry |
|---|---|---|---|
| Powering empty servers off, tuned open source against all servers on | 13–27% less energy | the lever | Google trace |
| Wattproof against tuned open source, identical servers | 5–11% more energy | against | Google trace |
| The same, two hardware generations | 9% less to 9% more | depends | Google trace |
| Minutes a week pods wait for a server, Wattproof against tuned open source | 0 vs up to 220 | for | First run |
| CPU-only rightsizing against VPA, once memory counts | 4–58% more energy | against | Google trace |
| GPU inference servers powered to demand, against all on | 8–52% less energy | for | Azure LLM |
| The same, against a reactive autoscaler | 2–12% more energy | trade-off | Azure LLM |
| Hours a week GPU replicas wait for a server, Wattproof against the reactive autoscaler | 0–4.5 vs 11–29 | for | Azure LLM |
Table 2. Simulated results to date. “Tuned open source” is the B1 baseline of §3, approximated; “Wattproof” is its planner (T). Fluid models: the Google rows replay 13,469 services over four scored weeks with memory as a second limit, the Azure rows one week of inference traffic, the waiting row every CPU run at boot times of 2–10 minutes. Ranges span the idle powers and load times swept. Each row links to its notebook entry; every number reproduces with go run ./cmd/wattproof-sim.
The energy claim that survives so far is narrow: power state on mixed hardware. With the idle power the servers' SPECpower results report, Wattproof uses the same as tuned open source there, or up to 9% less. If idle power is 30–50% of peak, the services decide: from 9% less to 9% more. The power profiles measured on the testbed decide which case applies to its servers.
On identical servers, a planner that keeps headroom and powers on before demand returns uses 5–11% more energy than a reactive setup. What it buys is service: in the CPU runs no pod ever waits for a server. On the GPU week Wattproof's replicas wait at most 4.5 hours a week, against 11–29 hours for the reactive autoscaler, which uses 2–12% less energy.
Rightsizing CPU alone lost to the Vertical Pod Autoscaler once memory was modelled: memory then decides how many servers stay on. The design changed accordingly (ADR-0013). Earlier entries that ignored memory carry a correction at the top.
What is about to be tested
Each technique below is a hypothesis, not a feature. The simulations give a first estimate, and some of it points against Wattproof. Each question is tested on the metered testbed, with the test and what counts as failure fixed before it runs. Any of them can come out against Wattproof, and the result is published either way.
- Q1
Does powering servers to demand use less wall energy than stock Kubernetes, without worse service?
- Test
- The registered experiment of §3: Wattproof against stock Kubernetes, in randomised blocks.
- Simulations expect
- A saving of 9–21% on the Google trace, larger where servers draw more when idle (§4).
- Counts against it
- Wall energy not lower, or p99 latency or minutes of SLO violation worse than a margin set in advance from the service level.
- When
- T0 + 6–9 weeks
- Q2
Does it also beat a tuned open-source setup?
- Test
- The same experiment, against B1. Tested only if Q1 holds.
- Simulations expect
- Nothing on identical servers; a small difference either way on mixed hardware (§4).
- Counts against it
- The same tests as Q1, against B1.
- When
- T0 + 6–9 weeks
- Q3
Does rightsizing that follows the hour of day add anything once power state runs?
- Test
- Experiment 2: Wattproof with resizing against Wattproof without it, and against the Vertical Pod Autoscaler. Memory is resized too if ADR-0013 is accepted.
- Simulations expect
- CPU-only rightsizing lost to VPA in every scenario once memory counted.
- Counts against it
- No less wall energy than VPA, or more CPU throttling, memory kills or evictions.
- When
- After Q1 and Q2, planned Feb 2027
- Q4
Is switching GPU inference servers off worth it?
- Test
- One GPU server measured at the wall: off, idle and serving, boot and model load timed, 20 power cycles each followed by a health check. The measured values go back into the simulation.
- Counts against it
- Fixed in advance: if the simulated saving with these inputs is under 5% of all servers on, GPU power state is dropped.
- When
- With the registered experiment, where an operator gives access
- Q5
Do GPU energy settings save as much at the wall as on the GPU board?
- Test
- On the same server: default settings, a vendor's published energy profile, and locked clocks per serving phase, compared by wall energy per generated token.
- Counts against it
- No lower wall energy per token, or token latency worse than a margin set in advance. The wall saving is reported next to the board saving.
- When
- With Q4
Roadmap
Dates after the testbed counts from T0, the day a metered testbed is handed over. That date is the critical path, and it is not the project's to set.
- Sep–Oct 2026
Design in public
Architecture, decision records, related work, the measurement method and the protocol draft.
- Oct 2026
Simulator on public traces
Google 2019 and Azure 2024 traces as inputs; the results in §4.
- Now
Report and measurement software
A rightsizing report that works on any cluster, the Redfish meter driver, and the power-profiling tools.
- T0 · target Nov 2026
Metered testbed
Two hardware generations, every server metered at class B or better. Calibration, profiles, an A/A pilot, outside review.
- T0 + 6–9 weeks
Registered experiment
Wattproof against B0 and B1, as registered. One GPU server is also measured at the wall.
- planned Jan 2027
v0.1 and the report
Signed release, the full report with confidence intervals, and the data with a DOI.
- planned Feb 2027
Experiment 2
Rightsizing that follows the day, against the Vertical Pod Autoscaler.
Notebook and results, kept apart
The lab notebook holds everything since day zero, exploratory runs and corrections included. Registered results hold only confirmatory experiments: the numbers the project claims. The two are never mixed.
Lab notebook
6 entries- GPU server power state on the Azure inference week Saves 8–52% against all on; the planner must smooth its input to avoid cycling servers too often.
- The simulator on Google's services, with memory Once memory counts, CPU-only rightsizing loses to VPA in every scenario.
- What Google's services reserve and use 13,469 services from the 2019 trace; peak usage is a median 0.68 of the CPU requested.
- How much the conclusions depend on the usage assumptioncorrected What holds whatever the ratio, and what does not.
- Simulating request shapingcorrected VPA captures most of the rightsizing gain in this model.
- First simulationcorrected No energy margin over tuned open source on identical servers: two hardware generations become a requirement.
Registered results
0None yet.
The first registered experiment runs on the metered testbed. Its result appears here on the date its protocol states, whatever it shows, with the data and its DOI.
Take part
Operators running Kubernetes on their own hardware: the rightsizing report will run in observe mode, change nothing, and tell you in servers, kWh and peak power what could be saved. Write if you would like to try it first.
Researchers in power measurement, statistics or scheduling: the protocol is open for comments until it is registered.
Contributors: the code is Apache-2.0, contributions are signed off under the Developer Certificate of Origin (CONTRIBUTING.md).