DISTRIBUTED SYSTEMS · IEEE 2023

One workload.
Five places.
A smarter choice.

Electricity prices change. Cooling costs change. GeoSched sends flexible jobs to the data center where the two add up to less.

5 sec simulation time step5 geographically distributed sites5 min resource sharing intervalInspect paper & source ↗
The problem

Data centers pay for power twice: to compute, then to cool.

Cloud providers run data centers around the world. Each one pays its local electricity price, and each one has to remove the heat its servers make. Cooling can take up to half of a data center's energy.

How much cooling costs depends on the weather. When it is cool outside, a site can blow outside air through the servers, which adds only 5–17% to the power bill. When it is hot, it has to run chillers, which add about a third.

So the same job costs a different amount in Singapore than in Oregon, and a different amount in January than in July. GeoSched is a scheduler that uses this: it sends each job that can wait a little to the site where running it costs least.

Playground

Pick a week and a job. See where it goes.

This one-job illustration uses the paper’s incremental-cost equations and weekly figure-derived profiles. Availability is a simplified toggle; the Python simulator checks individual tasks against machine CPU and memory. The cost includes cooling: outside air below 65°F (Table V), chillers above it (Eq. 7).

Shared availability · uncheck a site to exclude it from placement

The year

The cheapest site changes with the seasons.

Weekly figures for 2013. "All-in cost" is the power price times one plus the cooling overhead: what one MWh of computing really costs at that site.

How it works

Three steps for every job.

01

Profile every site

At five-second ticks, jobs arrive and complete. Every five minutes, sites publish snapshots of the CPU and memory available on each machine.

02

Price the job everywhere

Use predicted duration to price the extra compute power and cooling at each feasible site (Eqs. 9–10). Whole-site energy also counts idle power once.

03

Send it to the cheapest

Classes 0–1 choose the cheapest feasible snapshot. Classes 2–3 stay local. Reserve every task atomically; stale availability can cause queuing. The model does not guarantee an SLA.

Results

GeoSched against running every job locally.

The paper compares GeoSched with DumbSched, which runs every job where it arrives. Six 5-day simulations, starting on the 15th of January, March, May, July, September and November.

Paper · reported operating cost
−11.7%

The published study compares GeoSched with running jobs locally.

Paper · reported energy
−8.2%

These are reported paper values, not a new run of the source below.

NEW IMPLEMENTATION · SYNTHETIC DEMONSTRATION

Same workload. Different placement.

Compare local placement with GeoSched using the new task-level simulator. This small example uses 500 synthetic jobs, eight machines per site and one-hour windows. It is not a Google-trace reproduction.

Loading the saved demonstration…
SiteLocal costGeoSched costLocal energyGeoSched energyGeoSched completed
What this example can establish

It demonstrates scheduling and accounting under documented assumptions. It cannot validate the paper’s reported savings. Costs include idle power; all queues, running jobs and rejections are reported. No constants were tuned to match 11.7%.

Follow the experiment from input to result.

1. Prepare a trace

Provide successful jobs with arrival time, actual duration, an independently estimated duration, scheduling class and each task’s CPU/memory requests. Keep Google’s failed and killed jobs out. The bundled upstream one-day aggregate traces cannot recover individual task sizes without an extra approximation.

2. Match the paper’s experiment

Five sites use Table IV’s heterogeneous machine configurations. Use five-day windows starting at midnight on January 15, March 15, May 15, July 15, September 15 and November 15. Resource state is shared every 300 seconds, with five-second simulation steps.

3. Inspect assumptions

The paper omits some constants and policies. Default power constants, 20°C supply air, FIFO queues, first-fit packing and weekly profiles are declared assumptions. The runtime prediction is supplied as an input; the demo uses a fixed estimate. Data-transfer cost, node shutdown and enforceable SLA deadlines are excluded.

4. Check comparable work

Compare identical input jobs. Check completed, queued, running and rejected counts before interpreting savings. Integrate static and dynamic IT power plus cooling over the same observation window. Jobs continuing beyond that window remain marked running.

Why it matters

Location can reduce operating cost. Where it runs decides what it costs, and the answer changes every season.

Limits: the model covers server and cooling power, not the cost of moving data between sites, and it uses weekly averages for price and temperature. The paper names renewable energy, data transfer and richer service-level agreements as next steps. The new simulator makes its own assumptions visible; numerical reproduction remains unverified.