We make high-performance systems more efficient without giving up their latency and service-level goals. Our work combines fine-grained core management, power-aware scheduling, processor idle states, and hardware/software co-design for latency-critical cloud workloads. Recent results show how energy savings can be exposed as a systems problem rather than a hardware-only setting.