September 2026

Conference Paper

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

By:
Mcdaniel, Adam R; Jantz, Michael; Sharma, Ashesh; Martin, Steven; Abbott, Steve; Khandekar, Shreyas; Neth, Brandon; Alvarez, Bruno; Kashi, Aditya ; Elwasif, Wael R; Hernandez Mendoza, Oscar R
Page Number:
1-15
Book Title:
ISC High Performance 2026 Research Paper Proceedings (41st International Conference)
Publication Date:
September 11, 2026
Publisher Location:
IEEE, New Jersey, United States of America
Conference Name:
ISC High Performance 2026
Conference Location:
Hamburg, Germany
Conference Sponsor:
IEEE
View DOI Listing:
https://doi.org/10.23919/ISC.2026.11520492

Abstract

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.