This research demonstrates that applying custom, task-specific power limits to modern superchips is a highly effective strategy for saving significant amounts of GPU energy in high-performance computing. By meticulously analyzing the relationship between power, performance, and GPU energy consumption, the study proves that a one-size-fits-all approach is inefficient.
Achievement As supercomputers become exponentially more powerful, their energy consumption has emerged as a major operational and environmental concern. This research contributes to the understanding of measuring and analyzing the energy used by powerful computing chips GPUs, when running two different large-scale scientific simulations. By providing a detailed, application-level view of power traces and energy
High-performance computing systems consume vast amounts of energy, particularly when moving data between different parts of the machine. To address this challenge, a research team investigated a novel strategy for optimizing data transfers.
The Simplified Interface to Complex Memories (SICM) project delivers a powerful software solution that abstracts away the complexity of modern, multi-tiered memory systems. By providing a unified interface and automated data placement strategies, SICM allows scientific applications to achieve optimal performance without requiring developers to write complex, non-portable code. This work is critical for the future of high-performance computing, as it enhances developer productivity, boosts application efficiency, and provides a durable framework for harnessing the power of next-generation computer architectures. The result is a practical and effective tool that makes exascale systems more accessible and powerful for the entire scientific community.
The research team developed an intelligent, automated software solution that elegantly solves the complex problem of managing data in modern computers with multiple memory types. This framework helps applications run more efficiently on today's advanced hardware, improves how resources are utilized in multi-tasking environments, and simplifies data management without requiring any manual intervention from programmers. By making heterogeneous memory systems both powerful and easy to use, this work provides an essential enabling technology for next-generation architectures, including systems with high-bandwidth and disaggregated memories.
Summary: Automation and autonomy can enable revolutionary scientific advances by coordinating a diverse array of experimental and computational capabilities more efficiently and more effectively than current hands-on approaches. This experiment creates an autonomous system to plan and adaptively control additive manufacturing build processes. It involves multiple characterization modes, computation across the edge-to-center computing continuum, and multiple scientific user facilities. The objective of the autonomous additive manufacturing (AAM) system is to control the residual stress in a part to address a grand challenge – building parts that are ready and safe to use immediately (i.e., “born qualified”). The AAM system is deployed at ORNL’s Manufacturing Demonstration Facility (MDF), Spallation Neutron Source (SNS), and Oak Ridge Leadership Computing Facility (OLCF) as a cross-facility instrument-science workflow. Its INTERSECT architecture consists of science use case design patterns, a system of systems architecture, and a microservices architecture. For more details see: https://intersect-architecture.readthedocs.io/en/latest/examples/aam/.
This research introduces a data-efficient, AI-driven framework for making smarter scheduling decisions in High-Performance Computing. By using attention-based techniques and intelligent data sampling, the method effectively models the complex trade-off between performance and power. This work paves the way for more sustainable next-generation supercomputing systems that accelerate scientific discovery while minimizing operational costs.
By combining AI with molecular dynamics simulations, researchers at ORNL have developed a new tool to more accurately predict how plants and helpful microbes communicate and form partnerships at the most fundamental level. The new AI-powered workflow helps scientists identify which plant genes control the best microbial partnerships.
Two-and-a-half years after breaking the exascale barrier, the Frontier supercomputer at the Department of Energy’s Oak Ridge National Laboratory continues to set new standards for its computing speed and performance.
Click here for static version. Just before dawn, Scott Atchley woke up for the third time, took another sip of coffee and sat down at his computer to watch the next failure. It was the morning of May 27, 2022. Atchley and fellow scientists had spent months tuning and tweaking Frontier, the $600 million supercomputer installed