论文信息 - Towards a NoOps Model for WLCG

Towards a NoOps Model for WLCG

One of the most costly factors in providing a global computing infrastructure such as the WLCG is the human effort in deployment, integration, and operation of the distributed services supporting collaborative computing, data sharing and delivery, and analysis of extreme scale datasets. Furthermore, the time required to roll out global software updates, introduce new service components, or prototype novel systems requiring coordinated deployments across multiple facilities is often increased by communication latencies, staff availability, and in many cases expertise required for operations of bespoke services. While the WLCG (and distributed systems implemented throughout HEP) is a global service platform, it lacks the capability and flexibility of a modern platform-as-a-service including continuous integration/continuous delivery (CI/CD) methods, development-operations capabilities (DevOps, where developers assume a more direct role in the actual production infrastructure), and automation. Most importantly, tooling which reduces required training, bespoke service expertise, and the operational effort throughout the infrastructure, most notably at the resource endpoints (sites), is entirely absent in the current model. In this paper, we explore ideas and questions around potential NoOps models in this context: what is realistic given organizational policies and constraints? How should operational responsibility be organized across teams and facilities? What are the technical gaps? What are the social and cybersecurity challenges? Conversely what advantages does a NoOps model deliver for innovation and for accelerating the pace of delivery of new services needed for the HL-LHC era? We will describe initial work along these lines in the context of providing a data delivery network supporting IRIS-HEP DOMA R&D.

[1] Jiahui Chen,et al. Building the SLATE Platform , 2018, PEARC.

[2] Weisong Shi,et al. Edge Computing: Vision and Challenges , 2016, IEEE Internet of Things Journal.

[3] Eli Dart,et al. The Modern Research Data Portal: a design pattern for networked, data-intensive science , 2018, PeerJ Comput. Sci..

[4] Tadashi Maeno,et al. Operation of the ATLAS Distributed Computing , 2019 .

[5] Johannes Elmsheuser,et al. Overview of the ATLAS distributed computing system , 2019 .

[6] Eric A. Brewer,et al. Borg, Omega, and Kubernetes , 2016, ACM Queue.