This presentation will introduce Arcadia, a unified system designed to simulate compute, memory, and network performance of AI training clusters. By providing a multi-disciplinary performance analysis framework, Arcadia aims to facilitate the design and optimization of various system levels, including application, network, and hardware. This comprehensive system enables researchers and practitioners to gain valuable insights into the performance of future AI models and workloads on specific infrastructures, fostering data-driven decision-making processes and promoting the future evolution of models and hardware. Arcadia provides ability to simulate performance impact of scheduled operational tasks on AI-models that are running in production; helps an engineer to make job-aware decisions during day-to-day operational activity. Attendees will learn about the capabilities and potential impact of Arcadia in advancing the field of AI systems and infrastructure.
- WATCH NOW
- 2024 EVENTS
- PAST EVENTS
- 2023
- 2022
- February
- RTC @Scale 2022
- March
- Systems @Scale Spring 2022
- April
- Product @Scale Spring 2022
- May
- Data @Scale Spring 2022
- June
- Systems @Scale Summer 2022
- Networking @Scale Summer 2022
- August
- Reliability @Scale Summer 2022
- September
- AI @Scale 2022
- November
- Networking @Scale Fall 2022
- Video @Scale Fall 2022
- December
- Systems @Scale Winter 2022
- 2021
- 2020
- 2019
- 2018
- 2017
- 2016
- 2015
- Blog & Video Archive
- Speaker Submissions