BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//cfp.pydata.org//pydataglobal2025//speaker//XAYYZD
BEGIN:VEVENT
UID:pretalx-pydataglobal2025-TXYJHL@cfp.pydata.org
DTSTART:20251210T163000Z
DTEND:20251210T170000Z
DESCRIPTION:The proliferation of AI/ML workloads across commercial enterpri
 ses\, necessitates robust mechanisms to track\, inspect and analyze their 
 use of on-prem/cloud infrastructure. To that end\, effective insights are 
 crucial for optimizing cloud resource allocation with increasing workload 
 demand\, while mitigating cloud infrastructure costs and promoting operati
 onal stability.\n\nThis talk will outline an approach to systematically mo
 nitor\, inspect and analyze AI/ML workloads’ properties like runtime\, r
 esource demand/utilization and cost attribution tags . By implementing gra
 nular inspection across multi-player teams and projects\, organizations ca
 n gain actionable insights into resource bottlenecks\, identify opportunit
 ies for cost savings\, and enable AI/ML platform engineers to directly att
 ribute infrastructure costs to specific workloads. \n\nCost attribution of
  infrastructure usage by AI/ML workloads focuses on key metrics such as co
 mpute node group information\,  cpu usage seconds\, data transfer\, gpu al
 location \, memory and ephemeral storage utilization. It enables platform 
 administrators to identify competing workloads which lead to diminishing R
 OI. Answering questions from data scientists like "Why did my workload run
  for 6 hours today\, when it took only 2 hours yesterday" or "Why did my w
 orkload start 3 hours behind schedule?" also becomes easier.\n\nThrough ou
 r work on Metaflow\, we will showcase how we built a comprehensive framewo
 rk for transparent usage reporting\, cost attribution\, performance optimi
 zation\, and strategic planning for future AI/ML initiatives. Metaflow is 
 a human centric python library that enables seamless scaling and managemen
 t of AI/ML projects.\n\nUltimately\, a well-defined usage tracking system 
 empowers organizations to maximize the return on investment from their AI/
 ML endeavors while maintaining budgetary control and operational efficienc
 y. Platform engineers and administrators will be able to gain insights int
 o the following operational aspects of supporting a battle hardened ML Pla
 tform:\n\n1.Optimize resource allocation: Understand consumption patterns 
 to right-size clusters and allocate resources more efficiently\, reducing 
 idle time and preventing bottlenecks.\n\n2. Proactively manage capacity: F
 orecast future resource needs based on historical usage trends\, ensuring 
 the infrastructure can scale effectively with increasing workload demand.\
 n\n3. Facilitate strategic planning: Make informed decisions regarding fut
 ure infrastructure investments and scaling strategies.\n\n4.Diagnose workl
 oad execution delays: Identify resource contention\, queuing issues\, or i
 nsufficient capacity leading to delayed workload starts.\n\nData Scientist
 s on the other hand will gain clarity on factors that influence workload p
 erformance. Tuning them can lead to efficiencies in runtime and associated
  cost profiles.
DTSTAMP:20260719T005733Z
LOCATION:Machine Learning & AI
SUMMARY:Optimizing AI/ML Workloads: Resource Management and Cost Attributio
 n - Saurabh Garg
URL:https://cfp.pydata.org/pydataglobal2025/talk/TXYJHL/
END:VEVENT
END:VCALENDAR
