BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//cfp.pydata.org//pydataglobal2025//talk//VHX7E7
BEGIN:VEVENT
SUMMARY:Lessons learnt in optimizing a large-scale pandas application usin
 g Polars\, FireDucks and cuDF: Go Smart and Save More! - Sourav Saha
DTSTART:20251209T133000Z
DTEND:20251209T140000Z
DTSTAMP:20260914T183025Z
UID:pretalx-pydataglobal2025-VHX7E7@cfp.pydata.org
DESCRIPTION:In general\, a Data Scientist spends significant efforts in tr
 ansforming the raw data into a more digestible format before training an A
 I model or creating visualisations. Traditional tools such as pandas have 
 long been the linchpin in this process\, offering powerful capabilities bu
 t not without limitations. With numerous possible ways to write the same t
 hing in pandas\, often a user ends up selecting the uneconomical\, ineffic
 ient ones\, leading to large computational　costs　with the growth in da
 ta size. We introduce a couple of frequently occurring　intricate perform
 ance issues in pandas\, and what we have learnt in solving the same using 
 popular high-performance pandas alternatives: Polars\, FireDucks and cuDF.
  The talk intends to highlight one of the best practices (breaking out of 
 the loops) that one should follow while dealing with large-scale data anal
 ysis\, while demonstrating the key advantages of the high-performance pand
 as alternatives based on different scenarios.
LOCATION:Analytics\, Visualization & Decision Science
URL:https://cfp.pydata.org/pydataglobal2025/talk/VHX7E7/
END:VEVENT
END:VCALENDAR
