Project Description: Ingested and pre-processed over 5,000+ unformatted user streaming logs to extract patterns relating to user retention and viewer preferences. Re-enforced data reliability through Pandas by purging duplicated server logs and resolving missing feedback structures (NaN ratings) utilizing internal median calculations. Seamlessly orchestrated automated pipeline queries through DuckDB to synthesize rapid Dataframe outcomes without memory overheads. Queried complex metrics grouping massive segments by specific Account Subscription Tiers to discover maximum Watch-Time correlations. Measured multi-parameter user interactions, successfully ranking global streaming Content Genres based on raw view volume mapped against peak User Ratings.