Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

🎬 Netflix Data Analysis

A comprehensive Exploratory Data Analysis (EDA) project on the Netflix Movies and TV Shows dataset using Python. This project uncovers trends in Netflix's content library through data cleaning, preprocessing, and insightful visualizations.

🏅 This notebook received a Kaggle Bronze Medal for its quality of analysis and presentation.


📌 Project Overview

This project analyzes Netflix's catalog to answer questions such as:

  • Which countries produce the most Netflix content?
  • How has Netflix's content library grown over time?
  • What is the ratio of Movies to TV Shows?
  • What are the most common ratings?
  • Which genres dominate the platform?
  • How has Netflix expanded globally?

The project follows a complete EDA workflow including data cleaning, feature engineering, visualization, and insight generation.


📂 Dataset

  • Dataset: Netflix Movies and TV Shows
  • Source: Kaggle
  • Records: ~8,800 titles
  • Features: 12

Some important columns include:

  • Title
  • Type
  • Director
  • Cast
  • Country
  • Date Added
  • Release Year
  • Rating
  • Duration
  • Listed In (Genre)
  • Description

🛠️ Technologies Used

  • Python
  • Pandas
  • NumPy
  • Matplotlib
  • Seaborn
  • Jupyter Notebook

📊 Analysis Performed

The notebook includes:

  • Data Cleaning
  • Missing Value Analysis
  • Duplicate Detection
  • Data Type Conversion
  • Feature Engineering
  • Country-wise Analysis
  • Genre Analysis
  • Rating Distribution
  • Release Year Trends
  • Movies vs TV Shows Comparison
  • Data Visualization

📈 Key Insights

🎥 Movies dominate Netflix's catalog

Approximately 70% of Netflix's content consists of Movies, while 30% are TV Shows.

Movies vs TV Shows


🌍 United States produces the highest amount of Netflix content

The United States contributes significantly more titles than any other country, followed by India and the United Kingdom.

Top Countries


📅 Massive growth after 2015

Netflix experienced rapid content expansion between 2015–2019, reaching its highest number of releases before declining slightly in later years.

Netflix Content Over Years


📁 Project Structure

Netflix-data-Analysis/
│
├── images/
│   ├── movies_vs_tv.png
│   ├── top_countries.png
│   └── content_over_years.png
│
├── netflix-data-analysis.ipynb
├── netflix_titles.csv
└── README.md

🚀 How to Run

  1. Clone the repository
git clone https://github.com/RugvedBane/Netflix-data-Analysis.git
  1. Navigate into the project
cd Netflix-data-Analysis
  1. Install the required libraries
pip install pandas numpy matplotlib seaborn notebook
  1. Launch Jupyter Notebook
jupyter notebook
  1. Open
netflix-data-analysis.ipynb

📚 Skills Demonstrated

  • Exploratory Data Analysis (EDA)
  • Data Cleaning
  • Data Visualization
  • Feature Engineering
  • Statistical Analysis
  • Business Insight Generation
  • Python for Data Analysis

🏆 Achievement

🥉 Awarded a Kaggle Bronze Medal for this notebook.

The project was recognized for its analytical approach, visualizations, and presentation on Kaggle.


👨‍💻 Author

Rugved Bane


⭐ If you found this project useful, consider giving it a Star.

About

Exploratory Data Analysis on 8,800+ Netflix titles — content trends, genre patterns & release insights | Kaggle Bronze Medal

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages