Data Science
Finding answers inside data
What is it?
Every day, the world creates a huge amount of data: shopping records, phone locations, hospital records, bank payments, weather measurements, and social media posts. By itself, data is just a big pile of numbers and words. Data Science is the work of turning that pile into useful answers.
An everyday example
Imagine you own a small shop and you write down everything you sell for a whole year. A data scientist looks at your notebook and tells you: “You sell the most ice cream on hot Fridays, your Tuesday sales are weak, and customers who buy bread usually also buy milk. So put milk next to the bread!”
How does it work?
A data scientist usually follows these steps:
- 1
Ask a question.
For example: “Why are we losing customers?”
- 2
Collect data.
Gather records from databases, websites, sensors, or surveys.
- 3
Clean the data.Often the biggest part of the job
Real data is messy. Some values are missing, some are typed wrongly, and some are repeated. Cleaning is often the biggest part of the job.
- 4
Explore the data.
Look at averages, patterns, and charts to understand what is going on.
- 5
Analyze or build a model.
Use statistics or Machine Learning to find deeper patterns or make predictions.
- 6
Share the story.
Explain the results with simple charts and clear words so that decision-makers can act.
Before cleaning
| date | product | qty | problem |
|---|---|---|---|
| 4 Jul | Ice cream | 12 | |
| 4 Jul | Ice cream | 12 | Repeated row |
| 5 Jul | Bread | Missing value | |
| 5 Jul | Mlik | 4 | Typed wrongly |
After cleaning
| date | product | qty | what we did |
|---|---|---|---|
| 4 Jul | Ice cream | 12 | Copy removed |
| 5 Jul | Bread | 3 | Filled from the receipt |
| 5 Jul | Milk | 4 | Spelling fixed |
What skills and tools are used?
- Statistics (understanding numbers and chance)
- Programming (mostly Python)
- Databases (SQL)
- Charts and dashboards
What it looks like in Python
Here is how a data scientist could ask the shop's notebook two of those questions with Python and the Pandas library. You will write code like this in Module 1.
import pandas as pd
# One row per item sold: receipt_id, date, weekday, product, quantity
sales = pd.read_csv("shop_sales.csv")
# Which weekday sells the most ice cream?
ice_cream = sales[sales["product"] == "ice cream"]
print(ice_cream.groupby("weekday")["quantity"].sum().sort_values(ascending=False))
# What do customers buy together with bread?
baskets = sales.groupby("receipt_id")["product"].apply(list)
with_bread = baskets[baskets.apply(lambda items: "bread" in items)]
print(with_bread.explode().value_counts().drop("bread").head(3))Example code. The file shop_sales.csv stands in for the shop's year of sales.
Where is it used?
Business and Retail
Finding best-selling products, deciding prices, understanding what customers want
Banking and Finance
Spotting risky loans, detecting suspicious payments, analyzing markets
Healthcare
Studying which treatments work best, planning hospital beds and staff
Sports
Studying player performance, choosing game strategies
Government
Planning city traffic, taxes, public services, and population needs
Marketing
Measuring which advertisements work and who to show them to
Energy and Agriculture
Predicting electricity demand, estimating crop harvests
